ubuntu_dialogs_corpus

Referências:

Comboio

Use o seguinte comando para carregar esse conjunto de dados no TFDS:

ds = tfds.load('huggingface:ubuntu_dialogs_corpus/train')
  • Descrição :
Ubuntu Dialogue Corpus, a dataset containing almost 1 million multi-turn dialogues, with a total of over 7 million utterances and 100 million words. This provides a unique resource for research into building dialogue managers based on neural language models that can make use of large amounts of unlabeled data. The dataset has both the multi-turn property of conversations in the Dialog State Tracking Challenge datasets, and the unstructured nature of interactions from microblog services such as Twitter.
  • Licença : Nenhuma licença conhecida
  • Versão : 2.0.0
  • Divisões :
Dividir Exemplos
'train' 1.000.000
  • Características :
{
    "Context": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "Utterance": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "Label": {
        "dtype": "int32",
        "id": null,
        "_type": "Value"
    }
}

dev_test

Use o seguinte comando para carregar esse conjunto de dados no TFDS:

ds = tfds.load('huggingface:ubuntu_dialogs_corpus/dev_test')
  • Descrição :
Ubuntu Dialogue Corpus, a dataset containing almost 1 million multi-turn dialogues, with a total of over 7 million utterances and 100 million words. This provides a unique resource for research into building dialogue managers based on neural language models that can make use of large amounts of unlabeled data. The dataset has both the multi-turn property of conversations in the Dialog State Tracking Challenge datasets, and the unstructured nature of interactions from microblog services such as Twitter.
  • Licença : Nenhuma licença conhecida
  • Versão : 2.0.0
  • Divisões :
Dividir Exemplos
'test' 18920
'validation' 19560
  • Características :
{
    "Context": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "Ground Truth Utterance": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "Distractor_0": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "Distractor_1": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "Distractor_2": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "Distractor_3": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "Distractor_4": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "Distractor_5": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "Distractor_6": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "Distractor_7": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "Distractor_8": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    }
}
,

Referências:

Comboio

Use o seguinte comando para carregar esse conjunto de dados no TFDS:

ds = tfds.load('huggingface:ubuntu_dialogs_corpus/train')
  • Descrição :
Ubuntu Dialogue Corpus, a dataset containing almost 1 million multi-turn dialogues, with a total of over 7 million utterances and 100 million words. This provides a unique resource for research into building dialogue managers based on neural language models that can make use of large amounts of unlabeled data. The dataset has both the multi-turn property of conversations in the Dialog State Tracking Challenge datasets, and the unstructured nature of interactions from microblog services such as Twitter.
  • Licença : Nenhuma licença conhecida
  • Versão : 2.0.0
  • Divisões :
Dividir Exemplos
'train' 1.000.000
  • Características :
{
    "Context": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "Utterance": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "Label": {
        "dtype": "int32",
        "id": null,
        "_type": "Value"
    }
}

dev_test

Use o seguinte comando para carregar esse conjunto de dados no TFDS:

ds = tfds.load('huggingface:ubuntu_dialogs_corpus/dev_test')
  • Descrição :
Ubuntu Dialogue Corpus, a dataset containing almost 1 million multi-turn dialogues, with a total of over 7 million utterances and 100 million words. This provides a unique resource for research into building dialogue managers based on neural language models that can make use of large amounts of unlabeled data. The dataset has both the multi-turn property of conversations in the Dialog State Tracking Challenge datasets, and the unstructured nature of interactions from microblog services such as Twitter.
  • Licença : Nenhuma licença conhecida
  • Versão : 2.0.0
  • Divisões :
Dividir Exemplos
'test' 18920
'validation' 19560
  • Características :
{
    "Context": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "Ground Truth Utterance": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "Distractor_0": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "Distractor_1": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "Distractor_2": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "Distractor_3": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "Distractor_4": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "Distractor_5": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "Distractor_6": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "Distractor_7": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "Distractor_8": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    }
}