ubuntu_dialogs_corpus

Ссылки:

тренироваться

Используйте следующую команду, чтобы загрузить этот набор данных в TFDS:

ds = tfds.load('huggingface:ubuntu_dialogs_corpus/train')
  • Описание :
Ubuntu Dialogue Corpus, a dataset containing almost 1 million multi-turn dialogues, with a total of over 7 million utterances and 100 million words. This provides a unique resource for research into building dialogue managers based on neural language models that can make use of large amounts of unlabeled data. The dataset has both the multi-turn property of conversations in the Dialog State Tracking Challenge datasets, and the unstructured nature of interactions from microblog services such as Twitter.
  • Лицензия : Нет известной лицензии.
  • Версия : 2.0.0
  • Расколы :
Расколоть Примеры
'train' 1000000
  • Функции :
{
    "Context": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "Utterance": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "Label": {
        "dtype": "int32",
        "id": null,
        "_type": "Value"
    }
}

dev_test

Используйте следующую команду, чтобы загрузить этот набор данных в TFDS:

ds = tfds.load('huggingface:ubuntu_dialogs_corpus/dev_test')
  • Описание :
Ubuntu Dialogue Corpus, a dataset containing almost 1 million multi-turn dialogues, with a total of over 7 million utterances and 100 million words. This provides a unique resource for research into building dialogue managers based on neural language models that can make use of large amounts of unlabeled data. The dataset has both the multi-turn property of conversations in the Dialog State Tracking Challenge datasets, and the unstructured nature of interactions from microblog services such as Twitter.
  • Лицензия : Нет известной лицензии.
  • Версия : 2.0.0
  • Расколы :
Расколоть Примеры
'test' 18920
'validation' 19560
  • Функции :
{
    "Context": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "Ground Truth Utterance": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "Distractor_0": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "Distractor_1": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "Distractor_2": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "Distractor_3": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "Distractor_4": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "Distractor_5": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "Distractor_6": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "Distractor_7": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "Distractor_8": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    }
}