e2e_cleaned

설명 :

정리된 MR이 포함된 E2E NLG Challenge 데이터의 업데이트 릴리스입니다. E2E 데이터는 레스토랑 도메인의 화행 기반 MR(의미 표현)과 예측해야 하는 자연어 참조 최대 5개를 포함합니다.

추가 문서 : 코드가 있는 논문에서 탐색
홈페이지 : https://github.com/tuetschek/e2e-cleaning
소스 코드 : tfds.datasets.e2e_cleaned.Builder
버전 :
- 0.1.0 (기본값): 릴리스 정보가 없습니다.
다운로드 크기 : 13.92 MiB
데이터 세트 크기 : 14.70 MiB
자동 캐시 ( 문서 ): 예
분할 :

나뉘다	예
`'test'`	4,693
`'train'`	33,525
`'validation'`	4,299

기능 구조 :

FeaturesDict({
    'input_text': FeaturesDict({
        'table': Sequence({
            'column_header': string,
            'content': string,
            'row_number': int16,
        }),
    }),
    'target_text': string,
})

기능 문서 :

특징	수업	D타입
	풍모Dict
input_text	풍모Dict
입력_텍스트/테이블	순서
input_text/테이블/column_header	텐서	끈
input_text/테이블/콘텐츠	텐서	끈
입력_텍스트/테이블/행_번호	텐서	정수16
target_text	텐서	끈

감독 키 ( as_supervised 문서 참조): ('input_text', 'target_text')
그림 ( tfds.show_examples ): 지원되지 않습니다.
예 ( tfds.as_dataframe ):

인용 :

@inproceedings{dusek-etal-2019-semantic,
    title = "Semantic Noise Matters for Neural Natural Language Generation",
    author = "Du{\v{s} }ek, Ond{\v{r} }ej  and
      Howcroft, David M.  and
      Rieser, Verena",
    booktitle = "Proceedings of the 12th International Conference on Natural Language Generation",
    month = oct # "{--}" # nov,
    year = "2019",
    address = "Tokyo, Japan",
    publisher = "Association for Computational Linguistics",
    url = "https://www.aclweb.org/anthology/W19-8652",
    doi = "10.18653/v1/W19-8652",
    pages = "421--426",
    abstract = "Neural natural language generation (NNLG) systems are known for their pathological outputs, i.e. generating text which is unrelated to the input specification. In this paper, we show the impact of semantic noise on state-of-the-art NNLG models which implement different semantic control mechanisms. We find that cleaned data can improve semantic correctness by up to 97{\%}, while maintaining fluency. We also find that the most common error is omitting information, rather than hallucination.",
}

e2e_cleaned 컬렉션을 사용해 정리하기 내 환경설정을 기준으로 콘텐츠를 저장하고 분류하세요.

e2e_cleaned