xsum_factuality

Người giới thiệu:

xsum_factuality

Sử dụng lệnh sau để tải tập dữ liệu này trong TFDS:

ds = tfds.load('huggingface:xsum_factuality/xsum_factuality')
  • Sự miêu tả :
Neural abstractive summarization models are highly prone to hallucinate content that is unfaithful to the input
document. The popular metric such as ROUGE fails to show the severity of the problem. The dataset consists of
faithfulness and factuality annotations of abstractive summaries for the XSum dataset. We have crowdsourced 3 judgements
 for each of 500 x 5 document-system pairs. This will be a valuable resource to the abstractive summarization community.
Tách ra Ví dụ
'train' 5597
  • Đặc trưng :
{
    "bbcid": {
        "dtype": "int32",
        "id": null,
        "_type": "Value"
    },
    "system": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "summary": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "is_factual": {
        "num_classes": 2,
        "names": [
            "no",
            "yes"
        ],
        "names_file": null,
        "id": null,
        "_type": "ClassLabel"
    },
    "worker_id": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    }
}

xsum_trung thành

Sử dụng lệnh sau để tải tập dữ liệu này trong TFDS:

ds = tfds.load('huggingface:xsum_factuality/xsum_faithfulness')
  • Sự miêu tả :
Neural abstractive summarization models are highly prone to hallucinate content that is unfaithful to the input
document. The popular metric such as ROUGE fails to show the severity of the problem. The dataset consists of
faithfulness and factuality annotations of abstractive summaries for the XSum dataset. We have crowdsourced 3 judgements
 for each of 500 x 5 document-system pairs. This will be a valuable resource to the abstractive summarization community.
Tách ra Ví dụ
'train' 11185
  • Đặc trưng :
{
    "bbcid": {
        "dtype": "int32",
        "id": null,
        "_type": "Value"
    },
    "system": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "summary": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "hallucination_type": {
        "num_classes": 2,
        "names": [
            "intrinsic",
            "extrinsic"
        ],
        "names_file": null,
        "id": null,
        "_type": "ClassLabel"
    },
    "hallucinated_span_start": {
        "dtype": "int32",
        "id": null,
        "_type": "Value"
    },
    "hallucinated_span_end": {
        "dtype": "int32",
        "id": null,
        "_type": "Value"
    },
    "worker_id": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    }
}