TFDS now supports the Croissant 🥐 format! Read the documentation to know more.

opus_ubuntu

References:

as-bs

Use the following command to load this dataset in TFDS:

ds = tfds.load('huggingface:opus_ubuntu/as-bs')

Description:

A parallel corpus of Ubuntu localization files. Source: https://translations.launchpad.net
244 languages, 23,988 bitexts
total number of files: 30,959
total number of tokens: 29.84M
total number of sentence fragments: 7.73M

License: No known license
Version: 1.0.0
Splits:

Split	Examples
`'train'`	8583

Features:

{
    "id": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "translation": {
        "languages": [
            "as",
            "bs"
        ],
        "id": null,
        "_type": "Translation"
    }
}

az-cs

Use the following command to load this dataset in TFDS:

ds = tfds.load('huggingface:opus_ubuntu/az-cs')

Description:

A parallel corpus of Ubuntu localization files. Source: https://translations.launchpad.net
244 languages, 23,988 bitexts
total number of files: 30,959
total number of tokens: 29.84M
total number of sentence fragments: 7.73M

License: No known license
Version: 1.0.0
Splits:

Split	Examples
`'train'`	293

Features:

{
    "id": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "translation": {
        "languages": [
            "az",
            "cs"
        ],
        "id": null,
        "_type": "Translation"
    }
}

bg-de

Use the following command to load this dataset in TFDS:

ds = tfds.load('huggingface:opus_ubuntu/bg-de')

Description:

A parallel corpus of Ubuntu localization files. Source: https://translations.launchpad.net
244 languages, 23,988 bitexts
total number of files: 30,959
total number of tokens: 29.84M
total number of sentence fragments: 7.73M

License: No known license
Version: 1.0.0
Splits:

Split	Examples
`'train'`	184

Features:

{
    "id": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "translation": {
        "languages": [
            "bg",
            "de"
        ],
        "id": null,
        "_type": "Translation"
    }
}

br-es_PR

Use the following command to load this dataset in TFDS:

ds = tfds.load('huggingface:opus_ubuntu/br-es_PR')

Description:

A parallel corpus of Ubuntu localization files. Source: https://translations.launchpad.net
244 languages, 23,988 bitexts
total number of files: 30,959
total number of tokens: 29.84M
total number of sentence fragments: 7.73M

License: No known license
Version: 1.0.0
Splits:

Split	Examples
`'train'`	125

Features:

{
    "id": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "translation": {
        "languages": [
            "br",
            "es_PR"
        ],
        "id": null,
        "_type": "Translation"
    }
}

bn-ga

Use the following command to load this dataset in TFDS:

ds = tfds.load('huggingface:opus_ubuntu/bn-ga')

Description:

A parallel corpus of Ubuntu localization files. Source: https://translations.launchpad.net
244 languages, 23,988 bitexts
total number of files: 30,959
total number of tokens: 29.84M
total number of sentence fragments: 7.73M

License: No known license
Version: 1.0.0
Splits:

Split	Examples
`'train'`	7324

Features:

{
    "id": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "translation": {
        "languages": [
            "bn",
            "ga"
        ],
        "id": null,
        "_type": "Translation"
    }
}

br-hi

Use the following command to load this dataset in TFDS:

ds = tfds.load('huggingface:opus_ubuntu/br-hi')

Description:

A parallel corpus of Ubuntu localization files. Source: https://translations.launchpad.net
244 languages, 23,988 bitexts
total number of files: 30,959
total number of tokens: 29.84M
total number of sentence fragments: 7.73M

License: No known license
Version: 1.0.0
Splits:

Split	Examples
`'train'`	15551

Features:

{
    "id": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "translation": {
        "languages": [
            "br",
            "hi"
        ],
        "id": null,
        "_type": "Translation"
    }
}

br-la

Use the following command to load this dataset in TFDS:

ds = tfds.load('huggingface:opus_ubuntu/br-la')

Description:

A parallel corpus of Ubuntu localization files. Source: https://translations.launchpad.net
244 languages, 23,988 bitexts
total number of files: 30,959
total number of tokens: 29.84M
total number of sentence fragments: 7.73M

License: No known license
Version: 1.0.0
Splits:

Split	Examples
`'train'`	527

Features:

{
    "id": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "translation": {
        "languages": [
            "br",
            "la"
        ],
        "id": null,
        "_type": "Translation"
    }
}

bs-szl

Use the following command to load this dataset in TFDS:

ds = tfds.load('huggingface:opus_ubuntu/bs-szl')

Description:

A parallel corpus of Ubuntu localization files. Source: https://translations.launchpad.net
244 languages, 23,988 bitexts
total number of files: 30,959
total number of tokens: 29.84M
total number of sentence fragments: 7.73M

License: No known license
Version: 1.0.0
Splits:

Split	Examples
`'train'`	646

Features:

{
    "id": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "translation": {
        "languages": [
            "bs",
            "szl"
        ],
        "id": null,
        "_type": "Translation"
    }
}

br-uz

Use the following command to load this dataset in TFDS:

ds = tfds.load('huggingface:opus_ubuntu/br-uz')

Description:

A parallel corpus of Ubuntu localization files. Source: https://translations.launchpad.net
244 languages, 23,988 bitexts
total number of files: 30,959
total number of tokens: 29.84M
total number of sentence fragments: 7.73M

License: No known license
Version: 1.0.0
Splits:

Split	Examples
`'train'`	1416

Features:

{
    "id": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "translation": {
        "languages": [
            "br",
            "uz"
        ],
        "id": null,
        "_type": "Translation"
    }
}

br-yi

Use the following command to load this dataset in TFDS:

ds = tfds.load('huggingface:opus_ubuntu/br-yi')

Description:

A parallel corpus of Ubuntu localization files. Source: https://translations.launchpad.net
244 languages, 23,988 bitexts
total number of files: 30,959
total number of tokens: 29.84M
total number of sentence fragments: 7.73M

License: No known license
Version: 1.0.0
Splits:

Split	Examples
`'train'`	2799

Features:

{
    "id": {
        "dtype": "string",
        "id": null,
        "_type": "Value"
    },
    "translation": {
        "languages": [
            "br",
            "yi"
        ],
        "id": null,
        "_type": "Translation"
    }
}