Riferimenti:
en
Utilizzare il comando seguente per caricare questo set di dati in TFDS:
ds = tfds.load('huggingface:multi_eurlex/en')
- Descrizione :
MultiEURLEX comprises 65k EU laws in 23 official EU languages (some low-ish resource).
Each EU law has been annotated with EUROVOC concepts (labels) by the Publication Office of EU.
As with the English EURLEX, the goal is to predict the relevant EUROVOC concepts (labels);
this is multi-label classification task (given the text, predict multiple labels).
- Licenza : nessuna licenza conosciuta
- Versione : 1.0.0
- Divide :
Diviso | Esempi |
---|---|
'test' | 5000 |
'train' | 55000 |
'validation' | 5000 |
- Caratteristiche :
{
"celex_id": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"text": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"labels": {
"feature": {
"num_classes": 21,
"names": [
"100149",
"100160",
"100148",
"100147",
"100152",
"100143",
"100156",
"100158",
"100154",
"100153",
"100142",
"100145",
"100150",
"100162",
"100159",
"100144",
"100151",
"100157",
"100161",
"100146",
"100155"
],
"names_file": null,
"id": null,
"_type": "ClassLabel"
},
"length": -1,
"id": null,
"_type": "Sequence"
}
}
da
Utilizzare il comando seguente per caricare questo set di dati in TFDS:
ds = tfds.load('huggingface:multi_eurlex/da')
- Descrizione :
MultiEURLEX comprises 65k EU laws in 23 official EU languages (some low-ish resource).
Each EU law has been annotated with EUROVOC concepts (labels) by the Publication Office of EU.
As with the English EURLEX, the goal is to predict the relevant EUROVOC concepts (labels);
this is multi-label classification task (given the text, predict multiple labels).
- Licenza : nessuna licenza conosciuta
- Versione : 1.0.0
- Divide :
Diviso | Esempi |
---|---|
'test' | 5000 |
'train' | 55000 |
'validation' | 5000 |
- Caratteristiche :
{
"celex_id": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"text": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"labels": {
"feature": {
"num_classes": 21,
"names": [
"100149",
"100160",
"100148",
"100147",
"100152",
"100143",
"100156",
"100158",
"100154",
"100153",
"100142",
"100145",
"100150",
"100162",
"100159",
"100144",
"100151",
"100157",
"100161",
"100146",
"100155"
],
"names_file": null,
"id": null,
"_type": "ClassLabel"
},
"length": -1,
"id": null,
"_type": "Sequence"
}
}
de
Utilizzare il comando seguente per caricare questo set di dati in TFDS:
ds = tfds.load('huggingface:multi_eurlex/de')
- Descrizione :
MultiEURLEX comprises 65k EU laws in 23 official EU languages (some low-ish resource).
Each EU law has been annotated with EUROVOC concepts (labels) by the Publication Office of EU.
As with the English EURLEX, the goal is to predict the relevant EUROVOC concepts (labels);
this is multi-label classification task (given the text, predict multiple labels).
- Licenza : nessuna licenza conosciuta
- Versione : 1.0.0
- Divide :
Diviso | Esempi |
---|---|
'test' | 5000 |
'train' | 55000 |
'validation' | 5000 |
- Caratteristiche :
{
"celex_id": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"text": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"labels": {
"feature": {
"num_classes": 21,
"names": [
"100149",
"100160",
"100148",
"100147",
"100152",
"100143",
"100156",
"100158",
"100154",
"100153",
"100142",
"100145",
"100150",
"100162",
"100159",
"100144",
"100151",
"100157",
"100161",
"100146",
"100155"
],
"names_file": null,
"id": null,
"_type": "ClassLabel"
},
"length": -1,
"id": null,
"_type": "Sequence"
}
}
n.l
Utilizzare il comando seguente per caricare questo set di dati in TFDS:
ds = tfds.load('huggingface:multi_eurlex/nl')
- Descrizione :
MultiEURLEX comprises 65k EU laws in 23 official EU languages (some low-ish resource).
Each EU law has been annotated with EUROVOC concepts (labels) by the Publication Office of EU.
As with the English EURLEX, the goal is to predict the relevant EUROVOC concepts (labels);
this is multi-label classification task (given the text, predict multiple labels).
- Licenza : nessuna licenza conosciuta
- Versione : 1.0.0
- Divide :
Diviso | Esempi |
---|---|
'test' | 5000 |
'train' | 55000 |
'validation' | 5000 |
- Caratteristiche :
{
"celex_id": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"text": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"labels": {
"feature": {
"num_classes": 21,
"names": [
"100149",
"100160",
"100148",
"100147",
"100152",
"100143",
"100156",
"100158",
"100154",
"100153",
"100142",
"100145",
"100150",
"100162",
"100159",
"100144",
"100151",
"100157",
"100161",
"100146",
"100155"
],
"names_file": null,
"id": null,
"_type": "ClassLabel"
},
"length": -1,
"id": null,
"_type": "Sequence"
}
}
sv
Utilizzare il comando seguente per caricare questo set di dati in TFDS:
ds = tfds.load('huggingface:multi_eurlex/sv')
- Descrizione :
MultiEURLEX comprises 65k EU laws in 23 official EU languages (some low-ish resource).
Each EU law has been annotated with EUROVOC concepts (labels) by the Publication Office of EU.
As with the English EURLEX, the goal is to predict the relevant EUROVOC concepts (labels);
this is multi-label classification task (given the text, predict multiple labels).
- Licenza : nessuna licenza conosciuta
- Versione : 1.0.0
- Divide :
Diviso | Esempi |
---|---|
'test' | 5000 |
'train' | 42490 |
'validation' | 5000 |
- Caratteristiche :
{
"celex_id": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"text": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"labels": {
"feature": {
"num_classes": 21,
"names": [
"100149",
"100160",
"100148",
"100147",
"100152",
"100143",
"100156",
"100158",
"100154",
"100153",
"100142",
"100145",
"100150",
"100162",
"100159",
"100144",
"100151",
"100157",
"100161",
"100146",
"100155"
],
"names_file": null,
"id": null,
"_type": "ClassLabel"
},
"length": -1,
"id": null,
"_type": "Sequence"
}
}
bg
Utilizzare il comando seguente per caricare questo set di dati in TFDS:
ds = tfds.load('huggingface:multi_eurlex/bg')
- Descrizione :
MultiEURLEX comprises 65k EU laws in 23 official EU languages (some low-ish resource).
Each EU law has been annotated with EUROVOC concepts (labels) by the Publication Office of EU.
As with the English EURLEX, the goal is to predict the relevant EUROVOC concepts (labels);
this is multi-label classification task (given the text, predict multiple labels).
- Licenza : nessuna licenza conosciuta
- Versione : 1.0.0
- Divide :
Diviso | Esempi |
---|---|
'test' | 5000 |
'train' | 15986 |
'validation' | 5000 |
- Caratteristiche :
{
"celex_id": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"text": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"labels": {
"feature": {
"num_classes": 21,
"names": [
"100149",
"100160",
"100148",
"100147",
"100152",
"100143",
"100156",
"100158",
"100154",
"100153",
"100142",
"100145",
"100150",
"100162",
"100159",
"100144",
"100151",
"100157",
"100161",
"100146",
"100155"
],
"names_file": null,
"id": null,
"_type": "ClassLabel"
},
"length": -1,
"id": null,
"_type": "Sequence"
}
}
c.s
Utilizzare il comando seguente per caricare questo set di dati in TFDS:
ds = tfds.load('huggingface:multi_eurlex/cs')
- Descrizione :
MultiEURLEX comprises 65k EU laws in 23 official EU languages (some low-ish resource).
Each EU law has been annotated with EUROVOC concepts (labels) by the Publication Office of EU.
As with the English EURLEX, the goal is to predict the relevant EUROVOC concepts (labels);
this is multi-label classification task (given the text, predict multiple labels).
- Licenza : nessuna licenza conosciuta
- Versione : 1.0.0
- Divide :
Diviso | Esempi |
---|---|
'test' | 5000 |
'train' | 23187 |
'validation' | 5000 |
- Caratteristiche :
{
"celex_id": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"text": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"labels": {
"feature": {
"num_classes": 21,
"names": [
"100149",
"100160",
"100148",
"100147",
"100152",
"100143",
"100156",
"100158",
"100154",
"100153",
"100142",
"100145",
"100150",
"100162",
"100159",
"100144",
"100151",
"100157",
"100161",
"100146",
"100155"
],
"names_file": null,
"id": null,
"_type": "ClassLabel"
},
"length": -1,
"id": null,
"_type": "Sequence"
}
}
ora
Utilizzare il comando seguente per caricare questo set di dati in TFDS:
ds = tfds.load('huggingface:multi_eurlex/hr')
- Descrizione :
MultiEURLEX comprises 65k EU laws in 23 official EU languages (some low-ish resource).
Each EU law has been annotated with EUROVOC concepts (labels) by the Publication Office of EU.
As with the English EURLEX, the goal is to predict the relevant EUROVOC concepts (labels);
this is multi-label classification task (given the text, predict multiple labels).
- Licenza : nessuna licenza conosciuta
- Versione : 1.0.0
- Divide :
Diviso | Esempi |
---|---|
'test' | 5000 |
'train' | 7944 |
'validation' | 2500 |
- Caratteristiche :
{
"celex_id": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"text": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"labels": {
"feature": {
"num_classes": 21,
"names": [
"100149",
"100160",
"100148",
"100147",
"100152",
"100143",
"100156",
"100158",
"100154",
"100153",
"100142",
"100145",
"100150",
"100162",
"100159",
"100144",
"100151",
"100157",
"100161",
"100146",
"100155"
],
"names_file": null,
"id": null,
"_type": "ClassLabel"
},
"length": -1,
"id": null,
"_type": "Sequence"
}
}
pl
Utilizzare il comando seguente per caricare questo set di dati in TFDS:
ds = tfds.load('huggingface:multi_eurlex/pl')
- Descrizione :
MultiEURLEX comprises 65k EU laws in 23 official EU languages (some low-ish resource).
Each EU law has been annotated with EUROVOC concepts (labels) by the Publication Office of EU.
As with the English EURLEX, the goal is to predict the relevant EUROVOC concepts (labels);
this is multi-label classification task (given the text, predict multiple labels).
- Licenza : nessuna licenza conosciuta
- Versione : 1.0.0
- Divide :
Diviso | Esempi |
---|---|
'test' | 5000 |
'train' | 23197 |
'validation' | 5000 |
- Caratteristiche :
{
"celex_id": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"text": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"labels": {
"feature": {
"num_classes": 21,
"names": [
"100149",
"100160",
"100148",
"100147",
"100152",
"100143",
"100156",
"100158",
"100154",
"100153",
"100142",
"100145",
"100150",
"100162",
"100159",
"100144",
"100151",
"100157",
"100161",
"100146",
"100155"
],
"names_file": null,
"id": null,
"_type": "ClassLabel"
},
"length": -1,
"id": null,
"_type": "Sequence"
}
}
sc
Utilizzare il comando seguente per caricare questo set di dati in TFDS:
ds = tfds.load('huggingface:multi_eurlex/sk')
- Descrizione :
MultiEURLEX comprises 65k EU laws in 23 official EU languages (some low-ish resource).
Each EU law has been annotated with EUROVOC concepts (labels) by the Publication Office of EU.
As with the English EURLEX, the goal is to predict the relevant EUROVOC concepts (labels);
this is multi-label classification task (given the text, predict multiple labels).
- Licenza : nessuna licenza conosciuta
- Versione : 1.0.0
- Divide :
Diviso | Esempi |
---|---|
'test' | 5000 |
'train' | 22971 |
'validation' | 5000 |
- Caratteristiche :
{
"celex_id": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"text": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"labels": {
"feature": {
"num_classes": 21,
"names": [
"100149",
"100160",
"100148",
"100147",
"100152",
"100143",
"100156",
"100158",
"100154",
"100153",
"100142",
"100145",
"100150",
"100162",
"100159",
"100144",
"100151",
"100157",
"100161",
"100146",
"100155"
],
"names_file": null,
"id": null,
"_type": "ClassLabel"
},
"length": -1,
"id": null,
"_type": "Sequence"
}
}
sl
Utilizzare il comando seguente per caricare questo set di dati in TFDS:
ds = tfds.load('huggingface:multi_eurlex/sl')
- Descrizione :
MultiEURLEX comprises 65k EU laws in 23 official EU languages (some low-ish resource).
Each EU law has been annotated with EUROVOC concepts (labels) by the Publication Office of EU.
As with the English EURLEX, the goal is to predict the relevant EUROVOC concepts (labels);
this is multi-label classification task (given the text, predict multiple labels).
- Licenza : nessuna licenza conosciuta
- Versione : 1.0.0
- Divide :
Diviso | Esempi |
---|---|
'test' | 5000 |
'train' | 23184 |
'validation' | 5000 |
- Caratteristiche :
{
"celex_id": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"text": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"labels": {
"feature": {
"num_classes": 21,
"names": [
"100149",
"100160",
"100148",
"100147",
"100152",
"100143",
"100156",
"100158",
"100154",
"100153",
"100142",
"100145",
"100150",
"100162",
"100159",
"100144",
"100151",
"100157",
"100161",
"100146",
"100155"
],
"names_file": null,
"id": null,
"_type": "ClassLabel"
},
"length": -1,
"id": null,
"_type": "Sequence"
}
}
es
Utilizzare il comando seguente per caricare questo set di dati in TFDS:
ds = tfds.load('huggingface:multi_eurlex/es')
- Descrizione :
MultiEURLEX comprises 65k EU laws in 23 official EU languages (some low-ish resource).
Each EU law has been annotated with EUROVOC concepts (labels) by the Publication Office of EU.
As with the English EURLEX, the goal is to predict the relevant EUROVOC concepts (labels);
this is multi-label classification task (given the text, predict multiple labels).
- Licenza : nessuna licenza conosciuta
- Versione : 1.0.0
- Divide :
Diviso | Esempi |
---|---|
'test' | 5000 |
'train' | 52785 |
'validation' | 5000 |
- Caratteristiche :
{
"celex_id": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"text": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"labels": {
"feature": {
"num_classes": 21,
"names": [
"100149",
"100160",
"100148",
"100147",
"100152",
"100143",
"100156",
"100158",
"100154",
"100153",
"100142",
"100145",
"100150",
"100162",
"100159",
"100144",
"100151",
"100157",
"100161",
"100146",
"100155"
],
"names_file": null,
"id": null,
"_type": "ClassLabel"
},
"length": -1,
"id": null,
"_type": "Sequence"
}
}
fr
Utilizzare il comando seguente per caricare questo set di dati in TFDS:
ds = tfds.load('huggingface:multi_eurlex/fr')
- Descrizione :
MultiEURLEX comprises 65k EU laws in 23 official EU languages (some low-ish resource).
Each EU law has been annotated with EUROVOC concepts (labels) by the Publication Office of EU.
As with the English EURLEX, the goal is to predict the relevant EUROVOC concepts (labels);
this is multi-label classification task (given the text, predict multiple labels).
- Licenza : nessuna licenza conosciuta
- Versione : 1.0.0
- Divide :
Diviso | Esempi |
---|---|
'test' | 5000 |
'train' | 55000 |
'validation' | 5000 |
- Caratteristiche :
{
"celex_id": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"text": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"labels": {
"feature": {
"num_classes": 21,
"names": [
"100149",
"100160",
"100148",
"100147",
"100152",
"100143",
"100156",
"100158",
"100154",
"100153",
"100142",
"100145",
"100150",
"100162",
"100159",
"100144",
"100151",
"100157",
"100161",
"100146",
"100155"
],
"names_file": null,
"id": null,
"_type": "ClassLabel"
},
"length": -1,
"id": null,
"_type": "Sequence"
}
}
Esso
Utilizzare il comando seguente per caricare questo set di dati in TFDS:
ds = tfds.load('huggingface:multi_eurlex/it')
- Descrizione :
MultiEURLEX comprises 65k EU laws in 23 official EU languages (some low-ish resource).
Each EU law has been annotated with EUROVOC concepts (labels) by the Publication Office of EU.
As with the English EURLEX, the goal is to predict the relevant EUROVOC concepts (labels);
this is multi-label classification task (given the text, predict multiple labels).
- Licenza : nessuna licenza conosciuta
- Versione : 1.0.0
- Divide :
Diviso | Esempi |
---|---|
'test' | 5000 |
'train' | 55000 |
'validation' | 5000 |
- Caratteristiche :
{
"celex_id": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"text": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"labels": {
"feature": {
"num_classes": 21,
"names": [
"100149",
"100160",
"100148",
"100147",
"100152",
"100143",
"100156",
"100158",
"100154",
"100153",
"100142",
"100145",
"100150",
"100162",
"100159",
"100144",
"100151",
"100157",
"100161",
"100146",
"100155"
],
"names_file": null,
"id": null,
"_type": "ClassLabel"
},
"length": -1,
"id": null,
"_type": "Sequence"
}
}
punto
Utilizzare il comando seguente per caricare questo set di dati in TFDS:
ds = tfds.load('huggingface:multi_eurlex/pt')
- Descrizione :
MultiEURLEX comprises 65k EU laws in 23 official EU languages (some low-ish resource).
Each EU law has been annotated with EUROVOC concepts (labels) by the Publication Office of EU.
As with the English EURLEX, the goal is to predict the relevant EUROVOC concepts (labels);
this is multi-label classification task (given the text, predict multiple labels).
- Licenza : nessuna licenza conosciuta
- Versione : 1.0.0
- Divide :
Diviso | Esempi |
---|---|
'test' | 5000 |
'train' | 52370 |
'validation' | 5000 |
- Caratteristiche :
{
"celex_id": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"text": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"labels": {
"feature": {
"num_classes": 21,
"names": [
"100149",
"100160",
"100148",
"100147",
"100152",
"100143",
"100156",
"100158",
"100154",
"100153",
"100142",
"100145",
"100150",
"100162",
"100159",
"100144",
"100151",
"100157",
"100161",
"100146",
"100155"
],
"names_file": null,
"id": null,
"_type": "ClassLabel"
},
"length": -1,
"id": null,
"_type": "Sequence"
}
}
ro
Utilizzare il comando seguente per caricare questo set di dati in TFDS:
ds = tfds.load('huggingface:multi_eurlex/ro')
- Descrizione :
MultiEURLEX comprises 65k EU laws in 23 official EU languages (some low-ish resource).
Each EU law has been annotated with EUROVOC concepts (labels) by the Publication Office of EU.
As with the English EURLEX, the goal is to predict the relevant EUROVOC concepts (labels);
this is multi-label classification task (given the text, predict multiple labels).
- Licenza : nessuna licenza conosciuta
- Versione : 1.0.0
- Divide :
Diviso | Esempi |
---|---|
'test' | 5000 |
'train' | 15921 |
'validation' | 5000 |
- Caratteristiche :
{
"celex_id": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"text": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"labels": {
"feature": {
"num_classes": 21,
"names": [
"100149",
"100160",
"100148",
"100147",
"100152",
"100143",
"100156",
"100158",
"100154",
"100153",
"100142",
"100145",
"100150",
"100162",
"100159",
"100144",
"100151",
"100157",
"100161",
"100146",
"100155"
],
"names_file": null,
"id": null,
"_type": "ClassLabel"
},
"length": -1,
"id": null,
"_type": "Sequence"
}
}
et
Utilizzare il comando seguente per caricare questo set di dati in TFDS:
ds = tfds.load('huggingface:multi_eurlex/et')
- Descrizione :
MultiEURLEX comprises 65k EU laws in 23 official EU languages (some low-ish resource).
Each EU law has been annotated with EUROVOC concepts (labels) by the Publication Office of EU.
As with the English EURLEX, the goal is to predict the relevant EUROVOC concepts (labels);
this is multi-label classification task (given the text, predict multiple labels).
- Licenza : nessuna licenza conosciuta
- Versione : 1.0.0
- Divide :
Diviso | Esempi |
---|---|
'test' | 5000 |
'train' | 23126 |
'validation' | 5000 |
- Caratteristiche :
{
"celex_id": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"text": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"labels": {
"feature": {
"num_classes": 21,
"names": [
"100149",
"100160",
"100148",
"100147",
"100152",
"100143",
"100156",
"100158",
"100154",
"100153",
"100142",
"100145",
"100150",
"100162",
"100159",
"100144",
"100151",
"100157",
"100161",
"100146",
"100155"
],
"names_file": null,
"id": null,
"_type": "ClassLabel"
},
"length": -1,
"id": null,
"_type": "Sequence"
}
}
fi
Utilizzare il comando seguente per caricare questo set di dati in TFDS:
ds = tfds.load('huggingface:multi_eurlex/fi')
- Descrizione :
MultiEURLEX comprises 65k EU laws in 23 official EU languages (some low-ish resource).
Each EU law has been annotated with EUROVOC concepts (labels) by the Publication Office of EU.
As with the English EURLEX, the goal is to predict the relevant EUROVOC concepts (labels);
this is multi-label classification task (given the text, predict multiple labels).
- Licenza : nessuna licenza conosciuta
- Versione : 1.0.0
- Divide :
Diviso | Esempi |
---|---|
'test' | 5000 |
'train' | 42497 |
'validation' | 5000 |
- Caratteristiche :
{
"celex_id": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"text": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"labels": {
"feature": {
"num_classes": 21,
"names": [
"100149",
"100160",
"100148",
"100147",
"100152",
"100143",
"100156",
"100158",
"100154",
"100153",
"100142",
"100145",
"100150",
"100162",
"100159",
"100144",
"100151",
"100157",
"100161",
"100146",
"100155"
],
"names_file": null,
"id": null,
"_type": "ClassLabel"
},
"length": -1,
"id": null,
"_type": "Sequence"
}
}
eh
Utilizzare il comando seguente per caricare questo set di dati in TFDS:
ds = tfds.load('huggingface:multi_eurlex/hu')
- Descrizione :
MultiEURLEX comprises 65k EU laws in 23 official EU languages (some low-ish resource).
Each EU law has been annotated with EUROVOC concepts (labels) by the Publication Office of EU.
As with the English EURLEX, the goal is to predict the relevant EUROVOC concepts (labels);
this is multi-label classification task (given the text, predict multiple labels).
- Licenza : nessuna licenza conosciuta
- Versione : 1.0.0
- Divide :
Diviso | Esempi |
---|---|
'test' | 5000 |
'train' | 22664 |
'validation' | 5000 |
- Caratteristiche :
{
"celex_id": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"text": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"labels": {
"feature": {
"num_classes": 21,
"names": [
"100149",
"100160",
"100148",
"100147",
"100152",
"100143",
"100156",
"100158",
"100154",
"100153",
"100142",
"100145",
"100150",
"100162",
"100159",
"100144",
"100151",
"100157",
"100161",
"100146",
"100155"
],
"names_file": null,
"id": null,
"_type": "ClassLabel"
},
"length": -1,
"id": null,
"_type": "Sequence"
}
}
lt
Utilizzare il comando seguente per caricare questo set di dati in TFDS:
ds = tfds.load('huggingface:multi_eurlex/lt')
- Descrizione :
MultiEURLEX comprises 65k EU laws in 23 official EU languages (some low-ish resource).
Each EU law has been annotated with EUROVOC concepts (labels) by the Publication Office of EU.
As with the English EURLEX, the goal is to predict the relevant EUROVOC concepts (labels);
this is multi-label classification task (given the text, predict multiple labels).
- Licenza : nessuna licenza conosciuta
- Versione : 1.0.0
- Divide :
Diviso | Esempi |
---|---|
'test' | 5000 |
'train' | 23188 |
'validation' | 5000 |
- Caratteristiche :
{
"celex_id": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"text": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"labels": {
"feature": {
"num_classes": 21,
"names": [
"100149",
"100160",
"100148",
"100147",
"100152",
"100143",
"100156",
"100158",
"100154",
"100153",
"100142",
"100145",
"100150",
"100162",
"100159",
"100144",
"100151",
"100157",
"100161",
"100146",
"100155"
],
"names_file": null,
"id": null,
"_type": "ClassLabel"
},
"length": -1,
"id": null,
"_type": "Sequence"
}
}
lv
Utilizzare il comando seguente per caricare questo set di dati in TFDS:
ds = tfds.load('huggingface:multi_eurlex/lv')
- Descrizione :
MultiEURLEX comprises 65k EU laws in 23 official EU languages (some low-ish resource).
Each EU law has been annotated with EUROVOC concepts (labels) by the Publication Office of EU.
As with the English EURLEX, the goal is to predict the relevant EUROVOC concepts (labels);
this is multi-label classification task (given the text, predict multiple labels).
- Licenza : nessuna licenza conosciuta
- Versione : 1.0.0
- Divide :
Diviso | Esempi |
---|---|
'test' | 5000 |
'train' | 23208 |
'validation' | 5000 |
- Caratteristiche :
{
"celex_id": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"text": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"labels": {
"feature": {
"num_classes": 21,
"names": [
"100149",
"100160",
"100148",
"100147",
"100152",
"100143",
"100156",
"100158",
"100154",
"100153",
"100142",
"100145",
"100150",
"100162",
"100159",
"100144",
"100151",
"100157",
"100161",
"100146",
"100155"
],
"names_file": null,
"id": null,
"_type": "ClassLabel"
},
"length": -1,
"id": null,
"_type": "Sequence"
}
}
el
Utilizzare il comando seguente per caricare questo set di dati in TFDS:
ds = tfds.load('huggingface:multi_eurlex/el')
- Descrizione :
MultiEURLEX comprises 65k EU laws in 23 official EU languages (some low-ish resource).
Each EU law has been annotated with EUROVOC concepts (labels) by the Publication Office of EU.
As with the English EURLEX, the goal is to predict the relevant EUROVOC concepts (labels);
this is multi-label classification task (given the text, predict multiple labels).
- Licenza : nessuna licenza conosciuta
- Versione : 1.0.0
- Divide :
Diviso | Esempi |
---|---|
'test' | 5000 |
'train' | 55000 |
'validation' | 5000 |
- Caratteristiche :
{
"celex_id": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"text": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"labels": {
"feature": {
"num_classes": 21,
"names": [
"100149",
"100160",
"100148",
"100147",
"100152",
"100143",
"100156",
"100158",
"100154",
"100153",
"100142",
"100145",
"100150",
"100162",
"100159",
"100144",
"100151",
"100157",
"100161",
"100146",
"100155"
],
"names_file": null,
"id": null,
"_type": "ClassLabel"
},
"length": -1,
"id": null,
"_type": "Sequence"
}
}
mt
Utilizzare il comando seguente per caricare questo set di dati in TFDS:
ds = tfds.load('huggingface:multi_eurlex/mt')
- Descrizione :
MultiEURLEX comprises 65k EU laws in 23 official EU languages (some low-ish resource).
Each EU law has been annotated with EUROVOC concepts (labels) by the Publication Office of EU.
As with the English EURLEX, the goal is to predict the relevant EUROVOC concepts (labels);
this is multi-label classification task (given the text, predict multiple labels).
- Licenza : nessuna licenza conosciuta
- Versione : 1.0.0
- Divide :
Diviso | Esempi |
---|---|
'test' | 5000 |
'train' | 17521 |
'validation' | 5000 |
- Caratteristiche :
{
"celex_id": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"text": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"labels": {
"feature": {
"num_classes": 21,
"names": [
"100149",
"100160",
"100148",
"100147",
"100152",
"100143",
"100156",
"100158",
"100154",
"100153",
"100142",
"100145",
"100150",
"100162",
"100159",
"100144",
"100151",
"100157",
"100161",
"100146",
"100155"
],
"names_file": null,
"id": null,
"_type": "ClassLabel"
},
"length": -1,
"id": null,
"_type": "Sequence"
}
}
tutte_lingue
Utilizzare il comando seguente per caricare questo set di dati in TFDS:
ds = tfds.load('huggingface:multi_eurlex/all_languages')
- Descrizione :
MultiEURLEX comprises 65k EU laws in 23 official EU languages (some low-ish resource).
Each EU law has been annotated with EUROVOC concepts (labels) by the Publication Office of EU.
As with the English EURLEX, the goal is to predict the relevant EUROVOC concepts (labels);
this is multi-label classification task (given the text, predict multiple labels).
- Licenza : nessuna licenza conosciuta
- Versione : 1.0.0
- Divide :
Diviso | Esempi |
---|---|
'test' | 5000 |
'train' | 55000 |
'validation' | 5000 |
- Caratteristiche :
{
"celex_id": {
"dtype": "string",
"id": null,
"_type": "Value"
},
"text": {
"languages": [
"en",
"da",
"de",
"nl",
"sv",
"bg",
"cs",
"hr",
"pl",
"sk",
"sl",
"es",
"fr",
"it",
"pt",
"ro",
"et",
"fi",
"hu",
"lt",
"lv",
"el",
"mt"
],
"id": null,
"_type": "Translation"
},
"labels": {
"feature": {
"num_classes": 21,
"names": [
"100149",
"100160",
"100148",
"100147",
"100152",
"100143",
"100156",
"100158",
"100154",
"100153",
"100142",
"100145",
"100150",
"100162",
"100159",
"100144",
"100151",
"100157",
"100161",
"100146",
"100155"
],
"names_file": null,
"id": null,
"_type": "ClassLabel"
},
"length": -1,
"id": null,
"_type": "Sequence"
}
}