COMPRISE_Data13_Y-BBCTopics_V1.0
Description
The dataset contains the Yoruba data used to evaluate the performance of different weakly supervised learning techniques for text classification in low-resourced languages. The text was collected from the BBC news website for Yoruba and annotated by two native speakers. The data has seven categories based on the news headlines: Sports, Entertainment, Nigeria, Africa, World, Health, and Politics. It contains 1,908 sentences.
Resources
| Name |
Format |
Description |
Link |
|
0 |
|
http://data.europa.eu/88u/dataset/oai-zenodo-org-5998515 |
|
0 |
|
http://data.europa.eu/88u/dataset/oai-zenodo-org-5998515 |
Tags
- bbc
- news-topics
- yoruba
- text-classification