COMPRISE_Data10_MENYO-20K_V1.0

Description

A multi-domain parallel corpus for English-Yoruba language pair that can be used to benchmark machine translation systems. The dataset has 20,100 parallel sentences split into 10,070 train, 3,397 dev, and 6,633 test sentences.

Resources

Name Format Description Link
0 http://data.europa.eu/88u/dataset/oai-zenodo-org-5997745
0 http://data.europa.eu/88u/dataset/oai-zenodo-org-5997745

Tags

Topics

Categories