Spanish-German website parallel corpus (Processed)
Description
This is a parallel corpus of bilingual texts crawled from multilingual websites, which contains 2,840 TUs.
Period of crawling : 15/11/2016 - 23/01/2017.
A strict validation process was already followed for the source data, which resulted in discarding:
- TUs from crawled websites that do not comply to the PSI directive,
- TUs with more than 99% of mispelled tokens,
- TUs identified during the manual validation process and all the TUs from websites which error rate in the sample extracted for manual validation are strictly above the following thresholds:
50% of TUs with language identification errors,
50% of TUs with alignment errors,
50% of TUs with tokenization errors,
20% of TUs identified as machine translated content,
50% of TUs with translation errors.
This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated Translation (CEF.AT) actions SMART 2014/1074 and SMART 2015/1091. For further information on the project: http://lr-coordination.eu.
Resources
| Name |
Format |
Description |
Link |
|
57 |
|
https://elrc-share.eu/repository/browse/spanish-german-website-parallel-corpus-processed/552ec3ba313211e9a4d400155d0267062865c204b7664909b92f968ce897fc50/ |
Tags
- group-resources-for-language-technologies