Spanish-Italian website parallel corpus (Processed)

Description

This is a parallel corpus of bilingual texts crawled from multilingual websites, which contains 3,319 TUs. Date of crawling : 23/01/2017 A strict validation process was already followed for the source data, which resulted in discarding: - TUs from crawled websites that do not comply to the PSI directive, - TUs with more than 99% of mispelled tokens, - TUs identified during the manual validation process and all the TUs from websites which error rate in the sample extracted for manual validation is strictly above the following thresholds: 50% of TUs with language identification errors, 50% of TUs with alignment errors, 50% of TUs with tokenization errors, 20% of TUs identified as machine translated content, 50% of TUs with translation errors. This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated Translation (CEF.AT) actions SMART 2014/1074 and SMART 2015/1091. For further information on the project: http://lr-coordination.eu.

Resources

Name Format Description Link
57 https://elrc-share.eu/repository/browse/spanish-italian-website-parallel-corpus-processed/2f7d17cc312b11e9a4d400155d02670640ca323472fd4f56a9656062c58c5fb7/

Tags

  • group-resources-for-language-technologies

Topics

  • GOVE

Categories