Simple Italian sentences ranked by readability

Description

The dataset contains 500,000 sentences extracted from the Paisà corpus (https://www.corpusitaliano.it/) which have been selected for being easy to read according to four parameters: token number, average word length, depth of the parse tree and verb "arity". The sentences are ranked by readability.

Resources

Name Format Description Link
0 http://data.europa.eu/88u/dataset/oai-zenodo-org-2548585
0 http://data.europa.eu/88u/dataset/oai-zenodo-org-2548585

Tags

Topics

Categories