Estonian National Corpus 2017
Description
The corps was established in cooperation between the Estonian Language Institute and Lexical Computing Ltd. The size of the corps is 1.3 billion words. The base of the corps is the Estonian Language Joint Corps 2013, which was renewed by Lexical Computing Ltd. in 2017 on the order of the Estonian Language Institute.
The sub-corps are the Estonian language corps 1990-2008, the Estonian language web corps 2013, the Estonian language web corps 2017 and the Estonian Wikipedia 2017 corps. The content of the web housing is Estonian-language websites downloaded from the Internet. The programs described at http://corpus.tools have been used to create the housing: SpederLing, JustText, Chared, Onion and wiki2corpus. The housing is lemmatized, marked and connected by the EstNLTK analyser.
Resources
| Name |
Format |
Description |
Link |
|
0 |
|
https://www.sketchengine.co.uk/ |