Frequency lists of lemmas and word forms of the Compound Corps of Estonian 2021

Description

The frequency lists are generated on the basis of the subcases of the Compound Corps 2021 of the Estonian language (Estonian National Corpus 2021). The following subcases are available: Web Corp 2021 (Web 2021), Wikipedia 2021 (Wikipeadia 2021), DOAJ, News feeds 2014-2021 (Feeds 2014-2021), Literature (Literature). Thus, the corpus reflects the most recent use of language. In the subcorps Newsflows 2014-2021 and Literature there is also a material that has come from previous years. Housing capacity: - 944 907 713 words - 7,756,705 different lemmas - 857,784 lemmas above the frequency limit (ipm* 0.011, equivalent to 10 or more for ENC 2021). Lemmas are untreated, which means that - uppercase and lowercase statues are unplugged; - frequencies indicate the use of the individual word (compound verbs, noun phrases, etc. are shown by parts); - There may be foreign language words; - no word type considered ('hall' A and 'hall' S are together) **. * ipm (instances per million) indicates promille, or average occurrence per million, for a lemma or word. ** For Estonian, 'lempos', or lemma+word type, is not important because the types of words are already distinguished by their external shape and In the 'grey' example, the homonymous grey+S (hall night, sports hall) would still remain one of the cows. Refer to as: Hein, Indrek 2022. Frequency lists of lemmas and word forms of the Compound Corps 2021 of the Estonian language. Estonian Language Institute. DOI: 10.15155/3-00-0000-0000-0000-08D1FL

Resources

Name Format Description Link
0 https://www.eki.ee/tarkvara/
0 https://www.eki.ee/tarkvara/

Tags

Topics

Categories