NST N-gram - Norwegian Bokmål

Description

These n-grams are derived from parts of the Text Corpus from Nordic Language Technology AS (NST). The source material consists of 510 million words of running text. The n-grams are also available as an overview listing only the 1000 most frequent n-grams (n=1-6). In the full version, all the derived n-grams (n=1-6) are sorted alphabetically and by frequency, respectively. Frequency lists (unigrams) are also available separately.

Resources

Name Format Description Link
0 http://data.europa.eu/88u/dataset/sbr-3

Tags

  • språkbanken
  • korpus
  • språkteknologi
  • språkforskning
  • ngram

Topics

  • TECH

Categories