Replication data for: "The crystallization of language over time"

Description

This repository contains the data and R script accompanying the paper "The crystallization of language over time" (under review). The datasets are stored in txt (tab-delimited) format in the /data folder. The main datasets used for the analyses are ngrams_lemma.txt and ngrams_pos.txt, which contain lemma and part-of-speech trigrams and frequency information culled from the C-CLAMP corpus (1850-1999; Piersoul et al. 2021). The file trigrams_through_time.R contains the R code used to compute the trigrams' collocational strength and entropy and statistically model their evolution through time.

Resources

Name Format Description Link

Tags

  • corpus-linguistics
  • n-grams
  • mixed-models
  • entropy
  • collocational-strength-(δp)

Topics

  • SOCI
  • TECH

Categories