Code used to produce terms list in the work "NLP-Driven Electron Microscopy Ontology Development"

Description

This is a collection of code written by Maurice Curran that was used to process the Microscopy and Microanalysis conference proceeding corpus into word products described in the publication "NLP-Driven Electron Microscopy Ontology Development". The scripts are written in Python, to be used in the following order:1. SettingUpTextFiles.py and CopyingText.py to get the raw text files; 2. SentenceConversion.py; 3. reference_remover.py; 4. testing.py and testingavg.py; 5. SentenceCreator.py; 6. matscholar_model.py to get matscholar tags; 7. training_model_gensim.py to get gensim model;8. word2vecscript.py and gensim_visual.py;

Resources

Name Format Description Link
57 This zip file contains a set of scripts that extracts frequently occurring words from the conference proceedings of Microscopy & Microanalysis between the years of 2002 and 2019. https://data.nist.gov/od/ds/ark:/88434/mds2-3198/PythonFiles_Maurice_clean.zip

Tags

  • controlled-vocabulary
  • electron-microscopy
  • natural-language-processing
  • ontology
  • nlp

Topics

Categories