Source Code Embeddings

Description

A set of six pretrained fastText models for semantic representations of source code.  Each of the models has been trained on high-quality GitHub repositories where the primary language is one of Java, Python, C++, C#, C, PHP. For collecting training data 13.144 repositories were cloned, 2.402.790.348 lines of code were read out of 944,467,560 files and preprocessed, to finally produce a total of 944.467.560 tokens of clean training data.  For further details refer to the following paper:  Efstathiou, V.,  Spinellis, D., 2019. "Semantic Source Code Models Using Identifier Embeddings". In 16th International Conference on Mining Software Repositories: Data Showcase Track. MSR'19. 

Resources

Name Format Description Link
0 http://data.europa.eu/88u/dataset/oai-zenodo-org-2558730
0 http://data.europa.eu/88u/dataset/oai-zenodo-org-2558730
0 http://data.europa.eu/88u/dataset/oai-zenodo-org-2558730
0 http://data.europa.eu/88u/dataset/oai-zenodo-org-2558730
0 http://data.europa.eu/88u/dataset/oai-zenodo-org-2558730
0 http://data.europa.eu/88u/dataset/oai-zenodo-org-2558730
0 http://data.europa.eu/88u/dataset/oai-zenodo-org-2558730

Tags

  • fasttext,-code-semantics,-vector-space-models

Topics

Categories