Semantic Text Analyser BERT-like language model for formal language understanding
Description
SeTABERTa is a new multilingual langue model pertained from scratch using various Open Access text repositories: EU legislation, research articles, EU public documents and US patents. 2/3 of training data is English. The other part of data covers EU24 languages. The model was trained on JRC Big Data Platform. The model can be fine-tuned for other tasks.
The model is available on HuggingFace at https://huggingface.co/vidaud/SeTABERTa-mlm-v1 and can be loaded with FuggingFace transformers library.
Resources
| Name |
Format |
Description |
Link |
|
57 |
|
|