
A pilot study for BERT language modelling and morphological analysis for ancient and medieval Greek
- Author
- Pranaydeep Singh (UGent) , Gorik Rutten (UGent) and Els Lefever (UGent)
- Organization
- Project
- Abstract
- This paper presents a pilot study to automatic linguistic preprocessing of Ancient and Byzantine Greek, and morphological analysis more specifically. To this end, a novel subword-based BERT language model was trained on the basis of a varied corpus of Modern, Ancient and Post-classical Greek texts. Consequently, the obtained BERT embeddings were incorporated to train a fine-grained Part-of-Speech tagger for Ancient and Byzantine Greek. In addition, a corpus of Greek Epigrams was manually annotated and the resulting gold standard was used to evaluate the performance of the morphological analyser on Byzantine Greek. The experimental results show very good perplexity scores (4.9) for the BERT language model and state-of-the-art performance for the fine-grained Part-of-Speech tagger for in-domain data (treebanks containing a mixture of Classical and Medieval Greek), as well as for the newly created Byzantine Greek gold standard data set. The language models and associated code are made available for use at https://github.com/pranaydeeps/Ancient-Greek-BERT
- Keywords
- LT3
Downloads
-
2021.latechclfl-1.15.pdf
- full text (Published version)
- |
- open access
- |
- |
- 288.91 KB
Citation
Please use this url to cite or link to this publication: http://hdl.handle.net/1854/LU-8726146
- MLA
- Singh, Pranaydeep, et al. “A Pilot Study for BERT Language Modelling and Morphological Analysis for Ancient and Medieval Greek.” Proceedings of the 5th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature (LaTeCH-CLfL 2021), Association for Computational Linguistics, 2021, pp. 128–37.
- APA
- Singh, P., Rutten, G., & Lefever, E. (2021). A pilot study for BERT language modelling and morphological analysis for ancient and medieval Greek. Proceedings of the 5th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature (LaTeCH-CLfL 2021), 128–137. Association for Computational Linguistics.
- Chicago author-date
- Singh, Pranaydeep, Gorik Rutten, and Els Lefever. 2021. “A Pilot Study for BERT Language Modelling and Morphological Analysis for Ancient and Medieval Greek.” In Proceedings of the 5th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature (LaTeCH-CLfL 2021), 128–37. Association for Computational Linguistics.
- Chicago author-date (all authors)
- Singh, Pranaydeep, Gorik Rutten, and Els Lefever. 2021. “A Pilot Study for BERT Language Modelling and Morphological Analysis for Ancient and Medieval Greek.” In Proceedings of the 5th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature (LaTeCH-CLfL 2021), 128–137. Association for Computational Linguistics.
- Vancouver
- 1.Singh P, Rutten G, Lefever E. A pilot study for BERT language modelling and morphological analysis for ancient and medieval Greek. In: Proceedings of the 5th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature (LaTeCH-CLfL 2021). Association for Computational Linguistics; 2021. p. 128–37.
- IEEE
- [1]P. Singh, G. Rutten, and E. Lefever, “A pilot study for BERT language modelling and morphological analysis for ancient and medieval Greek,” in Proceedings of the 5th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature (LaTeCH-CLfL 2021), Online (Punta Cana, Dominican Republic), 2021, pp. 128–137.
@inproceedings{8726146, abstract = {{This paper presents a pilot study to automatic linguistic preprocessing of Ancient and Byzantine Greek, and morphological analysis more specifically. To this end, a novel subword-based BERT language model was trained on the basis of a varied corpus of Modern, Ancient and Post-classical Greek texts. Consequently, the obtained BERT embeddings were incorporated to train a fine-grained Part-of-Speech tagger for Ancient and Byzantine Greek. In addition, a corpus of Greek Epigrams was manually annotated and the resulting gold standard was used to evaluate the performance of the morphological analyser on Byzantine Greek. The experimental results show very good perplexity scores (4.9) for the BERT language model and state-of-the-art performance for the fine-grained Part-of-Speech tagger for in-domain data (treebanks containing a mixture of Classical and Medieval Greek), as well as for the newly created Byzantine Greek gold standard data set. The language models and associated code are made available for use at https://github.com/pranaydeeps/Ancient-Greek-BERT}}, author = {{Singh, Pranaydeep and Rutten, Gorik and Lefever, Els}}, booktitle = {{Proceedings of the 5th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature (LaTeCH-CLfL 2021)}}, isbn = {{9781954085916}}, keywords = {{LT3}}, language = {{eng}}, location = {{Online (Punta Cana, Dominican Republic)}}, pages = {{128--137}}, publisher = {{Association for Computational Linguistics}}, title = {{A pilot study for BERT language modelling and morphological analysis for ancient and medieval Greek}}, url = {{https://aclanthology.org/2021.latechclfl-1.15}}, year = {{2021}}, }