Advanced search
1 file | 620.45 KB Add to list

Linguistic annotation of Byzantine book epigrams : revisited

Colin Swaelens (UGent) , Ilse De Vos and Els Lefever (UGent)
Author
Organization
Project
Abstract
In the current surge of interest in large language models (LLM) within the field of natural language processing (NLP), the automatic assignment of linguistic information may seem straightforward. However, tasks like part-of-speech tagging, morphological analysis, and lemmatisation pose significant challenges for ancient languages such as Greek, Latin, and Sanskrit. A major issue with these languages is that they are examples of closed corpora, meaning the available data is finite and cannot be expanded. Furthermore, these corpora are relatively small compared to those for languages like English or Chinese. Smaller corpora provide less data for training, which - in the case of LLMs - typically results in reduced performance. This challenge is compounded by the morphologically richness of languages like Greek, making automatic linguistic annotation even more difficult. This article presents our recently developed techniques for part-of-speech tagging, morphological analysis, and lemmatisation for Byzantine Greek. These recently developed techniques are compared to existing algorithms, in order to assess whether these seemingly simple tasks benefit from complex solutions. This is followed by a discussion that highlights the challenges researchers have faced over the past fifty years of developing linguistic analysis tools for Greek.
Keywords
Natural Language Processing, Ancient Language Processing, Machine Learning, Linguistics, Greek

Downloads

  • (...).pdf
    • full text (Published version)
    • |
    • UGent only
    • |
    • PDF
    • |
    • 620.45 KB

Citation

Please use this url to cite or link to this publication:

MLA
Swaelens, Colin, et al. “Linguistic Annotation of Byzantine Book Epigrams : Revisited.” Computational Approaches to Ancient Greek and Latin, De Gruyter, 2026.
APA
Swaelens, C., De Vos, I., & Lefever, E. (2026). Linguistic annotation of Byzantine book epigrams : revisited. In Computational approaches to Ancient Greek and Latin. Berlin ; Boston: De Gruyter.
Chicago author-date
Swaelens, Colin, Ilse De Vos, and Els Lefever. 2026. “Linguistic Annotation of Byzantine Book Epigrams : Revisited.” In Computational Approaches to Ancient Greek and Latin. Berlin ; Boston: De Gruyter.
Chicago author-date (all authors)
Swaelens, Colin, Ilse De Vos, and Els Lefever. 2026. “Linguistic Annotation of Byzantine Book Epigrams : Revisited.” In Computational Approaches to Ancient Greek and Latin. Berlin ; Boston: De Gruyter.
Vancouver
1.
Swaelens C, De Vos I, Lefever E. Linguistic annotation of Byzantine book epigrams : revisited. In: Computational approaches to Ancient Greek and Latin. Berlin ; Boston: De Gruyter; 2026.
IEEE
[1]
C. Swaelens, I. De Vos, and E. Lefever, “Linguistic annotation of Byzantine book epigrams : revisited,” in Computational approaches to Ancient Greek and Latin, Berlin ; Boston: De Gruyter, 2026.
@incollection{01K7P68QRG9M7CGYJEYVB33JB9,
  abstract     = {{In the current surge of interest in large language models (LLM) within the field of natural language processing (NLP), the automatic assignment of linguistic information may seem straightforward. However, tasks like part-of-speech tagging, morphological analysis, and lemmatisation pose significant challenges for ancient languages such as Greek, Latin, and Sanskrit. A major issue with these languages is that they are examples of closed corpora, meaning the available data is finite and cannot be expanded. Furthermore, these corpora are relatively small compared to those for languages like English or Chinese. Smaller corpora provide less data for training, which - in the case of LLMs - typically results in reduced performance. This challenge is compounded by the morphologically richness of languages like Greek, making automatic linguistic annotation even more difficult. This article presents our  recently developed techniques for part-of-speech tagging, morphological analysis, and lemmatisation for Byzantine Greek. These recently developed techniques are compared to existing algorithms, in order to assess whether these seemingly simple tasks benefit from complex solutions. This is followed by a discussion that highlights the challenges researchers have faced over the past fifty years of developing linguistic analysis tools for Greek.}},
  author       = {{Swaelens, Colin and De Vos, Ilse and Lefever, Els}},
  booktitle    = {{Computational approaches to Ancient Greek and Latin}},
  issn         = {{2199-0255}},
  keywords     = {{Natural Language Processing,Ancient Language Processing,Machine Learning,Linguistics,Greek}},
  language     = {{eng}},
  publisher    = {{De Gruyter}},
  series       = {{Philologus Supplemente}},
  title        = {{Linguistic annotation of Byzantine book epigrams : revisited}},
  year         = {{2026}},
}