Advanced search
1 file | 457.54 KB Add to list

Lemmatisation of Medieval Greek : against the limits of transformer’s capabilities?

Colin Swaelens (UGent) , Pranaydeep Singh (UGent) , Ilse De Vos and Els Lefever (UGent)
Author
Organization
Project
Abstract
This paper presents preliminary experiments for the lemmatisation of unedited, Byzantine Greek epigrams. This type of Greek is quite different from its classical ancestor, mostly because of its orthographic inconsistencies. Existing lemmatisation algorithms display an accuracy drop of around 30pp when tested on these Byzantine book epigrams. We conducted seven different lemmatisation experiments, which were either transformer-based or based on neural edit-trees. The best performing lemmatiser was a hybrid method combining transformer-based embeddings with a dictionary look-up. We compare our results with existing lemmatisers, and provide a detailed error analysis revealing why unedited, Byzantine Greek is so challenging for lemmatisation.
Keywords
Natural Language Processing, Language Technology, Annotation, Lemmatisation, Byzantine Greek

Downloads

  • 2024.lrec-main.899.pdf
    • full text (Published version)
    • |
    • open access
    • |
    • PDF
    • |
    • 457.54 KB

Citation

Please use this url to cite or link to this publication:

MLA
Swaelens, Colin, et al. “Lemmatisation of Medieval Greek : Against the Limits of Transformer’s Capabilities?” Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), edited by Nicoletta Calzolari et al., ELRA, 2024, pp. 10293–302.
APA
Swaelens, C., Singh, P., De Vos, I., & Lefever, E. (2024). Lemmatisation of Medieval Greek : against the limits of transformer’s capabilities? In N. Calzolari, M.-Y. Kan, V. Hoste, A. Lenci, S. Sakti, & N. Xue (Eds.), Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) (pp. 10293–10302). ELRA.
Chicago author-date
Swaelens, Colin, Pranaydeep Singh, Ilse De Vos, and Els Lefever. 2024. “Lemmatisation of Medieval Greek : Against the Limits of Transformer’s Capabilities?” In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), edited by Nicoletta Calzolari, Min-Yen Kan, Veronique Hoste, Alessandro Lenci, Sakriani Sakti, and Nianwen Xue, 10293–302. ELRA.
Chicago author-date (all authors)
Swaelens, Colin, Pranaydeep Singh, Ilse De Vos, and Els Lefever. 2024. “Lemmatisation of Medieval Greek : Against the Limits of Transformer’s Capabilities?” In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), ed by. Nicoletta Calzolari, Min-Yen Kan, Veronique Hoste, Alessandro Lenci, Sakriani Sakti, and Nianwen Xue, 10293–10302. ELRA.
Vancouver
1.
Swaelens C, Singh P, De Vos I, Lefever E. Lemmatisation of Medieval Greek : against the limits of transformer’s capabilities? In: Calzolari N, Kan M-Y, Hoste V, Lenci A, Sakti S, Xue N, editors. Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024). ELRA; 2024. p. 10293–302.
IEEE
[1]
C. Swaelens, P. Singh, I. De Vos, and E. Lefever, “Lemmatisation of Medieval Greek : against the limits of transformer’s capabilities?,” in Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), Turin, Italy, 2024, pp. 10293–10302.
@inproceedings{01HYMT2PAHAVF998YC9R6AWM1Z,
  abstract     = {{This paper presents preliminary experiments for the lemmatisation of unedited, Byzantine Greek epigrams. This type of Greek is quite different from its classical ancestor, mostly because of its orthographic inconsistencies. Existing lemmatisation algorithms display an accuracy drop of around 30pp when tested on these Byzantine book epigrams. We conducted seven different lemmatisation experiments, which were either transformer-based or based on neural edit-trees. The best performing lemmatiser was a hybrid method combining transformer-based embeddings with a dictionary look-up. We compare our results with existing lemmatisers, and provide a detailed error analysis revealing why unedited, Byzantine Greek is so challenging for lemmatisation.
}},
  author       = {{Swaelens, Colin and Singh, Pranaydeep and De Vos, Ilse and Lefever, Els}},
  booktitle    = {{Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)}},
  editor       = {{Calzolari, Nicoletta and Kan, Min-Yen and Hoste, Veronique and Lenci, Alessandro and Sakti, Sakriani and Xue, Nianwen}},
  isbn         = {{9782493814104}},
  issn         = {{2951-2093}},
  keywords     = {{Natural Language Processing,Language Technology,Annotation,Lemmatisation,Byzantine Greek}},
  language     = {{eng}},
  location     = {{Turin, Italy}},
  pages        = {{10293--10302}},
  publisher    = {{ELRA}},
  title        = {{Lemmatisation of Medieval Greek : against the limits of transformer’s capabilities?}},
  url          = {{https://aclanthology.org/2024.lrec-main.899}},
  year         = {{2024}},
}