Advanced search
1 file | 9.79 MB Add to list

How relevant is part-of-speech information to compute similarity between Greek verses in a graph database?

Colin Swaelens (UGent) , Maxime Deforche (UGent) , Guy De Tré (UGent) , Ilse De Vos and Els Lefever (UGent)
Author
Organization
Project
Abstract
This paper presents the automatic linguistic analysis of the Database of Byzantine Book Epigrams (DBBE) on the one hand, and its representation and integration in a graph database on the other hand. Firstly, we provide a comprehensive description of the DBBE data we want to provide with a complete morphological analysis. The presented methodology explores the possibilities of fine-tuning the DBBErt transformer-based language model, which was trained on pre-Modern and Modern Greek. Secondly, the automatically annotated epigrams are integrated in a graph database, a new way to represent the relatedness of this entangled corpus. With the graph database, we can compute similarity between words, verses and epigrams. Given the scope of this paper, we computed a complete orthographic similarity between the verses, a similarity based on the automatically assigned part-of-speech information and a final similarity measure that combines both orthography and part-of-speech information. The results of these similarity measures provide scholars with new visual representations of relations between (parts of) texts, which is beneficial for new critical editions and commentaries.
Keywords
Byzantine Book Epigrams, Linguistic Annotation, Graph Database, Ancient Language Processing, Similarity Search

Downloads

  • paper4.pdf
    • full text (Accepted manuscript)
    • |
    • open access
    • |
    • PDF
    • |
    • 9.79 MB

Citation

Please use this url to cite or link to this publication:

MLA
Swaelens, Colin, et al. “How Relevant Is Part-of-Speech Information to Compute Similarity between Greek Verses in a Graph Database?” Proceedings of the First Workshop on Data-Driven Approaches to Ancient Languages (DAAL 2024), edited by Colin Swaelens et al., Language & Translation Technology Team (LT3), 2024, pp. 33–43.
APA
Swaelens, C., Deforche, M., De Tré, G., De Vos, I., & Lefever, E. (2024). How relevant is part-of-speech information to compute similarity between Greek verses in a graph database? In C. Swaelens, M. Deforche, I. De Vos, & E. Lefever (Eds.), Proceedings of the first workshop on Data-driven Approaches to Ancient Languages (DAAL 2024) (pp. 33–43). Ghent: Language & Translation Technology Team (LT3).
Chicago author-date
Swaelens, Colin, Maxime Deforche, Guy De Tré, Ilse De Vos, and Els Lefever. 2024. “How Relevant Is Part-of-Speech Information to Compute Similarity between Greek Verses in a Graph Database?” In Proceedings of the First Workshop on Data-Driven Approaches to Ancient Languages (DAAL 2024), edited by Colin Swaelens, Maxime Deforche, Ilse De Vos, and Els Lefever, 33–43. Ghent: Language & Translation Technology Team (LT3).
Chicago author-date (all authors)
Swaelens, Colin, Maxime Deforche, Guy De Tré, Ilse De Vos, and Els Lefever. 2024. “How Relevant Is Part-of-Speech Information to Compute Similarity between Greek Verses in a Graph Database?” In Proceedings of the First Workshop on Data-Driven Approaches to Ancient Languages (DAAL 2024), ed by. Colin Swaelens, Maxime Deforche, Ilse De Vos, and Els Lefever, 33–43. Ghent: Language & Translation Technology Team (LT3).
Vancouver
1.
Swaelens C, Deforche M, De Tré G, De Vos I, Lefever E. How relevant is part-of-speech information to compute similarity between Greek verses in a graph database? In: Swaelens C, Deforche M, De Vos I, Lefever E, editors. Proceedings of the first workshop on Data-driven Approaches to Ancient Languages (DAAL 2024). Ghent: Language & Translation Technology Team (LT3); 2024. p. 33–43.
IEEE
[1]
C. Swaelens, M. Deforche, G. De Tré, I. De Vos, and E. Lefever, “How relevant is part-of-speech information to compute similarity between Greek verses in a graph database?,” in Proceedings of the first workshop on Data-driven Approaches to Ancient Languages (DAAL 2024), Ghent, Belgium, 2024, pp. 33–43.
@inproceedings{01J1557VST458K7028TMF8PR9A,
  abstract     = {{This paper presents the automatic linguistic analysis of the Database of Byzantine Book Epigrams (DBBE) on the one hand, and its representation and integration in a graph database on the other hand. Firstly, we provide a comprehensive description of the DBBE data we want to provide with a complete morphological analysis. The presented methodology explores the possibilities of fine-tuning the DBBErt transformer-based language model, which was trained on pre-Modern and Modern Greek. Secondly, the automatically annotated epigrams are integrated in a graph database, a new way to represent the relatedness of this entangled corpus. With the graph database, we can compute similarity between words, verses and epigrams. Given the scope of this paper, we computed a complete orthographic similarity between the verses, a similarity based on the automatically assigned part-of-speech information and a final similarity measure that combines both orthography and part-of-speech information. The results of these similarity measures provide scholars with new visual representations of relations between (parts of) texts, which is beneficial for new critical editions and commentaries.}},
  author       = {{Swaelens, Colin and Deforche, Maxime and De Tré, Guy and De Vos, Ilse and Lefever, Els}},
  booktitle    = {{Proceedings of the first workshop on Data-driven Approaches to Ancient Languages (DAAL 2024)}},
  editor       = {{Swaelens, Colin and Deforche, Maxime and De Vos, Ilse and Lefever, Els}},
  isbn         = {{9789078848127}},
  keywords     = {{Byzantine Book Epigrams,Linguistic Annotation,Graph Database,Ancient Language Processing,Similarity Search}},
  language     = {{eng}},
  location     = {{Ghent, Belgium}},
  pages        = {{33--43}},
  publisher    = {{Language & Translation Technology Team (LT3)}},
  title        = {{How relevant is part-of-speech information to compute similarity between Greek verses in a graph database?}},
  year         = {{2024}},
}