Advanced search
1 file | 494.57 KB Add to list

The automatic determination of translation equivalents in lexicography : what works and what doesn’t?

Author
Organization
Project
Abstract
Cross-lingual embedding models act as facilitator of lexical knowledge transfer and offer many advantages, notably their applicability to low-resource and non-standard language pairs, making them a valuable tool for retrieving translation equivalents in lexicography. Despite their potential, these models have primarily been developed with a focus on Natural Language Processing (NLP), leading to significant issues, including flawed training and evaluation data, as well as inadequate evaluation metrics and procedures. In this paper, we introduce cross-lingual embedding models for lexicography, addressing the challenges and limitations inherent in the current NLP-focused research. We demonstrate the problematic aspects across three baseline cross-lingual embedding models and three language pairs and outline possible solutions. We show the importance of high-quality data, advocating that its role is vital compared to algorithmic optimisation in enhancing the effectiveness of these models.
Keywords
translation equivalent determination, cross-lingual embedding models, evaluation

Downloads

  • publisher version.pdf
    • full text (Published version)
    • |
    • open access
    • |
    • PDF
    • |
    • 494.57 KB

Citation

Please use this url to cite or link to this publication:

MLA
Denisová, Michaela, et al. “The Automatic Determination of Translation Equivalents in Lexicography : What Works and What Doesn’t?” Lexicography and Semantics : Proceedings of the XXI EURALEX International Congress, 8–12 October 2024, Cavtat, Croatia, edited by Kristina Š. Despot et al., Institute for the Croatian Language, 2024, pp. 305–16.
APA
Denisová, M., de Schryver, G.-M., & Rychlý, P. (2024). The automatic determination of translation equivalents in lexicography : what works and what doesn’t? In K. Š. Despot, A. Ostroški Anić, & I. Brač (Eds.), Lexicography and semantics : proceedings of the XXI EURALEX international congress, 8–12 October 2024, Cavtat, Croatia (pp. 305–316). Zagreb: Institute for the Croatian Language.
Chicago author-date
Denisová, Michaela, Gilles-Maurice de Schryver, and Pavel Rychlý. 2024. “The Automatic Determination of Translation Equivalents in Lexicography : What Works and What Doesn’t?” In Lexicography and Semantics : Proceedings of the XXI EURALEX International Congress, 8–12 October 2024, Cavtat, Croatia, edited by Kristina Š. Despot, Ana Ostroški Anić, and Ivana Brač, 305–16. Zagreb: Institute for the Croatian Language.
Chicago author-date (all authors)
Denisová, Michaela, Gilles-Maurice de Schryver, and Pavel Rychlý. 2024. “The Automatic Determination of Translation Equivalents in Lexicography : What Works and What Doesn’t?” In Lexicography and Semantics : Proceedings of the XXI EURALEX International Congress, 8–12 October 2024, Cavtat, Croatia, ed by. Kristina Š. Despot, Ana Ostroški Anić, and Ivana Brač, 305–316. Zagreb: Institute for the Croatian Language.
Vancouver
1.
Denisová M, de Schryver G-M, Rychlý P. The automatic determination of translation equivalents in lexicography : what works and what doesn’t? In: Despot KŠ, Ostroški Anić A, Brač I, editors. Lexicography and semantics : proceedings of the XXI EURALEX international congress, 8–12 October 2024, Cavtat, Croatia. Zagreb: Institute for the Croatian Language; 2024. p. 305–16.
IEEE
[1]
M. Denisová, G.-M. de Schryver, and P. Rychlý, “The automatic determination of translation equivalents in lexicography : what works and what doesn’t?,” in Lexicography and semantics : proceedings of the XXI EURALEX international congress, 8–12 October 2024, Cavtat, Croatia, Cavtat, Croatia, 2024, pp. 305–316.
@inproceedings{01JP4XG73RXXESH077X6N14J6Y,
  abstract     = {{Cross-lingual embedding models act as facilitator of lexical knowledge transfer and offer many advantages, notably their applicability to low-resource and non-standard language pairs, making them a valuable tool for retrieving translation equivalents in lexicography. Despite their potential, these models have primarily been developed with a focus on Natural Language Processing (NLP), leading to significant issues, including flawed training and evaluation data, as well as inadequate evaluation metrics and procedures. In this paper, we introduce cross-lingual embedding models for lexicography, addressing the challenges and limitations inherent in the current NLP-focused research. We demonstrate the problematic aspects across three baseline cross-lingual embedding models and three language pairs and outline possible solutions. We show the importance of high-quality data, advocating that its role is vital compared to algorithmic optimisation in enhancing the effectiveness of these models.}},
  author       = {{Denisová, Michaela and de Schryver, Gilles-Maurice and Rychlý, Pavel}},
  booktitle    = {{Lexicography and semantics : proceedings of the XXI EURALEX international congress, 8–12 October 2024, Cavtat, Croatia}},
  editor       = {{Despot, Kristina Š. and Ostroški Anić, Ana and Brač, Ivana}},
  isbn         = {{9789537967772}},
  keywords     = {{translation equivalent determination,cross-lingual embedding models,evaluation}},
  language     = {{eng}},
  location     = {{Cavtat, Croatia}},
  pages        = {{305--316}},
  publisher    = {{Institute for the Croatian Language}},
  title        = {{The automatic determination of translation equivalents in lexicography : what works and what doesn’t?}},
  url          = {{https://euralex.jezik.hr/wp-content/uploads/2021/09/Euralax-XXI-final-web.pdf}},
  year         = {{2024}},
}