The automatic determination of translation equivalents in lexicography : what works and what doesn’t?
- Author
- Michaela Denisová, Gilles-Maurice de Schryver (UGent) and Pavel Rychlý
- Organization
- Project
- Abstract
- Cross-lingual embedding models act as facilitator of lexical knowledge transfer and offer many advantages, notably their applicability to low-resource and non-standard language pairs, making them a valuable tool for retrieving translation equivalents in lexicography. Despite their potential, these models have primarily been developed with a focus on Natural Language Processing (NLP), leading to significant issues, including flawed training and evaluation data, as well as inadequate evaluation metrics and procedures. In this paper, we introduce cross-lingual embedding models for lexicography, addressing the challenges and limitations inherent in the current NLP-focused research. We demonstrate the problematic aspects across three baseline cross-lingual embedding models and three language pairs and outline possible solutions. We show the importance of high-quality data, advocating that its role is vital compared to algorithmic optimisation in enhancing the effectiveness of these models.
- Keywords
- translation equivalent determination, cross-lingual embedding models, evaluation
Downloads
-
publisher version.pdf
- full text (Published version)
- |
- open access
- |
- |
- 494.57 KB
Citation
Please use this url to cite or link to this publication: http://hdl.handle.net/1854/LU-01JP4XG73RXXESH077X6N14J6Y
- MLA
- Denisová, Michaela, et al. “The Automatic Determination of Translation Equivalents in Lexicography : What Works and What Doesn’t?” Lexicography and Semantics : Proceedings of the XXI EURALEX International Congress, 8–12 October 2024, Cavtat, Croatia, edited by Kristina Š. Despot et al., Institute for the Croatian Language, 2024, pp. 305–16.
- APA
- Denisová, M., de Schryver, G.-M., & Rychlý, P. (2024). The automatic determination of translation equivalents in lexicography : what works and what doesn’t? In K. Š. Despot, A. Ostroški Anić, & I. Brač (Eds.), Lexicography and semantics : proceedings of the XXI EURALEX international congress, 8–12 October 2024, Cavtat, Croatia (pp. 305–316). Zagreb: Institute for the Croatian Language.
- Chicago author-date
- Denisová, Michaela, Gilles-Maurice de Schryver, and Pavel Rychlý. 2024. “The Automatic Determination of Translation Equivalents in Lexicography : What Works and What Doesn’t?” In Lexicography and Semantics : Proceedings of the XXI EURALEX International Congress, 8–12 October 2024, Cavtat, Croatia, edited by Kristina Š. Despot, Ana Ostroški Anić, and Ivana Brač, 305–16. Zagreb: Institute for the Croatian Language.
- Chicago author-date (all authors)
- Denisová, Michaela, Gilles-Maurice de Schryver, and Pavel Rychlý. 2024. “The Automatic Determination of Translation Equivalents in Lexicography : What Works and What Doesn’t?” In Lexicography and Semantics : Proceedings of the XXI EURALEX International Congress, 8–12 October 2024, Cavtat, Croatia, ed by. Kristina Š. Despot, Ana Ostroški Anić, and Ivana Brač, 305–316. Zagreb: Institute for the Croatian Language.
- Vancouver
- 1.Denisová M, de Schryver G-M, Rychlý P. The automatic determination of translation equivalents in lexicography : what works and what doesn’t? In: Despot KŠ, Ostroški Anić A, Brač I, editors. Lexicography and semantics : proceedings of the XXI EURALEX international congress, 8–12 October 2024, Cavtat, Croatia. Zagreb: Institute for the Croatian Language; 2024. p. 305–16.
- IEEE
- [1]M. Denisová, G.-M. de Schryver, and P. Rychlý, “The automatic determination of translation equivalents in lexicography : what works and what doesn’t?,” in Lexicography and semantics : proceedings of the XXI EURALEX international congress, 8–12 October 2024, Cavtat, Croatia, Cavtat, Croatia, 2024, pp. 305–316.
@inproceedings{01JP4XG73RXXESH077X6N14J6Y,
abstract = {{Cross-lingual embedding models act as facilitator of lexical knowledge transfer and offer many advantages, notably their applicability to low-resource and non-standard language pairs, making them a valuable tool for retrieving translation equivalents in lexicography. Despite their potential, these models have primarily been developed with a focus on Natural Language Processing (NLP), leading to significant issues, including flawed training and evaluation data, as well as inadequate evaluation metrics and procedures. In this paper, we introduce cross-lingual embedding models for lexicography, addressing the challenges and limitations inherent in the current NLP-focused research. We demonstrate the problematic aspects across three baseline cross-lingual embedding models and three language pairs and outline possible solutions. We show the importance of high-quality data, advocating that its role is vital compared to algorithmic optimisation in enhancing the effectiveness of these models.}},
author = {{Denisová, Michaela and de Schryver, Gilles-Maurice and Rychlý, Pavel}},
booktitle = {{Lexicography and semantics : proceedings of the XXI EURALEX international congress, 8–12 October 2024, Cavtat, Croatia}},
editor = {{Despot, Kristina Š. and Ostroški Anić, Ana and Brač, Ivana}},
isbn = {{9789537967772}},
keywords = {{translation equivalent determination,cross-lingual embedding models,evaluation}},
language = {{eng}},
location = {{Cavtat, Croatia}},
pages = {{305--316}},
publisher = {{Institute for the Croatian Language}},
title = {{The automatic determination of translation equivalents in lexicography : what works and what doesn’t?}},
url = {{https://euralex.jezik.hr/wp-content/uploads/2021/09/Euralax-XXI-final-web.pdf}},
year = {{2024}},
}