Advanced search
1 file | 10.15 MB Add to list

IVESS : intelligent vocabulary and example selection for Spanish vocabulary learning

(2024)
Author
Promoter
(UGent) , (UGent) and (UGent)
Organization
Project
Abstract
In the rapidly evolving digital era of the 21st century, language and technology are growing closer together than ever. Language technology tools have become particularly adept at performing so-called “transactional” language operations. Translating text from one language into another, for example, is a transaction that machine translation tools such as DeepL and Google Translate will generally perform well. A direct result of this ever-improving technological support for language tasks is that the main reasons for people to learn a foreign/second language (L2) – such as travelling or boosting one’s career opportunities – become less incentivising. Yet, human conversations and social interactions in foreign languages do not exclusively revolve around performing practical language transactions, but also around establishing relationships. For technological applications to assist in building these relationships, they face tasks that appear much more difficult to achieve, such as interpreting facial expressions. It is therefore likely that relying on technological tools alone will not suffice for anyone who wants to establish sustainable international contacts and fully engage in a foreign culture: being able to understand and speak foreign languages remains – even in the current digital era – an indispensable human skill. What language technology tools do possess, however, is the ability to play the role of a valuable assistant in the L2 learning process. It is this research area, the interface between L2 learning and computer assistance, to which the present PhD dissertation will contribute. More specifically, we conduct research into how relevant vocabulary items and example sentences (that illustrate how these vocabulary items are used) can be automatically selected from large collections of texts (“corpora”) to facilitate the development of learning materials for L2 Spanish. To perform these operations in an “intelligent” way (i.e. more efficiently and more tailored to the needs of L2 learners), we resort to natural language processing (NLP) techniques as our source of “intelligence”, making this research fall in the scientific domain of “Intelligent Computer-Assisted Language Learning” (ICALL). In our experiments, we show (1) that computer-assisted methods for vocabulary retrieval are highly performant but require further research to ensure the error-free output necessary in pedagogical settings; (2) that automatic vocabulary selection methods show moderate to strong correlations with L2 learners’ intuitions and needs; and (3) that artificially generated example sentences are found more suitable by L2 learners than sentences retrieved from corpora. Regarding difficulty-based vocabulary selection, we present LexComSpaL2, a first-of-its-kind dataset that can be used to train individualised vocabulary difficulty classifiers. Finally, we show by means of three case studies that learning in an adaptive and interactive online learning environment (called an “ICALL ecosystem”) does not necessarily lead to a better user experience or an increased interest in language technology, but that it does enhance students’ insights into NLP and increase their confidence in the computer as a learning assistant.
Te midden van de razendsnelle evoluties in het 21ste-eeuwse digitale tijdperk groeien taal en technologie steeds meer naar elkaar toe. Vandaag de dag zijn taaltechnologietools erg performant geworden in taken waarbij een “transactie” met taal plaatsvindt. Machinevertalingtools zoals DeepL zijn bijvoorbeeld in staat om vertaaltransacties doorgaans tot een goed einde te brengen. Een direct gevolg van deze steeds betere technologische ondersteuning bij taaltaken is dat de redenen om een vreemde taal (L2) te leren – zoals reizen of carrière maken op het werk – aan aantrekkelijkheid inboeten. Nochtans is het uitvoeren van praktische taaltransacties niet het enige waar menselijke conversaties in vreemde talen rond draaien, deze sociale interacties dienen namelijk ook vaak om relaties op te bouwen. En wanneer we met technologische tools een bijdrage willen leveren aan het bouwen van dit soort relaties, blijkt dit veel moeilijker te verwezenlijken (denk maar aan de correcte interpretatie van gelaatsuitdrukkingen). De kans is dan ook groot dat uitsluitend vertrouwen op digitale tools onvoldoende zal blijken voor iedereen die duurzame internationale contacten wil leggen en zich ten volle wil ontplooien in een vreemde cultuur. In staat zijn om vreemde talen te begrijpen en spreken blijft dus ook in het huidige digitale tijdperk een onmisbare menselijke vaardigheid. Wat taaltechnologietools echter wél te bieden hebben, is de mogelijkheid om de rol van waardevolle assistent te vervullen tijdens het L2-leerproces. Het is aan dit onderzoeksgebied, het raakvlak tussen L2-verwerving en computerondersteuning, dat deze doctoraatsstudie wil bijdragen. Ons onderzoek spitst zich specifiek toe op de automatische selectie van relevante woordenschatitems en voorbeeldzinnen (die illustreren hoe de woordenschatitems gebruikt worden) uit grote verzamelingen teksten (“corpora”), met als doel om de ontwikkeling van leermaterialen voor L2 Spaans te vereenvoudigen. Om deze selectie op een “intelligente” manier te laten verlopen (d.w.z. efficiënter en meer op maat van de noden van L2-leerders), doen we een beroep op technieken uit de natuurlijke taalverwerking (afgekort in het Engels als NLP). Hierdoor landen we met ons onderzoek in het domein van “Intelligent Computer-Assisted Language Learning” (ICALL). In onze experimenten tonen we aan (1) dat computerondersteunde methodes om woordenschat te herkennen en uit de corpusteksten te halen erg performant zijn, maar dat verder onderzoek noodzakelijk is om tot de foutloze output te komen die vereist is in een pedagogische context; (2) dat automatische methodes voor woordenschatselectie een matige tot sterke correlatie vertonen met de intuïties en noden van L2-leerders; en (3) dat artificieel gegenereerde voorbeeldzinnen geschikter worden bevonden door L2-leerders dan zinnen uit corpora. Wat betreft woordenschatselectie op basis van moeilijkheidsniveau stellen we LexComSpaL2 voor, een dataset die – als eerste in z'n soort – gebruikt kan worden om gepersonaliseerde modellen te trainen die de moeilijkheidsgraad van woordenschat kunnen voorspellen. Tot slot tonen we via drie casestudies aan dat leren in een aangepaste en interactieve online leeromgeving (“ICALL-ecosysteem” gedoopt) niet noodzakelijk leidt tot een betere gebruikservaring of tot een grotere interesse in taaltechnologie, maar dat studenten op deze manier wel verbeterde inzichten in NLP verwerven en meer vertrouwen krijgen in de computer als leerassistent.

Downloads

  • JasperDegraeuwe PhD dissertation IVESS rev.pdf
    • full text (Published version)
    • |
    • open access
    • |
    • PDF
    • |
    • 10.15 MB

Citation

Please use this url to cite or link to this publication:

MLA
Degraeuwe, Jasper. IVESS : Intelligent Vocabulary and Example Selection for Spanish Vocabulary Learning. Ghent University. Faculty of Arts and Philosophy, 2024.
APA
Degraeuwe, J. (2024). IVESS : intelligent vocabulary and example selection for Spanish vocabulary learning. Ghent University. Faculty of Arts and Philosophy, Ghent, Belgium.
Chicago author-date
Degraeuwe, Jasper. 2024. “IVESS : Intelligent Vocabulary and Example Selection for Spanish Vocabulary Learning.” Ghent, Belgium: Ghent University. Faculty of Arts and Philosophy.
Chicago author-date (all authors)
Degraeuwe, Jasper. 2024. “IVESS : Intelligent Vocabulary and Example Selection for Spanish Vocabulary Learning.” Ghent, Belgium: Ghent University. Faculty of Arts and Philosophy.
Vancouver
1.
Degraeuwe J. IVESS : intelligent vocabulary and example selection for Spanish vocabulary learning. [Ghent, Belgium]: Ghent University. Faculty of Arts and Philosophy; 2024.
IEEE
[1]
J. Degraeuwe, “IVESS : intelligent vocabulary and example selection for Spanish vocabulary learning,” Ghent University. Faculty of Arts and Philosophy, Ghent, Belgium, 2024.
@phdthesis{01J8TMB4TKFW0VG82NQGG87R1X,
  abstract     = {{In the rapidly evolving digital era of the 21st century, language and technology are growing closer together than ever. Language technology tools have become particularly adept at performing so-called “transactional” language operations. Translating text from one language into another, for example, is a transaction that machine translation tools such as DeepL and Google Translate will generally perform well. A direct result of this ever-improving technological support for language tasks is that the main reasons for people to learn a foreign/second language (L2) – such as travelling or boosting one’s career opportunities – become less incentivising. Yet, human conversations and social interactions in foreign languages do not exclusively revolve around performing practical language transactions, but also around establishing relationships. For technological applications to assist in building these relationships, they face tasks that appear much more difficult to achieve, such as interpreting facial expressions. It is therefore likely that relying on technological tools alone will not suffice for anyone who wants to establish sustainable international contacts and fully engage in a foreign culture: being able to understand and speak foreign languages remains – even in the current digital era – an indispensable human skill. What language technology tools do possess, however, is the ability to play the role of a valuable assistant in the L2 learning process. It is this research area, the interface between L2 learning and computer assistance, to which the present PhD dissertation will contribute. More specifically, we conduct research into how relevant vocabulary items and example sentences (that illustrate how these vocabulary items are used) can be automatically selected from large collections of texts (“corpora”) to facilitate the development of learning materials for L2 Spanish. To perform these operations in an “intelligent” way (i.e. more efficiently and more tailored to the needs of L2 learners), we resort to natural language processing (NLP) techniques as our source of “intelligence”, making this research fall in the scientific domain of “Intelligent Computer-Assisted Language Learning” (ICALL). In our experiments, we show (1) that computer-assisted methods for vocabulary retrieval are highly performant but require further research to ensure the error-free output necessary in pedagogical settings; (2) that automatic vocabulary selection methods show moderate to strong correlations with L2 learners’ intuitions and needs; and (3) that artificially generated example sentences are found more suitable by L2 learners than sentences retrieved from corpora. Regarding difficulty-based vocabulary selection, we present LexComSpaL2, a first-of-its-kind dataset that can be used to train individualised vocabulary difficulty classifiers. Finally, we show by means of three case studies that learning in an adaptive and interactive online learning environment (called an “ICALL ecosystem”) does not necessarily lead to a better user experience or an increased interest in language technology, but that it does enhance students’ insights into NLP and increase their confidence in the computer as a learning assistant.}},
  author       = {{Degraeuwe, Jasper}},
  language     = {{eng}},
  pages        = {{XVIII, 317}},
  publisher    = {{Ghent University. Faculty of Arts and Philosophy}},
  school       = {{Ghent University}},
  title        = {{IVESS : intelligent vocabulary and example selection for Spanish vocabulary learning}},
  year         = {{2024}},
}