Advanced search
1 file | 988.71 KB Add to list

How useful are corpus-based methods for extrapolating psycholinguistic variables?

Pawel Mandera (UGent) , Emmanuel Keuleers (UGent) and Marc Brysbaert (UGent)
Author
Organization
Abstract
Subjective ratings for age of acquisition, concreteness, affective valence, and many other variables are an important element of psycholinguistic research. However, even for well-studied languages, ratings usually cover just a small part of the vocabulary. A possible solution involves using corpora to build a semantic similarity space and to apply machine learning techniques to extrapolate existing ratings to previously unrated words. We conduct a systematic comparison of two extrapolation techniques: k-nearest neighbours, and random forest, in combination with semantic spaces built using latent seman- tic analysis, topic model, a hyperspace analogue to language (HAL)-like model, and a skip-gram model. A variant of the k-nearest neighbours method used with skip-gram word vectors gives the most accurate predictions but the random forest method has an advantage of being able to easily incorporate additional predictors. We evaluate the usefulness of the methods by exploring how much of the human perform- ance in a lexical decision task can be explained by extrapolated ratings for age of acquisition and how precisely we can assign words to discrete categories based on extrapolated ratings. We find that at least some of the extrapolation methods may introduce artefacts to the data and produce results that could lead to different conclusions that would be reached based on the human ratings. From a practical point of view, the usefulness of ratings extrapolated with the described methods may be limited.
Keywords
AGE-OF-ACQUISITION, LATENT SEMANTIC ANALYSIS, LEXICAL COOCCURRENCE, ENGLISH WORDS, NORMS, CONCRETENESS, PROJECT, RATINGS, MODELS, LEMMAS, Semantic models, Human ratings, Machine learning

Downloads

  • (...).pdf
    • full text
    • |
    • UGent only
    • |
    • PDF
    • |
    • 988.71 KB

Citation

Please use this url to cite or link to this publication:

MLA
Mandera, Pawel, Emmanuel Keuleers, and Marc Brysbaert. “How Useful Are Corpus-based Methods for Extrapolating Psycholinguistic Variables?” QUARTERLY JOURNAL OF EXPERIMENTAL PSYCHOLOGY 68.8 (2015): 1623–1642. Print.
APA
Mandera, P., Keuleers, E., & Brysbaert, M. (2015). How useful are corpus-based methods for extrapolating psycholinguistic variables? QUARTERLY JOURNAL OF EXPERIMENTAL PSYCHOLOGY, 68(8), 1623–1642.
Chicago author-date
Mandera, Pawel, Emmanuel Keuleers, and Marc Brysbaert. 2015. “How Useful Are Corpus-based Methods for Extrapolating Psycholinguistic Variables?” Quarterly Journal of Experimental Psychology 68 (8): 1623–1642.
Chicago author-date (all authors)
Mandera, Pawel, Emmanuel Keuleers, and Marc Brysbaert. 2015. “How Useful Are Corpus-based Methods for Extrapolating Psycholinguistic Variables?” Quarterly Journal of Experimental Psychology 68 (8): 1623–1642.
Vancouver
1.
Mandera P, Keuleers E, Brysbaert M. How useful are corpus-based methods for extrapolating psycholinguistic variables? QUARTERLY JOURNAL OF EXPERIMENTAL PSYCHOLOGY. 2015;68(8):1623–42.
IEEE
[1]
P. Mandera, E. Keuleers, and M. Brysbaert, “How useful are corpus-based methods for extrapolating psycholinguistic variables?,” QUARTERLY JOURNAL OF EXPERIMENTAL PSYCHOLOGY, vol. 68, no. 8, pp. 1623–1642, 2015.
@article{5878575,
  abstract     = {Subjective ratings for age of acquisition, concreteness, affective valence, and many other variables are an important element of psycholinguistic research. However, even for well-studied languages, ratings usually cover just a small part of the vocabulary. A possible solution involves using corpora to build a semantic similarity space and to apply machine learning techniques to extrapolate existing ratings to previously unrated words. We conduct a systematic comparison of two extrapolation techniques: k-nearest neighbours, and random forest, in combination with semantic spaces built using latent seman- tic analysis, topic model, a hyperspace analogue to language (HAL)-like model, and a skip-gram model. A variant of the k-nearest neighbours method used with skip-gram word vectors gives the most accurate predictions but the random forest method has an advantage of being able to easily incorporate additional predictors. We evaluate the usefulness of the methods by exploring how much of the human perform- ance in a lexical decision task can be explained by extrapolated ratings for age of acquisition and how precisely we can assign words to discrete categories based on extrapolated ratings. We find that at least some of the extrapolation methods may introduce artefacts to the data and produce results that could lead to different conclusions that would be reached based on the human ratings. From a practical point of view, the usefulness of ratings extrapolated with the described methods may be limited.},
  author       = {Mandera, Pawel and Keuleers, Emmanuel and Brysbaert, Marc},
  issn         = {1747-0218},
  journal      = {QUARTERLY JOURNAL OF EXPERIMENTAL PSYCHOLOGY},
  keywords     = {AGE-OF-ACQUISITION,LATENT SEMANTIC ANALYSIS,LEXICAL COOCCURRENCE,ENGLISH WORDS,NORMS,CONCRETENESS,PROJECT,RATINGS,MODELS,LEMMAS,Semantic models,Human ratings,Machine learning},
  language     = {eng},
  number       = {8},
  pages        = {1623--1642},
  title        = {How useful are corpus-based methods for extrapolating psycholinguistic variables?},
  url          = {http://dx.doi.org/10.1080/17470218.2014.988735},
  volume       = {68},
  year         = {2015},
}

Altmetric
View in Altmetric
Web of Science
Times cited: