How useful are corpus-based methods for extrapolating psycholinguistic variables?

Q J Exp Psychol (Hove)

a Department of Experimental Psychology , Ghent University, Ghent , Belgium.

Published: November 2016

Subjective ratings for age of acquisition, concreteness, affective valence, and many other variables are an important element of psycholinguistic research. However, even for well-studied languages, ratings usually cover just a small part of the vocabulary. A possible solution involves using corpora to build a semantic similarity space and to apply machine learning techniques to extrapolate existing ratings to previously unrated words. We conduct a systematic comparison of two extrapolation techniques: k-nearest neighbours, and random forest, in combination with semantic spaces built using latent semantic analysis, topic model, a hyperspace analogue to language (HAL)-like model, and a skip-gram model. A variant of the k-nearest neighbours method used with skip-gram word vectors gives the most accurate predictions but the random forest method has an advantage of being able to easily incorporate additional predictors. We evaluate the usefulness of the methods by exploring how much of the human performance in a lexical decision task can be explained by extrapolated ratings for age of acquisition and how precisely we can assign words to discrete categories based on extrapolated ratings. We find that at least some of the extrapolation methods may introduce artefacts to the data and produce results that could lead to different conclusions that would be reached based on the human ratings. From a practical point of view, the usefulness of ratings extrapolated with the described methods may be limited.

Download full-text PDF

Source
http://dx.doi.org/10.1080/17470218.2014.988735DOI Listing

Publication Analysis

Top Keywords

ratings age
8
age acquisition
8
k-nearest neighbours
8
random forest
8
extrapolated ratings
8
ratings
7
corpus-based methods
4
methods extrapolating
4
extrapolating psycholinguistic
4
psycholinguistic variables?
4

Similar Publications

Want AI Summaries of new PubMed Abstracts delivered to your In-box?

Enter search terms and have AI summaries delivered each week - change queries or unsubscribe any time!