☆ 4.2 Article

How useful are corpus-based methods for extrapolating psycholinguistic variables?

QUARTERLY JOURNAL OF EXPERIMENTAL PSYCHOLOGY (2015)

Journal

QUARTERLY JOURNAL OF EXPERIMENTAL PSYCHOLOGY

Volume 68, Issue 8, Pages 1623-1642

Publisher

SAGE PUBLICATIONS LTD

DOI: 10.1080/17470218.2014.988735

Keywords

Semantic models; Human ratings; Machine learning

Funding

Odysseus grant from the Government of Flanders

Ask authors/readers for more resources

Protocol

Community support

Reagent

Community support

Abstract

Subjective ratings for age of acquisition, concreteness, affective valence, and many other variables are an important element of psycholinguistic research. However, even for well-studied languages, ratings usually cover just a small part of the vocabulary. A possible solution involves using corpora to build a semantic similarity space and to apply machine learning techniques to extrapolate existing ratings to previously unrated words. We conduct a systematic comparison of two extrapolation techniques: k-nearest neighbours, and random forest, in combination with semantic spaces built using latent semantic analysis, topic model, a hyperspace analogue to language (HAL)-like model, and a skip-gram model. A variant of the k-nearest neighbours method used with skip-gram word vectors gives the most accurate predictions but the random forest method has an advantage of being able to easily incorporate additional predictors. We evaluate the usefulness of the methods by exploring how much of the human performance in a lexical decision task can be explained by extrapolated ratings for age of acquisition and how precisely we can assign words to discrete categories based on extrapolated ratings. We find that at least some of the extrapolation methods may introduce artefacts to the data and produce results that could lead to different conclusions that would be reached based on the human ratings. From a practical point of view, the usefulness of ratings extrapolated with the described methods may be limited.

How useful are corpus-based methods for extrapolating psycholinguistic variables?

Journal

QUARTERLY JOURNAL OF EXPERIMENTAL PSYCHOLOGY

Publisher

SAGE PUBLICATIONS LTD

Keywords

Categories

Funding

Ask authors/readers for more resources

Protocol

Reagent

Authors

I am an author on this paper

Reviews

Primary Rating

Secondary Ratings

Novelty

Significance

Scientific rigor

Rate this paper

Recommended

How useful are corpus-based methods for extrapolating psycholinguistic variables?

Journal

QUARTERLY JOURNAL OF EXPERIMENTAL PSYCHOLOGY

Publisher

SAGE PUBLICATIONS LTD

Keywords

Categories

Funding

Ask authors/readers for more resources

Protocol

Reagent

Authors

I am an author on this paper

Reviews

Primary Rating

Secondary Ratings

Novelty

Significance

Scientific rigor

Rate this paper

Recommended

Export Citation

Share Paper