Back to search

University of Illinois at Urbana-Champaign

Compiling contextualized lists of frequent vocabulary from user- supplied corpora using natural language processing techniques

Abstract

dc:description

Since there are thousands of words to learn in a new language, one common challenge for language learners and teachers is knowing which vocabulary items to prioritize over the others and, in general, setting vocabulary-learning goals. Within vocabulary teaching research, one approach has been to focus on lists of the most common vocabulary. West (1953) proposed a list of the 2000 most frequent word families in English that, it was argued, were most important for learners to master. Along the same lines, Coxhead (2000) offered a list of the most common words in academic English known as the Academic Word List (AWL). Arguing that AWL did not adequately reflect the learners’ specialized vocabulary needs, however, corpus linguists began to develop wordlists in specialized subject areas with an English for Specific Purposes (ESP) perspective for students in Business, Engineering, Medical, and Law majors and so on. A central theme in almost all previous endeavors to develop better wordlists has been the notion of 'representativeness'—the extent to which a wordlist 'represents' the language needs of leaners. In this study, it is proposed that an alternative way to maximize representativeness in a wordlist is to enable users to compile a wordlist from any text or corpus that is of interest to them and to provide the means of compiling a wordlist using that text. Using Natural Language Toolkit (NLTK), this study shows how a few Natural Language Processing (NLP) techniques may be used to compile a list of the most common words in the Europarl corpus along with retrieving example sentences from the corpus for each word. This new approach can have applications for both language leaners as well as for the purposes of preparing instructional materials in an ESP setting.

Degree

thesis:*
Name thesis:degree_name
M.A.
Level thesis:degree_level
Thesis
Discipline thesis:degree_discipline
Teaching of English Sec Lang
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2016

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Abdar, Omid
Contributors dc:contributor
  • Sadler, Randall
  • Schwartz, Lane

Subjects

dc:subject × 1

Rights

dc:rights
Statement dc:rights
  • Copyright 2016 Omid Abdar
Language dc:language
en

Identifiers

dc:identifier.*
Handle dc:identifier
http://hdl.handle.net/2142/92955

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Abdar, Omid. Compiling contextualized lists of frequent vocabulary from user- supplied corpora using natural language processing techniques. Thesis thesis, University of Illinois at Urbana-Champaign, 2016. http://hdl.handle.net/2142/92955