{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/92955"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/92955","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Compiling contextualized lists of frequent vocabulary from user- supplied corpora using natural language processing techniques","abstract":"Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2018-08-01","abstract_html":"Submission published under a 24 month embargo labeled &#x27;Closed Access&#x27;, the embargo will last until 2018-08-01","abstract_has_math":false,"creators":["Abdar, Omid"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.A.","degree_level":"Thesis","degree_discipline":"Teaching of English Sec Lang","degree_department":null,"school":null,"contributors":["Sadler, Randall","Schwartz, Lane"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2016,"date_issued":"2016-11-10T18:27:52Z","date_published":"2016-11-10T18:27:52Z","updated_at":"2026-07-22T22:26:35Z","subjects":["English for Specific Purposes, Vocabulary, Wordlists, Natural Language Processing"],"languages":["en"],"rights":["Copyright 2016 Omid Abdar"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/92955","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Sadler, Randall","Schwartz, Lane"]},{"key":"dc:creator","label":"Author","values":["Abdar, Omid"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2016-11-10T18:27:52Z","2018-11-11T10:15:32Z","2016-07-15","2016-08"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Teaching of English Sec Lang"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.A."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["English for Specific Purposes, Vocabulary, Wordlists, Natural Language Processing"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2016 Omid Abdar"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/92955"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2018-08-01","The student, Omid Abdar, accepted the attached license on 2016-07-14 at 15:51.","The student, Omid Abdar, submitted this Thesis for approval on 2016-07-14 at 16:00.","This Thesis was approved for publication on 2016-07-15 at 17:00.","DSpace SAF Submission Ingestion Package generated from Vireo submission #9948 on 2016-11-10 at 12:20:49","Made available in DSpace on 2016-11-10T18:27:52Z (GMT). No. of bitstreams: 2 ABDAR-THESIS-2016.pdf: 1388826 bytes, checksum: a5d95839aa608a1b9f4e3bbd7c996971 (MD5) LICENSE.txt: 4207 bytes, checksum: 4d43d26be895d9ee373421263076a266 (MD5) Previous issue date: 2016-07-15","Since there are thousands of words to learn in a new language, one common challenge for language learners and teachers is knowing which vocabulary items to prioritize over the others and, in general, setting vocabulary-learning goals. Within vocabulary teaching research, one approach has been to focus on lists of the most common vocabulary. West (1953) proposed a list of the 2000 most frequent word families in English that, it was argued, were most important for learners to master. Along the same lines, Coxhead (2000) offered a list of the most common words in academic English known as the Academic Word List (AWL). Arguing that AWL did not adequately reflect the learners’ specialized vocabulary needs, however, corpus linguists began to develop wordlists in specialized subject areas with an English for Specific Purposes (ESP) perspective for students in Business, Engineering, Medical, and Law majors and so on. A central theme in almost all previous endeavors to develop better wordlists has been the notion of 'representativeness'—the extent to which a wordlist 'represents' the language needs of leaners. In this study, it is proposed that an alternative way to maximize representativeness in a wordlist is to enable users to compile a wordlist from any text or corpus that is of interest to them and to provide the means of compiling a wordlist using that text. Using Natural Language Toolkit (NLTK), this study shows how a few Natural Language Processing (NLP) techniques may be used to compile a list of the most common words in the Europarl corpus along with retrieving example sentences from the corpus for each word. This new approach can have applications for both language leaners as well as for the purposes of preparing instructional materials in an ESP setting.","Embargo set by: Seth Robbins for item 95375 Lift date: 2018-11-10T18:28:02Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 95375 on 2018-11-11T10:15:32Z."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Compiling contextualized lists of frequent vocabulary from user- supplied corpora using natural language processing techniques"]}]}],"canonical_facts":{"dc:contributor":["Sadler, Randall","Schwartz, Lane"],"dc:creator":["Abdar, Omid"],"dc:date":["2016-11-10T18:27:52Z","2018-11-11T10:15:32Z","2016-07-15","2016-08"],"dc:description":["Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2018-08-01","The student, Omid Abdar, accepted the attached license on 2016-07-14 at 15:51.","The student, Omid Abdar, submitted this Thesis for approval on 2016-07-14 at 16:00.","This Thesis was approved for publication on 2016-07-15 at 17:00.","DSpace SAF Submission Ingestion Package generated from Vireo submission #9948 on 2016-11-10 at 12:20:49","Made available in DSpace on 2016-11-10T18:27:52Z (GMT). No. of bitstreams: 2 ABDAR-THESIS-2016.pdf: 1388826 bytes, checksum: a5d95839aa608a1b9f4e3bbd7c996971 (MD5) LICENSE.txt: 4207 bytes, checksum: 4d43d26be895d9ee373421263076a266 (MD5) Previous issue date: 2016-07-15","Since there are thousands of words to learn in a new language, one common challenge for language learners and teachers is knowing which vocabulary items to prioritize over the others and, in general, setting vocabulary-learning goals. Within vocabulary teaching research, one approach has been to focus on lists of the most common vocabulary. West (1953) proposed a list of the 2000 most frequent word families in English that, it was argued, were most important for learners to master. Along the same lines, Coxhead (2000) offered a list of the most common words in academic English known as the Academic Word List (AWL). Arguing that AWL did not adequately reflect the learners’ specialized vocabulary needs, however, corpus linguists began to develop wordlists in specialized subject areas with an English for Specific Purposes (ESP) perspective for students in Business, Engineering, Medical, and Law majors and so on. A central theme in almost all previous endeavors to develop better wordlists has been the notion of 'representativeness'—the extent to which a wordlist 'represents' the language needs of leaners. In this study, it is proposed that an alternative way to maximize representativeness in a wordlist is to enable users to compile a wordlist from any text or corpus that is of interest to them and to provide the means of compiling a wordlist using that text. Using Natural Language Toolkit (NLTK), this study shows how a few Natural Language Processing (NLP) techniques may be used to compile a list of the most common words in the Europarl corpus along with retrieving example sentences from the corpus for each word. This new approach can have applications for both language leaners as well as for the purposes of preparing instructional materials in an ESP setting.","Embargo set by: Seth Robbins for item 95375 Lift date: 2018-11-10T18:28:02Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 95375 on 2018-11-11T10:15:32Z."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/92955"],"dc:language":["en"],"dc:rights":["Copyright 2016 Omid Abdar"],"dc:subject":["English for Specific Purposes, Vocabulary, Wordlists, Natural Language Processing"],"dc:title":["Compiling contextualized lists of frequent vocabulary from user- supplied corpora using natural language processing techniques"],"dc:type":["text"],"thesis:degree_discipline":["Teaching of English Sec Lang"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.A."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:26:35Z"}