Back to results

Virginia Tech

Automatic Lexicon Generation for Unsupervised Part-of-Speech Tagging Using Only Unannotated Text

Abstract

dc:description.abstract

With the growing number of textual resources available, the ability to understand them becomes critical. An essential first step in understanding these sources is the ability to identify the parts-of-speech in each sentence. The goal of this research is to propose, improve, and implement an algorithm capable of finding terms (words in a corpus) that are used in similar ways--a term categorizer. Such a term categorizer can be used to find a particular part-of-speech, i.e. nouns in a corpus, and generate a lexicon. The proposed work is not dependent on any external sources of information, such as dictionaries, and it shows a significant improvement (~30%) over an existing method of categorization. More importantly, the proposed algorithm can be applied as a component of an unsupervised part-of-speech tagger, making it truly unsupervised, requiring only unannotated text. The algorithm is discussed in detail, along with its background, and its performance. Experimentation shows that the proposed algorithm performs within 3% of the baseline, the Penn-TreeBank Lexicon.

Degree

thesis:*
Name thesis:degree_name
Master of Science
Level thesis:degree_level
masters
Discipline thesis:degree_discipline
Computer Science
Department dc:contributor.department
Computer Science
Grantor dc:publisher
Virginia Tech
Year dc:date.issued
1999

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Pereira, Dennis V.
Chair dc:contributor.committeechair
  • Egyhazy, Csaba J.
Committee members dc:contributor.committeemember
  • Belli, Gabriella M.
  • Frakes, William B.

Subjects

dc:subject × 5

Rights

dc:rights
Statement dc:rights
  • In Copyright

Identifiers

dc:identifier.*
Dc Identifier Other
etd-08242004-012316
OAI identifier oai:identifier
oai:vtechworks.lib.vt.edu:10919/10094

Chain of custody

source
Harvested from
Virginia Tech
Base URL
vtechworks.lib.vt.edu/oai/request
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Pereira, Dennis V.. Automatic Lexicon Generation for Unsupervised Part-of-Speech Tagging Using Only Unannotated Text. masters thesis, Virginia Tech, 1999. http://hdl.handle.net/10919/10094