Back to results

University of Washington

Unsupervised Morphological Word Clustering

Abstract

dc:description.abstract

This thesis describes a system which clusters the words of a given lexicon into conflation sets (sets of morphologically related words). The word clustering is based on clustering of suffixes, which, in turn, is based on stem-suffix co-occurrence frequencies. The suffix clustering is performed as a clique clustering of a weighted undirected graph with the suffixes as vertices; the edges weights are calculated as similarity measure between the suffix signatures of the vertices according to the proposed metric. The clustering that yields the lowest lexicon compression ratio is considered the optimum. In addition, the hypothesis that the lowest compression ratio suffix clustering yields the best word clustering is tested. The system is tested on the CELEX English, German and Dutch lexicons and its performance is evaluated against the set of conflation classes extracted from the CELEX morphological database (Baayen, et al., 1993). The system performance is compared to that of other systems: Morfessor (Creutz and Lagus, 2005), Linguistica (Goldsmith, 2000), and χ^2 - significance test based word clustering approach (Moon et al., 2009).

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Lushtak, Sergei A.
Advisor dc:contributor.advisor
  • Levow, Gina-Anne

Subjects

dc:subject × 1

Rights

dc:rights
Statement dc:rights
  • Copyright is held by the individual authors.
Language dc:language.iso
en_US

Identifiers

dc:identifier.*
Handle dc:identifier.uri
http://hdl.handle.net/1773/22453
OAI identifier oai:identifier
oai:digital.lib.washington.edu:1773/22453

Chain of custody

source
Harvested from
University of Washington
Base URL
digital.lib.washington.edu/server/oai/request
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Lushtak, Sergei A.. Unsupervised Morphological Word Clustering. 2013. http://hdl.handle.net/1773/22453