University of Illinois at Urbana-Champaign
An Evaluation of Text Classification Methods for Literary Study
Abstract
dc:descriptionSome of our conclusions are consistent with what are obtained in topic classification, such as Odds Ratio does not improve SVM performance and stop word removal might harm classification. Some conclusions contradict previous results, such as SVM does not beat naive Bayes in both cases. Some findings are new to this area---SVM and naive Bayes select top features in different frequency ranges; stemming might harm feature selection methods. These experiment results provide new insights to the relation between classification methods, feature engineering options and non-topic document properties. They also provide guidance for classification method selection in literary text classification applications.
Degree
thesis:*- Name thesis:degree_name
- Ph.D.
- Level thesis:degree_level
- Dissertation
- Discipline thesis:degree_discipline
- Library and Information Science
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2015
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Yu, Bei
- Contributors dc:contributor
-
- Linda Smith
Subjects
dc:subject × 1Rights
- Language dc:language
- eng
Identifiers
dc:identifier.*- Identifier
- (MiAaPQ)AAI3250350
- OAI identifier oai:identifier
- oai:www.ideals.illinois.edu:2142/81543