University of Nevada, Las Vegas
Thesaurus-aided learning for rule-based categorization of Ocr texts
Abstract
dc:description.abstractThe question posed in this thesis is whether the effectiveness of the rule-based approach to automatic text categorization on OCR collections can be improved by using domain-specific thesauri. A rule-based categorizer was constructed consisting of a C++ program called C-KANT which consults documents and creates a program which can be executed by the CLIPS expert system shell. A series of tests using domain-specific thesauri revealed that a query expansion approach to rule-based automatic text categorization using domain-dependent thesauri will not improve the categorization of OCR texts. Although some improvement to categorization could be made using rules over a mixture of thesauri, the improvements were not significantly large.
Degree
thesis:*- Name thesis:degree_name
- Master of Science (MS)
- Level thesis:degree_level
- Thesis
- Discipline thesis:degree_discipline
- Computer Science
- Grantor dc:publisher
- University of Nevada, Las Vegas
- Year
- 2003
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Coombs, Jeffrey Scott
- Contributors dc:contributor
-
- Kazem Tagva
Rights
dc:rights- Statement dc:rights
-
- IN COPYRIGHT. For more information about this rights statement, please visit http://rightsstatements.org/vocab/InC/1.0/
- Language dc:language
- English
Identifiers
dc:identifier.*- Identifier
- https://oasis.library.unlv.edu/rtds/1614
- OAI identifier oai:identifier
- oai:oasis.library.unlv.edu:rtds-2613