University of Nevada, Las Vegas
Predictor of OCR accuracy using statistical techniques
Abstract
dc:description.abstractSystems that predict optical character recognition (OCR) accuracy of an input image by a given OCR system were developed. Seven features associated with image defects were identified and utilized. Two kinds of nonparametric classification engines, the nearest neighbor rule-based and neural network-based, were implemented. The performance of these systems were compared to an old heuristic-based system using a cost model of a large-scale document conversion process and a test data set consisting of 502 pages. The results show that the performance of new classifiers were better than that of the heuristic-based system. The neural network-based system outperformed the nearest-neighbor-based system. These new systems can be used to reduce the cost of a large-scale document conversion process by discriminating good quality pages for OCR from degraded images for manual data entry.
Degree
thesis:*- Name thesis:degree_name
- Master of Science (MS)
- Level thesis:degree_level
- Thesis
- Discipline thesis:degree_discipline
- Computer Science
- Grantor dc:publisher
- University of Nevada, Las Vegas
- Year
- 1996
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Gonzalez, Juan Manuel
Rights
dc:rights- Statement dc:rights
-
- IN COPYRIGHT. For more information about this rights statement, please visit http://rightsstatements.org/vocab/InC/1.0/
- Language dc:language
- English
Identifiers
dc:identifier.*- Identifier
- https://oasis.library.unlv.edu/rtds/589
- OAI identifier oai:identifier
- oai:oasis.library.unlv.edu:rtds-1588