{"id":{"repo_id":"unlv","oai_identifier":"oai:oasis.library.unlv.edu:rtds-1436"},"canonical_url":"https://search.dev.ndltd.org/etd/unlv/oai:oasis.library.unlv.edu:rtds-1436","repository":{"repo_id":"unlv","name":"University of Nevada - Las Vegas","base_url":"https://oasis.library.unlv.edu/do/oai/"},"display":{"title":"Autotag: A tool for creating structured document collections from printed materials","abstract":"Today's optical character recognition (OCR) devices ordinarily are not capable of delimiting or \"marking up\" specific structural information about the document such as the title, its authors, and titles of sections. Such information appears in the OCR device output, but would require a human to go through the output to locate the information. This type of information is highly useful for information retrieval (IR), allowing users much more flexibility in making queries of a retrieval system. This thesis will describe the design, implementation, and evaluation of a software system called Autotag. This system will automatically markup structural information in OCR-generated text. It will also establish a mapping between objects in page images and their corresponding ASCII representation. This mapping can then be used to design flexible image-based interfaces for information retrieval related applications.","abstract_html":"Today&#x27;s optical character recognition (OCR) devices ordinarily are not capable of delimiting or &quot;marking up&quot; specific structural information about the document such as the title, its authors, and titles of sections. Such information appears in the OCR device output, but would require a human to go through the output to locate the information. This type of information is highly useful for information retrieval (IR), allowing users much more flexibility in making queries of a retrieval system. This thesis will describe the design, implementation, and evaluation of a software system called Autotag. This system will automatically markup structural information in OCR-generated text. It will also establish a mapping between objects in page images and their corresponding ASCII representation. This mapping can then be used to design flexible image-based interfaces for information retrieval related applications.","abstract_has_math":false,"creators":["Condit, Allen S"],"institution":"University of Nevada, Las Vegas","degree_name":"Master of Science (MS)","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":1994,"date_issued":"1994-01-01T08:00:00Z","date_published":"1994-01-01T08:00:00Z","updated_at":"2026-07-24T05:24:31Z","subjects":[],"languages":["English"],"rights":["IN COPYRIGHT. For more information about this rights statement, please visit http://rightsstatements.org/vocab/InC/1.0/"],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["https://oasis.library.unlv.edu/rtds/437"],"render_values":[{"text":"https://oasis.library.unlv.edu/rtds/437","href":"https://oasis.library.unlv.edu/rtds/437","code":true}]}]},"links":{"outbound_url":"https://doi.org/10.25669/hhva-1mg0","outbound_label":"DOI","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["Condit, Allen S"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:publisher","label":"Institution","values":["University of Nevada, Las Vegas"]},{"key":"dc:type","label":"Dc Type","values":["Text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science (MS)"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["English"]},{"key":"dc:rights","label":"Dc Rights","values":["IN COPYRIGHT. For more information about this rights statement, please visit http://rightsstatements.org/vocab/InC/1.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["10.25669/hhva-1mg0","https://oasis.library.unlv.edu/rtds/437","https://oasis.library.unlv.edu/context/rtds/article/1436/viewcontent/uc.pdf"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Today's optical character recognition (OCR) devices ordinarily are not capable of delimiting or \"marking up\" specific structural information about the document such as the title, its authors, and titles of sections. Such information appears in the OCR device output, but would require a human to go through the output to locate the information. This type of information is highly useful for information retrieval (IR), allowing users much more flexibility in making queries of a retrieval system. This thesis will describe the design, implementation, and evaluation of a software system called Autotag. This system will automatically markup structural information in OCR-generated text. It will also establish a mapping between objects in page images and their corresponding ASCII representation. This mapping can then be used to design flexible image-based interfaces for information retrieval related applications."]},{"key":"dc:format","label":"Dc Format","values":["pdf"]},{"key":"dc:title","label":"Title","values":["Autotag: A tool for creating structured document collections from printed materials"]}]}],"canonical_facts":{"dc:creator":["Condit, Allen S"],"dc:description.abstract":["Today's optical character recognition (OCR) devices ordinarily are not capable of delimiting or \"marking up\" specific structural information about the document such as the title, its authors, and titles of sections. Such information appears in the OCR device output, but would require a human to go through the output to locate the information. This type of information is highly useful for information retrieval (IR), allowing users much more flexibility in making queries of a retrieval system. This thesis will describe the design, implementation, and evaluation of a software system called Autotag. This system will automatically markup structural information in OCR-generated text. It will also establish a mapping between objects in page images and their corresponding ASCII representation. This mapping can then be used to design flexible image-based interfaces for information retrieval related applications."],"dc:format":["pdf"],"dc:identifier":["10.25669/hhva-1mg0","https://oasis.library.unlv.edu/rtds/437","https://oasis.library.unlv.edu/context/rtds/article/1436/viewcontent/uc.pdf"],"dc:language":["English"],"dc:publisher":["University of Nevada, Las Vegas"],"dc:rights":["IN COPYRIGHT. For more information about this rights statement, please visit http://rightsstatements.org/vocab/InC/1.0/"],"dc:title":["Autotag: A tool for creating structured document collections from printed materials"],"dc:type":["Text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["Master of Science (MS)"]},"updated_at":"2026-07-24T05:24:31Z"}