{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/117739"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/117739","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"The influence of optical character recognition quality on the robustness of semantic encoding","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2023-04-12 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2023-04-12 without embargo terms","abstract_has_math":false,"creators":["Jiang, Ming"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Informatics","degree_department":null,"school":null,"contributors":["Downie, J. Stephen","Renear, Allen","Underwood, Ted","Kilicoglu, Halil","LeBlanc, Zoe"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-12","date_published":"2022-12","updated_at":"2026-07-22T22:24:56Z","subjects":["Optical Character Recognition","Word Embeddings","Semantic Encoding","Large Language Models","Robustness","Hathitrust","Digital Humanities","Digital Libraries","Data Curation"],"languages":["en","eng"],"rights":["Copyright 2022 Ming Jiang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/117739","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Downie, J. Stephen","Renear, Allen","Underwood, Ted","Kilicoglu, Halil","LeBlanc, Zoe"]},{"key":"dc:creator","label":"Author","values":["Jiang, Ming"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022-12","2022-10-27"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Informatics"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Optical Character Recognition","Word Embeddings","Semantic Encoding","Large Language Models","Robustness","Hathitrust","Digital Humanities","Digital Libraries","Data Curation"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2022 Ming Jiang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/117739"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2023-04-12 without embargo terms","The student, Ming Jiang, accepted the attached license on 2022-10-21 at 14:40.","The student, Ming Jiang, submitted this Dissertation for approval on 2022-10-21 at 14:55.","This Dissertation was approved for publication on 2022-10-27 at 09:57.","DSpace SAF Submission Ingestion Package generated from Vireo submission #18532 on 2023-04-12 at 07:24:23","Historical textual collections, digitized by machine scanning and optical character recognition (OCR), offer unique opportunities for exploring and disseminating heritage knowledge. Research innovations in this field, including recent advances in natural language processing (NLP), have been widely promoted as promising new tools for supporting research on these collections. Unfortunately, the inevitable OCR noise in these digitized materials challenges the performance of advanced NLP techniques, which are generally built for born-digital corpora. Moreover, the black-box NLP further makes it hard to understand the effects of OCR errors on NLP algorithms. This dissertation concentrates on the problem mentioned above, with a specific focus on the robustness of word embedding techniques such as word2vec, BERT, etc. for semantic encoding of OCR'd texts. We explore the problem through three interrelated parts of the studies. The first two parts compare various word embedding technologies to capture their latent characteristics on texts with OCR quality issues; Part I examines document-level encoding; Part II investigates sentence- and word-level encoding. Finally, the last part analyzes the effect of different levels of OCR noise on a specific word embedding methodology. Experimental results show that: (1) fine-tuned BERT outperforms pre-trained BERT when encoding OCR'd texts; (2) BERT-based dynamic embeddings are more sensitive to OCR errors than static embeddings in encoding words and sentences; (3) coarse-grained encoding (e.g., document-level) mitigates OCR noise interference on word embeddings, while fine-grained encoding (e.g., word-level) reduces the robustness of word embeddings to OCR noise; (4) OCR noise in unseen testing data can reduce embedding performance and downstream outcomes, while noise in the training corpus can benefit embedding robustness; and, (5) OCR noise does matter in scientific relation classification. Following our results, we recommend that scholars analyze their data with regard to both text granularity and data quality in training and testing corpora, in order to select the appropriate embedding tool for their analyses."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["The influence of optical character recognition quality on the robustness of semantic encoding"]}]}],"canonical_facts":{"dc:contributor":["Downie, J. Stephen","Renear, Allen","Underwood, Ted","Kilicoglu, Halil","LeBlanc, Zoe"],"dc:creator":["Jiang, Ming"],"dc:date":["2022-12","2022-10-27"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2023-04-12 without embargo terms","The student, Ming Jiang, accepted the attached license on 2022-10-21 at 14:40.","The student, Ming Jiang, submitted this Dissertation for approval on 2022-10-21 at 14:55.","This Dissertation was approved for publication on 2022-10-27 at 09:57.","DSpace SAF Submission Ingestion Package generated from Vireo submission #18532 on 2023-04-12 at 07:24:23","Historical textual collections, digitized by machine scanning and optical character recognition (OCR), offer unique opportunities for exploring and disseminating heritage knowledge. Research innovations in this field, including recent advances in natural language processing (NLP), have been widely promoted as promising new tools for supporting research on these collections. Unfortunately, the inevitable OCR noise in these digitized materials challenges the performance of advanced NLP techniques, which are generally built for born-digital corpora. Moreover, the black-box NLP further makes it hard to understand the effects of OCR errors on NLP algorithms. This dissertation concentrates on the problem mentioned above, with a specific focus on the robustness of word embedding techniques such as word2vec, BERT, etc. for semantic encoding of OCR'd texts. We explore the problem through three interrelated parts of the studies. The first two parts compare various word embedding technologies to capture their latent characteristics on texts with OCR quality issues; Part I examines document-level encoding; Part II investigates sentence- and word-level encoding. Finally, the last part analyzes the effect of different levels of OCR noise on a specific word embedding methodology. Experimental results show that: (1) fine-tuned BERT outperforms pre-trained BERT when encoding OCR'd texts; (2) BERT-based dynamic embeddings are more sensitive to OCR errors than static embeddings in encoding words and sentences; (3) coarse-grained encoding (e.g., document-level) mitigates OCR noise interference on word embeddings, while fine-grained encoding (e.g., word-level) reduces the robustness of word embeddings to OCR noise; (4) OCR noise in unseen testing data can reduce embedding performance and downstream outcomes, while noise in the training corpus can benefit embedding robustness; and, (5) OCR noise does matter in scientific relation classification. Following our results, we recommend that scholars analyze their data with regard to both text granularity and data quality in training and testing corpora, in order to select the appropriate embedding tool for their analyses."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/117739"],"dc:language":["en","eng"],"dc:rights":["Copyright 2022 Ming Jiang"],"dc:subject":["Optical Character Recognition","Word Embeddings","Semantic Encoding","Large Language Models","Robustness","Hathitrust","Digital Humanities","Digital Libraries","Data Curation"],"dc:title":["The influence of optical character recognition quality on the robustness of semantic encoding"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Informatics"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:56Z"}