{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/116178"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/116178","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Pitfalls and possibilities: What NLP systems are missing out on","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-15 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2022-11-15 without embargo terms","abstract_has_math":false,"creators":["Park, Hyunji (Hayley)"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Linguistics","degree_department":null,"school":null,"contributors":["Schwartz, Lane","Hockenmaier, Julia","Ji, Heng","Tyers, Francis M."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-08","date_published":"2022-08","updated_at":"2026-07-22T22:24:55Z","subjects":["natural language processing","computational linguistics","morphology","deep learning"],"languages":["en","eng"],"rights":["2022 Hyunji (Hayley) Park"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/116178","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Schwartz, Lane","Hockenmaier, Julia","Ji, Heng","Tyers, Francis M."]},{"key":"dc:creator","label":"Author","values":["Park, Hyunji (Hayley)"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022-08","2022-07-05"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Linguistics"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["natural language processing","computational linguistics","morphology","deep learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["2022 Hyunji (Hayley) Park"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/116178"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-15 without embargo terms","The student, Hyunji (Hayley) Park, accepted the attached license on 2022-06-30 at 14:43.","The student, Hyunji (Hayley) Park, submitted this Dissertation for approval on 2022-06-30 at 14:55.","This Dissertation was approved for publication on 2022-07-05 at 13:04.","DSpace SAF Submission Ingestion Package generated from Vireo submission #18125 on 2022-11-15 at 17:38:04","Despite recent advancements in natural language processing (NLP), there are still many areas in NLP that need much progress. In particular, this dissertation presents three studies, where careful consideration of datasets challenges the existing NLP methods. First, we present a study that augments the existing data to investigate the effect of morphology on LSTM language modeling. With NLP research disproportionally dedicated to English and a few other morphologically poor languages, the effect of morphology is clearly under-studied with a couple of previous papers disagreeing on the interaction between morphology and language modeling difficulty. By compiling a parallel Bible corpus and a linguistic typology database that represent morphological typology, we show that morphological complexity makes a language harder to model and affects the effectiveness of subword segmentation methods such as BPE. Next, we develop the first dependency treebank for St. Lawrence Island Yupik to show that morphology interacts with syntax in the polysynthetic language in the context of dependency parsing. We argue that the Universal Dependencies (UD) guidelines, which focus on word-level annotations, should be extended to morpheme-level annotations to better serve morphologically rich languages. Finally, we present a study on long document classification in English using Transformers, focusing on the validity of evaluation methods available for this newly developed task. By providing a comprehensive evaluation of existing models’ relative efficacy against various datasets and baselines, we show that existing models often fail to outperform simple baseline models and yield inconsistent performance across the datasets. The findings emphasize that future studies should consider comprehensive baselines and datasets that better represent the task of long document classification to develop robust models. In all, this dissertation sheds light on areas in NLP that need further investigation and emphasize the importance of careful consideration of the datasets involved."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Pitfalls and possibilities: What NLP systems are missing out on"]}]}],"canonical_facts":{"dc:contributor":["Schwartz, Lane","Hockenmaier, Julia","Ji, Heng","Tyers, Francis M."],"dc:creator":["Park, Hyunji (Hayley)"],"dc:date":["2022-08","2022-07-05"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-15 without embargo terms","The student, Hyunji (Hayley) Park, accepted the attached license on 2022-06-30 at 14:43.","The student, Hyunji (Hayley) Park, submitted this Dissertation for approval on 2022-06-30 at 14:55.","This Dissertation was approved for publication on 2022-07-05 at 13:04.","DSpace SAF Submission Ingestion Package generated from Vireo submission #18125 on 2022-11-15 at 17:38:04","Despite recent advancements in natural language processing (NLP), there are still many areas in NLP that need much progress. In particular, this dissertation presents three studies, where careful consideration of datasets challenges the existing NLP methods. First, we present a study that augments the existing data to investigate the effect of morphology on LSTM language modeling. With NLP research disproportionally dedicated to English and a few other morphologically poor languages, the effect of morphology is clearly under-studied with a couple of previous papers disagreeing on the interaction between morphology and language modeling difficulty. By compiling a parallel Bible corpus and a linguistic typology database that represent morphological typology, we show that morphological complexity makes a language harder to model and affects the effectiveness of subword segmentation methods such as BPE. Next, we develop the first dependency treebank for St. Lawrence Island Yupik to show that morphology interacts with syntax in the polysynthetic language in the context of dependency parsing. We argue that the Universal Dependencies (UD) guidelines, which focus on word-level annotations, should be extended to morpheme-level annotations to better serve morphologically rich languages. Finally, we present a study on long document classification in English using Transformers, focusing on the validity of evaluation methods available for this newly developed task. By providing a comprehensive evaluation of existing models’ relative efficacy against various datasets and baselines, we show that existing models often fail to outperform simple baseline models and yield inconsistent performance across the datasets. The findings emphasize that future studies should consider comprehensive baselines and datasets that better represent the task of long document classification to develop robust models. In all, this dissertation sheds light on areas in NLP that need further investigation and emphasize the importance of careful consideration of the datasets involved."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/116178"],"dc:language":["en","eng"],"dc:rights":["2022 Hyunji (Hayley) Park"],"dc:subject":["natural language processing","computational linguistics","morphology","deep learning"],"dc:title":["Pitfalls and possibilities: What NLP systems are missing out on"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Linguistics"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:55Z"}