{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/120501"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/120501","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Towards the processing of semantic idiomaticity in natural language","abstract":"Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2025-05-01","abstract_html":"Submission published under a 24 month embargo labeled &#x27;Closed Access&#x27;, the embargo will last until 2025-05-01","abstract_has_math":false,"creators":["Zeng, Ziheng"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Bhat, Suma","Varshney, Lav","Hasegawa-Johnson, Mark","Viswanath, Pramod"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2023,"date_issued":"2023-05","date_published":"2023-05","updated_at":"2026-07-22T22:24:57Z","subjects":["Natural Language Processing","Idiomatic Expression","Non-compositional Language","Idiom Identification","Paraphrase Generation","Representation Learning","Commonsense Knowledge Graph"],"languages":["en","eng"],"rights":["Copyright 2023 Ziheng Zeng"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/120501","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Bhat, Suma","Varshney, Lav","Hasegawa-Johnson, Mark","Viswanath, Pramod"]},{"key":"dc:creator","label":"Author","values":["Zeng, Ziheng"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2023-05","2023-03-24"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Natural Language Processing","Idiomatic Expression","Non-compositional Language","Idiom Identification","Paraphrase Generation","Representation Learning","Commonsense Knowledge Graph"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2023 Ziheng Zeng"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/120501"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2025-05-01","The student, Ziheng Zeng, accepted the attached license on 2023-03-23 at 14:42.","The student, Ziheng Zeng, submitted this Dissertation for approval on 2023-03-23 at 14:50.","This Dissertation was approved for publication on 2023-03-24 at 11:20.","DSpace SAF Submission Ingestion Package generated from Vireo submission #18892 on 2023-09-01 at 17:20:12","Pre-trained language models (PTLMs) are the cornerstone of recent breakthroughs in natural language processing (NLP). They enable machines to comprehend, analyze, and utilize human language in a multitude of applications. Despite their impressive abilities, PTLMs face challenges when dealing with idiomatic expressions (IEs), a common form of figurative language. IEs exhibit semantic non-compositionality, meaning that their figurative meaning cannot be derived from their constituent words, and contextual ambiguity, meaning that they can be interpreted literally depending on the context. This dissertation aims to explore the challenges posed by IEs to PTLMs and to advance the development of more effective NLP systems capable of processing semantic idiomaticity both explicitly and implicitly. To address the problem explicitly, we propose methods to detect and remove IEs from natural sentences and thus directly reduce the influences from IEs. Specifically, we investigate an IE identification approach that detects and localizes expressions that exhibit non-compositionality and an IE paraphrasing method that replaces the expressions with their literal counterpart phrases. To address the issue implicitly, we tackle the more fundamental issue of IE comprehension by enabling effective IE representation and introducing commonsense knowledge to PTLMs. We study methods to enhance PTLMs' ability to produce semantically meaningful and contextually appropriate embeddings for IEs. Additionally, we create an IE-centered commonsense knowledge graph and train PTLMs to become commonsense knowledge models capable of understanding and generating inferential knowledge on IE uses. Our research shows that with enhanced innate IE comprehension ability, PTLMs' performance on IE processing tasks and IE-related language understanding tasks would improve naturally. By pursuing both explicit and implicit approaches to processing semantic idiomaticity, our research has accomplished the overarching goal of extending the foundational capacity of PTLMs to perform natural language understanding in the presence of IEs."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Towards the processing of semantic idiomaticity in natural language"]}]}],"canonical_facts":{"dc:contributor":["Bhat, Suma","Varshney, Lav","Hasegawa-Johnson, Mark","Viswanath, Pramod"],"dc:creator":["Zeng, Ziheng"],"dc:date":["2023-05","2023-03-24"],"dc:description":["Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2025-05-01","The student, Ziheng Zeng, accepted the attached license on 2023-03-23 at 14:42.","The student, Ziheng Zeng, submitted this Dissertation for approval on 2023-03-23 at 14:50.","This Dissertation was approved for publication on 2023-03-24 at 11:20.","DSpace SAF Submission Ingestion Package generated from Vireo submission #18892 on 2023-09-01 at 17:20:12","Pre-trained language models (PTLMs) are the cornerstone of recent breakthroughs in natural language processing (NLP). They enable machines to comprehend, analyze, and utilize human language in a multitude of applications. Despite their impressive abilities, PTLMs face challenges when dealing with idiomatic expressions (IEs), a common form of figurative language. IEs exhibit semantic non-compositionality, meaning that their figurative meaning cannot be derived from their constituent words, and contextual ambiguity, meaning that they can be interpreted literally depending on the context. This dissertation aims to explore the challenges posed by IEs to PTLMs and to advance the development of more effective NLP systems capable of processing semantic idiomaticity both explicitly and implicitly. To address the problem explicitly, we propose methods to detect and remove IEs from natural sentences and thus directly reduce the influences from IEs. Specifically, we investigate an IE identification approach that detects and localizes expressions that exhibit non-compositionality and an IE paraphrasing method that replaces the expressions with their literal counterpart phrases. To address the issue implicitly, we tackle the more fundamental issue of IE comprehension by enabling effective IE representation and introducing commonsense knowledge to PTLMs. We study methods to enhance PTLMs' ability to produce semantically meaningful and contextually appropriate embeddings for IEs. Additionally, we create an IE-centered commonsense knowledge graph and train PTLMs to become commonsense knowledge models capable of understanding and generating inferential knowledge on IE uses. Our research shows that with enhanced innate IE comprehension ability, PTLMs' performance on IE processing tasks and IE-related language understanding tasks would improve naturally. By pursuing both explicit and implicit approaches to processing semantic idiomaticity, our research has accomplished the overarching goal of extending the foundational capacity of PTLMs to perform natural language understanding in the presence of IEs."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/120501"],"dc:language":["en","eng"],"dc:rights":["Copyright 2023 Ziheng Zeng"],"dc:subject":["Natural Language Processing","Idiomatic Expression","Non-compositional Language","Idiom Identification","Paraphrase Generation","Representation Learning","Commonsense Knowledge Graph"],"dc:title":["Towards the processing of semantic idiomaticity in natural language"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:57Z"}