{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/124388"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/124388","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Knowledge representation and behavior understanding with pre-trained language models","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-09-16 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2024-09-16 without embargo terms","abstract_has_math":false,"creators":["Jiang, Minhao"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Han, Jiawei"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-05","date_published":"2024-05","updated_at":"2026-07-22T22:25:00Z","subjects":["Pre-trained Language Models","Knowledge Representation","Data Contamination"],"languages":["en","eng"],"rights":["Copyright 2024 Minhao Jiang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/124388","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Han, Jiawei"]},{"key":"dc:creator","label":"Author","values":["Jiang, Minhao"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2024-05","2024-05-01"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Pre-trained Language Models","Knowledge Representation","Data Contamination"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2024 Minhao Jiang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/124388"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-09-16 without embargo terms","The student, Minhao Jiang, accepted the attached license on 2024-04-23 at 16:48.","The student, Minhao Jiang, submitted this Thesis for approval on 2024-04-23 at 18:38.","This Thesis was approved for publication on 2024-05-01 at 10:09.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20593 on 2024-09-16 at 00:36:04","The advent of pre-trained language models has marked a significant milestone in the realm of computational linguistics and data mining, showcasing remarkable performances across diverse domains. These models, endowed with vast training corpora and formidable capacity for semantic knowledge representation, have transformed the landscape of different downstream applications. Among the various applications, the automatic construction and completion of taxonomies have shown to be beneficial in enhancing numerous downstream tasks and minimizing human labor in domain-specific taxonomy development. In this work, we first introduce a taxonomy completion framework that effectively leverages pre-trained language models to extract structural and semantic information from the existing taxonomy to significantly boost the performance of current taxonomy expansion and completion frameworks. On the other hand, even though the performances of pre-trained language models are very high in many datasets, the implications of data contamination during the pre-training stage of language models are still unclear in the current literature. Given their demonstrated prowess in enhancing task performance across diverse downstream applications, concerns arise regarding the authenticity of these capabilities, potentially inflated by the inadvertent inclusion of evaluation datasets within pre-training corpora. Through meticulous experimental investigation, this study endeavors to elucidate the effects of data contamination, emphasizing the imperative for more precise definitions and stringent methodologies to fortify LLMs against such vulnerabilities."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Knowledge representation and behavior understanding with pre-trained language models"]}]}],"canonical_facts":{"dc:contributor":["Han, Jiawei"],"dc:creator":["Jiang, Minhao"],"dc:date":["2024-05","2024-05-01"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-09-16 without embargo terms","The student, Minhao Jiang, accepted the attached license on 2024-04-23 at 16:48.","The student, Minhao Jiang, submitted this Thesis for approval on 2024-04-23 at 18:38.","This Thesis was approved for publication on 2024-05-01 at 10:09.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20593 on 2024-09-16 at 00:36:04","The advent of pre-trained language models has marked a significant milestone in the realm of computational linguistics and data mining, showcasing remarkable performances across diverse domains. These models, endowed with vast training corpora and formidable capacity for semantic knowledge representation, have transformed the landscape of different downstream applications. Among the various applications, the automatic construction and completion of taxonomies have shown to be beneficial in enhancing numerous downstream tasks and minimizing human labor in domain-specific taxonomy development. In this work, we first introduce a taxonomy completion framework that effectively leverages pre-trained language models to extract structural and semantic information from the existing taxonomy to significantly boost the performance of current taxonomy expansion and completion frameworks. On the other hand, even though the performances of pre-trained language models are very high in many datasets, the implications of data contamination during the pre-training stage of language models are still unclear in the current literature. Given their demonstrated prowess in enhancing task performance across diverse downstream applications, concerns arise regarding the authenticity of these capabilities, potentially inflated by the inadvertent inclusion of evaluation datasets within pre-training corpora. Through meticulous experimental investigation, this study endeavors to elucidate the effects of data contamination, emphasizing the imperative for more precise definitions and stringent methodologies to fortify LLMs against such vulnerabilities."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/124388"],"dc:language":["en","eng"],"dc:rights":["Copyright 2024 Minhao Jiang"],"dc:subject":["Pre-trained Language Models","Knowledge Representation","Data Contamination"],"dc:title":["Knowledge representation and behavior understanding with pre-trained language models"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:00Z"}