{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/120435"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/120435","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Are language models leaking personal information? Memorization vs. Association","abstract":"Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2025-05-01","abstract_html":"Submission published under a 24 month embargo labeled &#x27;U of I Access&#x27;, the embargo will last until 2025-05-01","abstract_has_math":false,"creators":["Shao, Hanyin"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Chang, Kevin Chen-Chuan"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2023,"date_issued":"2023-05","date_published":"2023-05","updated_at":"2026-07-22T22:24:57Z","subjects":["Natural Language Processing","Nlp"],"languages":["en","eng"],"rights":["Copyright 2023 Hanyin Shao"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/120435","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Chang, Kevin Chen-Chuan"]},{"key":"dc:creator","label":"Author","values":["Shao, Hanyin"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2023-05","2023-05-02"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Natural Language Processing","Nlp"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2023 Hanyin Shao"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/120435"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2025-05-01","The student, Hanyin Shao, accepted the attached license on 2023-04-28 at 00:38.","The student, Hanyin Shao, submitted this Thesis for approval on 2023-04-28 at 00:47.","This Thesis was approved for publication on 2023-05-02 at 15:36.","DSpace SAF Submission Ingestion Package generated from Vireo submission #19229 on 2023-09-01 at 17:15:16","Large pre-trained language models (PLMs) have transformed the field of natural language processing (NLP) in recent years. PLMs have become the basis for various state-of-the-art NLP systems. Despite the great success of PLMs solving a wide range of NLP tasks, there is rising concern about privacy risks brought with PLMs. For example, recent studies show that PLMs memorize a great portion of training data, including sensitive information, while the information may be leaked unintentionally and utilized by malicious adversaries. In this thesis, we evaluate whether PLMs are prone to leaking personal information and discuss possible reasons behind the privacy leakage. Specifically, we attempt to query PLMs for a target email address with contexts of the email address or prompts containing the owner’s name. We find that PLMs do leak personal information mainly due to memorization. However, the risk of specific personal information being extracted by attackers is low because the models are weak at associating personal identifying information with its owner. We also try to quantify PLMs’ capability of association to help validate the safety of PLM in terms of privacy preserving."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Are language models leaking personal information? Memorization vs. Association"]}]}],"canonical_facts":{"dc:contributor":["Chang, Kevin Chen-Chuan"],"dc:creator":["Shao, Hanyin"],"dc:date":["2023-05","2023-05-02"],"dc:description":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2025-05-01","The student, Hanyin Shao, accepted the attached license on 2023-04-28 at 00:38.","The student, Hanyin Shao, submitted this Thesis for approval on 2023-04-28 at 00:47.","This Thesis was approved for publication on 2023-05-02 at 15:36.","DSpace SAF Submission Ingestion Package generated from Vireo submission #19229 on 2023-09-01 at 17:15:16","Large pre-trained language models (PLMs) have transformed the field of natural language processing (NLP) in recent years. PLMs have become the basis for various state-of-the-art NLP systems. Despite the great success of PLMs solving a wide range of NLP tasks, there is rising concern about privacy risks brought with PLMs. For example, recent studies show that PLMs memorize a great portion of training data, including sensitive information, while the information may be leaked unintentionally and utilized by malicious adversaries. In this thesis, we evaluate whether PLMs are prone to leaking personal information and discuss possible reasons behind the privacy leakage. Specifically, we attempt to query PLMs for a target email address with contexts of the email address or prompts containing the owner’s name. We find that PLMs do leak personal information mainly due to memorization. However, the risk of specific personal information being extracted by attackers is low because the models are weak at associating personal identifying information with its owner. We also try to quantify PLMs’ capability of association to help validate the safety of PLM in terms of privacy preserving."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/120435"],"dc:language":["en","eng"],"dc:rights":["Copyright 2023 Hanyin Shao"],"dc:subject":["Natural Language Processing","Nlp"],"dc:title":["Are language models leaking personal information? Memorization vs. Association"],"dc:type":["text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:57Z"}