{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/129900"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/129900","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Differential privacy in the era of generative AI: promises and challenges","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-20 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2025-10-20 without embargo terms","abstract_has_math":false,"creators":["Wu, Fan"],"institution":"University of Illinois Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Forsyth, David A.","Chandrasekaran, Varun","Forsyth, David A","Wang, Gang","Peng, Hao","Kohno, Tadayoshi"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-05-30","date_published":"2025-05-30","updated_at":"2026-07-22T22:25:06Z","subjects":["Differential Privacy","Machine Learning","Generative Ai","Large Language Models"],"languages":["en","eng"],"rights":["Copyright 2025 Fan Wu"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/129900","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Forsyth, David A.","Chandrasekaran, Varun","Forsyth, David A","Wang, Gang","Peng, Hao","Kohno, Tadayoshi"]},{"key":"dc:creator","label":"Author","values":["Wu, Fan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-05-30","2025-08"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Differential Privacy","Machine Learning","Generative Ai","Large Language Models"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Fan Wu"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/129900"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-20 without embargo terms","The student, Fan Wu, accepted the attached license on 2025-05-29 at 18:03.","The student, Fan Wu, submitted this Dissertation for approval on 2025-05-29 at 18:11.","This Dissertation was approved for publication on 2025-05-30 at 10:15.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22316 on 2025-10-20 at 20:14:44","Large language models (LLMs) are seeing rapid development and widespread deployment. As these models become increasingly capable and are deployed across diverse domains involving sensitive data, privacy concerns have intensified. Their inadvertently memorizing and leaking private information creates significant privacy risks when they are fine-tuned with user data or deployed as interactive agents. This thesis addresses the critical privacy challenges emerging in the era of generative AI, with a particular focus on protecting training data privacy in LLMs across various learning paradigms and application scenarios, as well as understanding what protection we actually offer. As a central tool, we leverage and scrutinize differential privacy (DP). Concretely, we develop a novel DP framework for language model alignment through preference tuning (RLHF), formalize new privacy definitions for multi-user training data scenarios, and critically examine DP-SGD—the workhorse algorithm for DP LLM training—and reveal an alarming variance in its empirical privacy protection. Together, these contributions advance both the practical applications and fundamental understanding of differential privacy in LLMs, providing researchers and practitioners with new tools and insights to navigate the landscape of privacy in generative AI."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Differential privacy in the era of generative AI: promises and challenges"]}]}],"canonical_facts":{"dc:contributor":["Forsyth, David A.","Chandrasekaran, Varun","Forsyth, David A","Wang, Gang","Peng, Hao","Kohno, Tadayoshi"],"dc:creator":["Wu, Fan"],"dc:date":["2025-05-30","2025-08"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-20 without embargo terms","The student, Fan Wu, accepted the attached license on 2025-05-29 at 18:03.","The student, Fan Wu, submitted this Dissertation for approval on 2025-05-29 at 18:11.","This Dissertation was approved for publication on 2025-05-30 at 10:15.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22316 on 2025-10-20 at 20:14:44","Large language models (LLMs) are seeing rapid development and widespread deployment. As these models become increasingly capable and are deployed across diverse domains involving sensitive data, privacy concerns have intensified. Their inadvertently memorizing and leaking private information creates significant privacy risks when they are fine-tuned with user data or deployed as interactive agents. This thesis addresses the critical privacy challenges emerging in the era of generative AI, with a particular focus on protecting training data privacy in LLMs across various learning paradigms and application scenarios, as well as understanding what protection we actually offer. As a central tool, we leverage and scrutinize differential privacy (DP). Concretely, we develop a novel DP framework for language model alignment through preference tuning (RLHF), formalize new privacy definitions for multi-user training data scenarios, and critically examine DP-SGD—the workhorse algorithm for DP LLM training—and reveal an alarming variance in its empirical privacy protection. Together, these contributions advance both the practical applications and fundamental understanding of differential privacy in LLMs, providing researchers and practitioners with new tools and insights to navigate the landscape of privacy in generative AI."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/129900"],"dc:language":["en","eng"],"dc:rights":["Copyright 2025 Fan Wu"],"dc:subject":["Differential Privacy","Machine Learning","Generative Ai","Large Language Models"],"dc:title":["Differential privacy in the era of generative AI: promises and challenges"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:06Z"}