{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/129193"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/129193","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Crafting safe human-centric agents with risk intelligence","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2025-10-19 without embargo terms","abstract_has_math":false,"creators":["Sun, Chenkai"],"institution":"University of Illinois Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Ji, Heng","Zhai, ChengXiang","Han, Jiawei","Bendersky, Michael","Small, Kevin"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-04-18","date_published":"2025-04-18","updated_at":"2026-07-22T22:25:04Z","subjects":["AI Safety","Large Language Models","AI Alignment","Personalization","Social Simulation","Efficiency"],"languages":["en","eng"],"rights":["Copyright 2025 Chenkai Sun"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/129193","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Ji, Heng","Zhai, ChengXiang","Han, Jiawei","Bendersky, Michael","Small, Kevin"]},{"key":"dc:creator","label":"Author","values":["Sun, Chenkai"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-04-18","2025-05"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["AI Safety","Large Language Models","AI Alignment","Personalization","Social Simulation","Efficiency"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Chenkai Sun"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/129193"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","The student, Chenkai Sun, accepted the attached license on 2025-04-17 at 17:14.","The student, Chenkai Sun, submitted this Dissertation for approval on 2025-04-17 at 17:19.","This Dissertation was approved for publication on 2025-04-18 at 10:30.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21745 on 2025-10-19 at 18:09:18","Safety has become a critical concern in digital communication environments, where harmful content can propagate rapidly and compromise societal well-being. As artificial intelligence increasingly serves as an automated content creation and distribution mechanism, the risks of adverse algorithmic impacts are amplified when these systems lack awareness of how their outputs influence audiences. This dissertation addresses this challenge by developing an integrated framework for communication risk management that enables systems to assess, anticipate, and mitigate potential negative consequences. The framework comprises four interconnected research contributions. First, we establish the foundations for risk assessment through a novel task formulation and dataset that captures how identical messages affect diverse user personas differently. This approach transcends traditional content moderation by evaluating information safety and appropriateness for different population groups, creating a measurement capability essential for risk-aware AI systems. Furthermore, we address the challenge of evaluating risk for users with minimal digital footprints, often referred to as \"lurkers\". By leveraging social graphs constructed through large language models, this work presents a solution for more accurately predicting opinions from such users, expanding the applicability of personalized agents in risk management. Another cornerstone of this dissertation is the development of efficient strategies for language model personalization in risk assessment. Through hierarchical and collaborative data refinement techniques, our Persona-DB approach achieves high-accuracy personalization with substantially reduced data retrieval requirements. This work bridges the gap between risk management and practical deployment by addressing the computational costs of personalization. Lastly, we extend risk assessment beyond immediate impacts to encompass temporal dynamics and cascading effects. By employing language models as social simulators, our framework projects how content might influence populations over time, enabling the anticipation of long-term consequences that traditional safety approaches neglect. This capability not only enhances risk assessment but also provides a mechanism for aligning content-generating AI with broader safety considerations."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Crafting safe human-centric agents with risk intelligence"]}]}],"canonical_facts":{"dc:contributor":["Ji, Heng","Zhai, ChengXiang","Han, Jiawei","Bendersky, Michael","Small, Kevin"],"dc:creator":["Sun, Chenkai"],"dc:date":["2025-04-18","2025-05"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","The student, Chenkai Sun, accepted the attached license on 2025-04-17 at 17:14.","The student, Chenkai Sun, submitted this Dissertation for approval on 2025-04-17 at 17:19.","This Dissertation was approved for publication on 2025-04-18 at 10:30.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21745 on 2025-10-19 at 18:09:18","Safety has become a critical concern in digital communication environments, where harmful content can propagate rapidly and compromise societal well-being. As artificial intelligence increasingly serves as an automated content creation and distribution mechanism, the risks of adverse algorithmic impacts are amplified when these systems lack awareness of how their outputs influence audiences. This dissertation addresses this challenge by developing an integrated framework for communication risk management that enables systems to assess, anticipate, and mitigate potential negative consequences. The framework comprises four interconnected research contributions. First, we establish the foundations for risk assessment through a novel task formulation and dataset that captures how identical messages affect diverse user personas differently. This approach transcends traditional content moderation by evaluating information safety and appropriateness for different population groups, creating a measurement capability essential for risk-aware AI systems. Furthermore, we address the challenge of evaluating risk for users with minimal digital footprints, often referred to as \"lurkers\". By leveraging social graphs constructed through large language models, this work presents a solution for more accurately predicting opinions from such users, expanding the applicability of personalized agents in risk management. Another cornerstone of this dissertation is the development of efficient strategies for language model personalization in risk assessment. Through hierarchical and collaborative data refinement techniques, our Persona-DB approach achieves high-accuracy personalization with substantially reduced data retrieval requirements. This work bridges the gap between risk management and practical deployment by addressing the computational costs of personalization. Lastly, we extend risk assessment beyond immediate impacts to encompass temporal dynamics and cascading effects. By employing language models as social simulators, our framework projects how content might influence populations over time, enabling the anticipation of long-term consequences that traditional safety approaches neglect. This capability not only enhances risk assessment but also provides a mechanism for aligning content-generating AI with broader safety considerations."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/129193"],"dc:language":["en","eng"],"dc:rights":["Copyright 2025 Chenkai Sun"],"dc:subject":["AI Safety","Large Language Models","AI Alignment","Personalization","Social Simulation","Efficiency"],"dc:title":["Crafting safe human-centric agents with risk intelligence"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:04Z"}