University of Illinois Urbana-Champaign
Crafting safe human-centric agents with risk intelligence
Abstract
dc:descriptionSafety has become a critical concern in digital communication environments, where harmful content can propagate rapidly and compromise societal well-being. As artificial intelligence increasingly serves as an automated content creation and distribution mechanism, the risks of adverse algorithmic impacts are amplified when these systems lack awareness of how their outputs influence audiences. This dissertation addresses this challenge by developing an integrated framework for communication risk management that enables systems to assess, anticipate, and mitigate potential negative consequences. The framework comprises four interconnected research contributions. First, we establish the foundations for risk assessment through a novel task formulation and dataset that captures how identical messages affect diverse user personas differently. This approach transcends traditional content moderation by evaluating information safety and appropriateness for different population groups, creating a measurement capability essential for risk-aware AI systems. Furthermore, we address the challenge of evaluating risk for users with minimal digital footprints, often referred to as "lurkers". By leveraging social graphs constructed through large language models, this work presents a solution for more accurately predicting opinions from such users, expanding the applicability of personalized agents in risk management. Another cornerstone of this dissertation is the development of efficient strategies for language model personalization in risk assessment. Through hierarchical and collaborative data refinement techniques, our Persona-DB approach achieves high-accuracy personalization with substantially reduced data retrieval requirements. This work bridges the gap between risk management and practical deployment by addressing the computational costs of personalization. Lastly, we extend risk assessment beyond immediate impacts to encompass temporal dynamics and cascading effects. By employing language models as social simulators, our framework projects how content might influence populations over time, enabling the anticipation of long-term consequences that traditional safety approaches neglect. This capability not only enhances risk assessment but also provides a mechanism for aligning content-generating AI with broader safety considerations.
Degree
thesis:*- Name thesis:degree_name
- Ph.D.
- Level thesis:degree_level
- Dissertation
- Discipline thesis:degree_discipline
- Computer Science
- Grantor
- University of Illinois Urbana-Champaign
- Year dc:date
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Sun, Chenkai
- Contributors dc:contributor
-
- Ji, Heng
- Zhai, ChengXiang
- Han, Jiawei
- Bendersky, Michael
- Small, Kevin
Subjects
dc:subject × 6Rights
dc:rights- Statement dc:rights
-
- Copyright 2025 Chenkai Sun
- Language dc:language
- en, eng
Identifiers
dc:identifier.*- Handle dc:identifier
- https://hdl.handle.net/2142/129193