{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/129175"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/129175","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Examining large language models for safety and robustness through the lens of social science","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2025-10-19 without embargo terms","abstract_has_math":false,"creators":["Jeoung, Sullam"],"institution":"University of Illinois Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Information Sciences","degree_department":null,"school":null,"contributors":["Diesner, Jana","Kilicoglu, Halil","Bosh, Nigel","Wang, Haohan"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-03-07","date_published":"2025-03-07","updated_at":"2026-07-22T22:25:04Z","subjects":["Large Language Model","Responsible AI"],"languages":["en","eng"],"rights":["Copyright 2025 Sullam Jeoung"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/129175","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Diesner, Jana","Kilicoglu, Halil","Bosh, Nigel","Wang, Haohan"]},{"key":"dc:creator","label":"Author","values":["Jeoung, Sullam"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-03-07","2025-05"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Information Sciences"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Large Language Model","Responsible AI"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Sullam Jeoung"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/129175"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","The student, Sullam Jeoung, accepted the attached license on 2025-03-05 at 10:51.","The student, Sullam Jeoung, submitted this Dissertation for approval on 2025-03-05 at 10:55.","This Dissertation was approved for publication on 2025-03-07 at 10:36.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21662 on 2025-10-19 at 18:09:09","Large language models have demonstrated remarkable capabilities, often achieving human-like performance levels and significantly impacting our daily lives. However, these models can perpetuate and amplify harmful stereotypes and biases associated with socio-demographic representations, potentially generating discriminatory content that adversely affects individuals and communities. Given these risks and their broader societal implications, ensuring the safety and robustness of these models through the identification and mitigation of harmful stereotypes has become imperative. This dissertation presents comprehensive methodologies to address these challenges by integrating insights from social science, psychology, and cognitive studies with methods from natural language processing. First, we present a framework to assess human-like stereotypical patterns in large language models (LLMs), drawing upon established psychological theories of how individuals develop stereotypes toward various social groups. This theoretically-grounded approach provides construct validity in defining and measuring stereotypes. The framework incorporates three key dimensions: warmth-competence analysis, keyword- reasoning patterns, and emotional-behavioral responses. Through clustering analysis, keyword extraction, and reasoning pattern evaluation of LLM responses, we examine how these models align with or deviate from documented human behavioral patterns. Our findings reveal that LLMs demonstrate nuanced perceptions of social groups, consistent with psychological research highlighting the multifaceted nature of stereotypes. Notably, the models’ reasoning patterns, particularly regarding groups’ economic status, demonstrate a nuanced awareness of societal disparities. Second, we propose methods to examine causal sensitivity of language models on socio-demographic attributes. This is based on a controlled experimental framework that uses name frequency analysis from U.S. Census data and systematic evaluation of model predictions through causal graphs. Our findings show that less frequent first names lead to divergent model predictions, highlighting the need for careful demographic consideration in dataset design to ensure fair and consistent model performance across different name representations. Third, we investigate how LLMs exhibit and inflate political stereotypes through the lens of cognitive biases and representative heuristics. We analyze LLMs’ responses using two key theoretical frameworks: ’kernel of truth’ (whether stereotypes reflect empirical realities) and ’representative heuristics’ (whether models overemphasize representative attributes of target groups), comparing model outputs with actual human responses across various political topics. Our findings show that while LLMs can accurately mimic certain political positions, they tend to exaggerate these positions compared to empirical human responses, suggesting a vulnerability to stereotypical thinking similar to human cognitive biases. This implies the need for careful consideration of cognitive bias frameworks in developing and deploying language models, particularly in politically sensitive contexts, and demonstrates the potential effectiveness of prompt-based mitigation strategies in reducing stereotypical responses. Overall, this dissertation enhances the understanding and safety of LLMs by proposing a framework to assess human-like stereotypes, methods to evaluate causal sensitivity based on socio-demographic attributes, and an analysis of political stereotypes through cognitive bias frameworks. By integrating insights from psychology and social sciences with computational methods this research makes a contribution to the ongoing discourse on ethical AI deployment, highlighting the necessity of understanding and addressing biases in language models to promote fairness and reduce discriminatory outcomes."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Examining large language models for safety and robustness through the lens of social science"]}]}],"canonical_facts":{"dc:contributor":["Diesner, Jana","Kilicoglu, Halil","Bosh, Nigel","Wang, Haohan"],"dc:creator":["Jeoung, Sullam"],"dc:date":["2025-03-07","2025-05"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","The student, Sullam Jeoung, accepted the attached license on 2025-03-05 at 10:51.","The student, Sullam Jeoung, submitted this Dissertation for approval on 2025-03-05 at 10:55.","This Dissertation was approved for publication on 2025-03-07 at 10:36.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21662 on 2025-10-19 at 18:09:09","Large language models have demonstrated remarkable capabilities, often achieving human-like performance levels and significantly impacting our daily lives. However, these models can perpetuate and amplify harmful stereotypes and biases associated with socio-demographic representations, potentially generating discriminatory content that adversely affects individuals and communities. Given these risks and their broader societal implications, ensuring the safety and robustness of these models through the identification and mitigation of harmful stereotypes has become imperative. This dissertation presents comprehensive methodologies to address these challenges by integrating insights from social science, psychology, and cognitive studies with methods from natural language processing. First, we present a framework to assess human-like stereotypical patterns in large language models (LLMs), drawing upon established psychological theories of how individuals develop stereotypes toward various social groups. This theoretically-grounded approach provides construct validity in defining and measuring stereotypes. The framework incorporates three key dimensions: warmth-competence analysis, keyword- reasoning patterns, and emotional-behavioral responses. Through clustering analysis, keyword extraction, and reasoning pattern evaluation of LLM responses, we examine how these models align with or deviate from documented human behavioral patterns. Our findings reveal that LLMs demonstrate nuanced perceptions of social groups, consistent with psychological research highlighting the multifaceted nature of stereotypes. Notably, the models’ reasoning patterns, particularly regarding groups’ economic status, demonstrate a nuanced awareness of societal disparities. Second, we propose methods to examine causal sensitivity of language models on socio-demographic attributes. This is based on a controlled experimental framework that uses name frequency analysis from U.S. Census data and systematic evaluation of model predictions through causal graphs. Our findings show that less frequent first names lead to divergent model predictions, highlighting the need for careful demographic consideration in dataset design to ensure fair and consistent model performance across different name representations. Third, we investigate how LLMs exhibit and inflate political stereotypes through the lens of cognitive biases and representative heuristics. We analyze LLMs’ responses using two key theoretical frameworks: ’kernel of truth’ (whether stereotypes reflect empirical realities) and ’representative heuristics’ (whether models overemphasize representative attributes of target groups), comparing model outputs with actual human responses across various political topics. Our findings show that while LLMs can accurately mimic certain political positions, they tend to exaggerate these positions compared to empirical human responses, suggesting a vulnerability to stereotypical thinking similar to human cognitive biases. This implies the need for careful consideration of cognitive bias frameworks in developing and deploying language models, particularly in politically sensitive contexts, and demonstrates the potential effectiveness of prompt-based mitigation strategies in reducing stereotypical responses. Overall, this dissertation enhances the understanding and safety of LLMs by proposing a framework to assess human-like stereotypes, methods to evaluate causal sensitivity based on socio-demographic attributes, and an analysis of political stereotypes through cognitive bias frameworks. By integrating insights from psychology and social sciences with computational methods this research makes a contribution to the ongoing discourse on ethical AI deployment, highlighting the necessity of understanding and addressing biases in language models to promote fairness and reduce discriminatory outcomes."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/129175"],"dc:language":["en","eng"],"dc:rights":["Copyright 2025 Sullam Jeoung"],"dc:subject":["Large Language Model","Responsible AI"],"dc:title":["Examining large language models for safety and robustness through the lens of social science"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Information Sciences"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:04Z"}