{"id":{"repo_id":"penn","oai_identifier":"oai:repository.upenn.edu:20.500.14332/62349"},"canonical_url":"https://search.dev.ndltd.org/etd/penn/oai:repository.upenn.edu:20.500.14332/62349","repository":{"repo_id":"penn","name":"University of Pennsylvania","base_url":"https://repository.upenn.edu/server/oai/request"},"display":{"title":"SAFEGUARDING AI SYSTEMS AGAINST UNEXPECTED INPUTS","abstract":"Artificial intelligence systems powered by deep neural networks have achieved remarkable success across a broad range of applications. However, perturbations such as natural image corruptions or crafted malicious queries, can cause significant performance degradation. This poses severe risks in safety-critical applications, such as autonomous driving and clinical decision-making. A key vulnerability of machine learning models is their inability to handle data outside the training distribution or knowledge. When facing unseen or otherwise challenging inputs, models often make incorrect decisions without warning users. This thesis improves the safety of machine learning systems by building three stages for handling challenging/unexpected inputs: (1) rejecting unexpected inputs with an explanation, (2) providing statistical guarantees on rejection, and (3) enabling models to adapt to challenging inputs. We consider two distinct scenarios: models with known training distributions (e.g., in cyber-physical systems) where challenges are out-of-distribution data, and models with unknown training distributions (e.g., large language models in a multilingual context) where challenges are defined by standards like harmful content across languages. For cyber-physical systems, we develop memory based prototypes that characterize the training distribution for out of distribution detection and provide statistical guarantees for a window based detector. We then leverage these prototypes to adapt the model to inputs from new distributions. For multilingual large language models, we design a reasoning enabled guardrail to shield against unsafe multilingual prompts. Finally, we study challenging inputs arising from natural distribution shift in a clinical application: acne lesion classification.","abstract_html":"Artificial intelligence systems powered by deep neural networks have achieved remarkable success across a broad range of applications. However, perturbations such as natural image corruptions or crafted malicious queries, can cause significant performance degradation. This poses severe risks in safety-critical applications, such as autonomous driving and clinical decision-making. A key vulnerability of machine learning models is their inability to handle data outside the training distribution or knowledge. When facing unseen or otherwise challenging inputs, models often make incorrect decisions without warning users. This thesis improves the safety of machine learning systems by building three stages for handling challenging/unexpected inputs: (1) rejecting unexpected inputs with an explanation, (2) providing statistical guarantees on rejection, and (3) enabling models to adapt to challenging inputs. We consider two distinct scenarios: models with known training distributions (e.g., in cyber-physical systems) where challenges are out-of-distribution data, and models with unknown training distributions (e.g., large language models in a multilingual context) where challenges are defined by standards like harmful content across languages. For cyber-physical systems, we develop memory based prototypes that characterize the training distribution for out of distribution detection and provide statistical guarantees for a window based detector. We then leverage these prototypes to adapt the model to inputs from new distributions. For multilingual large language models, we design a reasoning enabled guardrail to shield against unsafe multilingual prompts. Finally, we study challenging inputs arising from natural distribution shift in a clinical application: acne lesion classification.","abstract_has_math":false,"creators":["Yang, Yahan"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Lee, Insup"],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025","date_published":"2025","updated_at":"2026-07-24T03:47:40Z","subjects":["Computer Sciences","Data Science"],"languages":["en"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://repository.upenn.edu/handle/20.500.14332/62349","outbound_label":"Repository record","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Lee, Insup"]},{"key":"dc:creator","label":"Author","values":["Yang, Yahan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2026-01-29T17:22:03Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2026-01-29T17:22:03Z"]},{"key":"dc:date.issued","label":"Date","values":["2025"]},{"key":"dc:type","label":"Dc Type","values":["Dissertation/Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer Sciences","Data Science"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://repository.upenn.edu/handle/20.500.14332/62349"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["2025"]},{"key":"dc:description.abstract","label":"Abstract","values":["Artificial intelligence systems powered by deep neural networks have achieved remarkable success across a broad range of applications. However, perturbations such as natural image corruptions or crafted malicious queries, can cause significant performance degradation. This poses severe risks in safety-critical applications, such as autonomous driving and clinical decision-making. A key vulnerability of machine learning models is their inability to handle data outside the training distribution or knowledge. When facing unseen or otherwise challenging inputs, models often make incorrect decisions without warning users. This thesis improves the safety of machine learning systems by building three stages for handling challenging/unexpected inputs: (1) rejecting unexpected inputs with an explanation, (2) providing statistical guarantees on rejection, and (3) enabling models to adapt to challenging inputs. We consider two distinct scenarios: models with known training distributions (e.g., in cyber-physical systems) where challenges are out-of-distribution data, and models with unknown training distributions (e.g., large language models in a multilingual context) where challenges are defined by standards like harmful content across languages. For cyber-physical systems, we develop memory based prototypes that characterize the training distribution for out of distribution detection and provide statistical guarantees for a window based detector. We then leverage these prototypes to adapt the model to inputs from new distributions. For multilingual large language models, we design a reasoning enabled guardrail to shield against unsafe multilingual prompts. Finally, we study challenging inputs arising from natural distribution shift in a clinical application: acne lesion classification."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["Doctor of Philosophy (PhD)"]},{"key":"dc:title","label":"Title","values":["SAFEGUARDING AI SYSTEMS AGAINST UNEXPECTED INPUTS"]}]}],"canonical_facts":{"dc:contributor.advisor":["Lee, Insup"],"dc:creator":["Yang, Yahan"],"dc:date.accessioned":["2026-01-29T17:22:03Z"],"dc:date.available":["2026-01-29T17:22:03Z"],"dc:date.issued":["2025"],"dc:description":["2025"],"dc:description.abstract":["Artificial intelligence systems powered by deep neural networks have achieved remarkable success across a broad range of applications. However, perturbations such as natural image corruptions or crafted malicious queries, can cause significant performance degradation. This poses severe risks in safety-critical applications, such as autonomous driving and clinical decision-making. A key vulnerability of machine learning models is their inability to handle data outside the training distribution or knowledge. When facing unseen or otherwise challenging inputs, models often make incorrect decisions without warning users. This thesis improves the safety of machine learning systems by building three stages for handling challenging/unexpected inputs: (1) rejecting unexpected inputs with an explanation, (2) providing statistical guarantees on rejection, and (3) enabling models to adapt to challenging inputs. We consider two distinct scenarios: models with known training distributions (e.g., in cyber-physical systems) where challenges are out-of-distribution data, and models with unknown training distributions (e.g., large language models in a multilingual context) where challenges are defined by standards like harmful content across languages. For cyber-physical systems, we develop memory based prototypes that characterize the training distribution for out of distribution detection and provide statistical guarantees for a window based detector. We then leverage these prototypes to adapt the model to inputs from new distributions. For multilingual large language models, we design a reasoning enabled guardrail to shield against unsafe multilingual prompts. Finally, we study challenging inputs arising from natural distribution shift in a clinical application: acne lesion classification."],"dc:description.degree":["Doctor of Philosophy (PhD)"],"dc:identifier.uri":["https://repository.upenn.edu/handle/20.500.14332/62349"],"dc:language.iso":["en"],"dc:subject":["Computer Sciences","Data Science"],"dc:title":["SAFEGUARDING AI SYSTEMS AGAINST UNEXPECTED INPUTS"],"dc:type":["Dissertation/Thesis"]},"updated_at":"2026-07-24T03:47:40Z"}