Back to results

University of Pennsylvania

SAFEGUARDING AI SYSTEMS AGAINST UNEXPECTED INPUTS

Abstract

dc:description.abstract

Artificial intelligence systems powered by deep neural networks have achieved remarkable success across a broad range of applications. However, perturbations such as natural image corruptions or crafted malicious queries, can cause significant performance degradation. This poses severe risks in safety-critical applications, such as autonomous driving and clinical decision-making. A key vulnerability of machine learning models is their inability to handle data outside the training distribution or knowledge. When facing unseen or otherwise challenging inputs, models often make incorrect decisions without warning users. This thesis improves the safety of machine learning systems by building three stages for handling challenging/unexpected inputs: (1) rejecting unexpected inputs with an explanation, (2) providing statistical guarantees on rejection, and (3) enabling models to adapt to challenging inputs. We consider two distinct scenarios: models with known training distributions (e.g., in cyber-physical systems) where challenges are out-of-distribution data, and models with unknown training distributions (e.g., large language models in a multilingual context) where challenges are defined by standards like harmful content across languages. For cyber-physical systems, we develop memory based prototypes that characterize the training distribution for out of distribution detection and provide statistical guarantees for a window based detector. We then leverage these prototypes to adapt the model to inputs from new distributions. For multilingual large language models, we design a reasoning enabled guardrail to shield against unsafe multilingual prompts. Finally, we study challenging inputs arising from natural distribution shift in a clinical application: acne lesion classification.

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Yang, Yahan
Advisor dc:contributor.advisor
  • Lee, Insup

Subjects

dc:subject × 2

Rights

Language dc:language.iso
en

Identifiers

dc:identifier.*
Repository record dc:identifier.uri
https://repository.upenn.edu/handle/20.500.14332/62349
OAI identifier oai:identifier
oai:repository.upenn.edu:20.500.14332/62349

Chain of custody

source
Harvested from
University of Pennsylvania
Base URL
repository.upenn.edu/server/oai/request
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Yang, Yahan. SAFEGUARDING AI SYSTEMS AGAINST UNEXPECTED INPUTS. 2025. https://repository.upenn.edu/handle/20.500.14332/62349