University of Illinois - Chicago
Improved Out-of-Distribution Detection Using Segmented Images and Prompt-Only Text Reasoning
Abstract
dc:descriptionThe problem of Out-of-Distribution (OOD) detection has been thoroughly researched but continues to underperform in "near-OOD" settings, where OOD data may be very similar or inseparable from in-distribution (ID) data. The problem is that current state-of-the-art OOD detection methods fail to learn and utilize ID class-representative discriminative features effectively, while simultaneously demanding inaccessible amounts of resources. This thesis presents two OOD detection methods that exhibit state-of-the-art performance in near-OOD settings, one in the image data domain via the use of multi-view cross-attention processing of segmented images, and one in the text data domain using efficient and accessible prompt-only detection via alignment and scoring. The first approach leverages the segmented data and segmentation models to employ a multi-view method for image-based OOD detection, denoted as Cross-view Attention of Segmented views for OOD Detection (CASOD). Through the use of a pre-trained model and a novel cross-view correlation attention fusion architecture, discriminative features are learned across the original image and a foreground and background view, resulting in a highly informative ID class-relevant feature space. Utilizing distance-based OOD detection methods, CASOD achieves state-of-the-art performance over previous OOD detection baselines across a number of academically- or publicly-available datasets, including ImageNet, NINCO, SSB-Hard, iNaturalist, Textures, OpenImage-O, Places365, Species, and SUN. In particular, OOD detection performance on near-OOD datasets is shown to significantly improve. The second approach utilizes the inherent knowledge and reasoning capabilities in large language models (LLMs) to solve the task of prompt-only OOD detection, which requires only the use of text-prompting for OOD detection with no access to LLM components, logits, or outputs and no fine-tuning. Through prompting LLMs to perform lexical and semantic alignment before giving an OOD score for the input, OOD detection on academically- or publicly-available near-OOD datasets, including the Banking, CLINC, StackOverflow, 20NewsGroups, Dbpedia, and Snips datasets, can be significantly improved. Experiments demonstrate that this method, denoted as Alignment-based Thresholding for Prompt-only OOD detection (ATPO), not only outperforms the previous prompt-only OOD detection method but also outperforms strong traditional OOD detection methods with access to LLM features and logits.
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Alexander Politowicz (24400118)
Subjects
dc:subject × 1Rights
dc:rights- Statement dc:rights
-
- In Copyright
- Open Access after 2028-05-01
Identifiers
dc:identifier.*- DOI dc:identifier
- https://doi.org/10.25417/uic.32995172.v1
- OAI identifier oai:identifier
- oai:figshare.com:article/32995172