Back to results

National University of Singapore

EFFORTS TOWARD TRUSTWORTHY MACHINE LEARNING: MITIGATING OVERCONFIDENCE, HALLUCINATION, AND MODALITY BIAS

Abstract

dc:description.abstract

Trustworthy machine learning is critical for safe deployment of AI systems in high-stakes domains. Despite strong performance, models remain prone to reliability issues such as overconfidence, hallucinations, and modality bias. This thesis addresses these challenges through post-hoc methods and targeted fine-tuning strategies. First, we mitigate overconfidence in discriminative models by introducing a neighborhood agreement metric that leverages foundation model representations, offering more reliable confidence estimates than softmax scores. We further explore self-supervised probing tasks to capture semantic signals of correctness, improving calibration and failure detection on in- and out-of-distribution data. Second, we tackle hallucinations in large vision-language models (LVLMs) by proposing CLIP-guided decoding, which evaluates candidates during beam search to produce more faithful and visually grounded outputs. Finally, we study modality bias, revealing “blind faith in text” across LVLMs and showing that supervised fine-tuning with text augmentation reduces this imbalance. Together, these studies advance trustworthy AI.

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • DENG AILIN

Subjects

dc:subject × 3

Rights

dc:rights

Chain of custody

source
Harvested from
National University of Singapore
Base URL
scholarbank.nus.edu.sg/oai/request
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

DENG AILIN. EFFORTS TOWARD TRUSTWORTHY MACHINE LEARNING: MITIGATING OVERCONFIDENCE, HALLUCINATION, AND MODALITY BIAS. 2025.