National University of Singapore
EFFORTS TOWARD TRUSTWORTHY MACHINE LEARNING: MITIGATING OVERCONFIDENCE, HALLUCINATION, AND MODALITY BIAS
Abstract
dc:description.abstractTrustworthy machine learning is critical for safe deployment of AI systems in high-stakes domains. Despite strong performance, models remain prone to reliability issues such as overconfidence, hallucinations, and modality bias. This thesis addresses these challenges through post-hoc methods and targeted fine-tuning strategies. First, we mitigate overconfidence in discriminative models by introducing a neighborhood agreement metric that leverages foundation model representations, offering more reliable confidence estimates than softmax scores. We further explore self-supervised probing tasks to capture semantic signals of correctness, improving calibration and failure detection on in- and out-of-distribution data. Second, we tackle hallucinations in large vision-language models (LVLMs) by proposing CLIP-guided decoding, which evaluates candidates during beam search to produce more faithful and visually grounded outputs. Finally, we study modality bias, revealing “blind faith in text” across LVLMs and showing that supervised fine-tuning with text augmentation reduces this imbalance. Together, these studies advance trustworthy AI.
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- DENG AILIN