{"id":{"repo_id":"nus","oai_identifier":"oai:scholarbank.nus.edu.sg:10635/314766"},"canonical_url":"https://search.dev.ndltd.org/etd/nus/oai:scholarbank.nus.edu.sg:10635/314766","repository":{"repo_id":"nus","name":"National University of Singapore","base_url":"https://scholarbank.nus.edu.sg/oai/request"},"display":{"title":"EFFORTS TOWARD TRUSTWORTHY MACHINE LEARNING: MITIGATING OVERCONFIDENCE, HALLUCINATION, AND MODALITY BIAS","abstract":"Trustworthy machine learning is critical for safe deployment of AI systems in high-stakes domains. Despite strong performance, models remain prone to reliability issues such as overconfidence, hallucinations, and modality bias. This thesis addresses these challenges through post-hoc methods and targeted fine-tuning strategies. First, we mitigate overconfidence in discriminative models by introducing a neighborhood agreement metric that leverages foundation model representations, offering more reliable confidence estimates than softmax scores. We further explore self-supervised probing tasks to capture semantic signals of correctness, improving calibration and failure detection on in- and out-of-distribution data. Second, we tackle hallucinations in large vision-language models (LVLMs) by proposing CLIP-guided decoding, which evaluates candidates during beam search to produce more faithful and visually grounded outputs. Finally, we study modality bias, revealing “blind faith in text” across LVLMs and showing that supervised fine-tuning with text augmentation reduces this imbalance. Together, these studies advance trustworthy AI.","abstract_html":"Trustworthy machine learning is critical for safe deployment of AI systems in high-stakes domains. Despite strong performance, models remain prone to reliability issues such as overconfidence, hallucinations, and modality bias. This thesis addresses these challenges through post-hoc methods and targeted fine-tuning strategies. First, we mitigate overconfidence in discriminative models by introducing a neighborhood agreement metric that leverages foundation model representations, offering more reliable confidence estimates than softmax scores. We further explore self-supervised probing tasks to capture semantic signals of correctness, improving calibration and failure detection on in- and out-of-distribution data. Second, we tackle hallucinations in large vision-language models (LVLMs) by proposing CLIP-guided decoding, which evaluates candidates during beam search to produce more faithful and visually grounded outputs. Finally, we study modality bias, revealing “blind faith in text” across LVLMs and showing that supervised fine-tuning with text augmentation reduces this imbalance. Together, these studies advance trustworthy AI.","abstract_has_math":false,"creators":["DENG AILIN"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-05-26","date_published":"2025-05-26","updated_at":"2026-07-24T03:30:34Z","subjects":["safety","reliable machine learning","trustworthy machine learning"],"languages":[],"rights":[],"rights_urls":["https://scholarbank.nus.edu.sg/bitstreams/96e66d35-85d2-41d1-97c4-6b868cc3e7de/download"],"identifier_entries":[]},"links":{"outbound_url":null,"outbound_label":null,"outbound_source":null},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["DENG AILIN"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2025-05-26"]},{"key":"dc:relation.isreferencedby","label":"Dc Relation Isreferencedby","values":["https://scholarbank.nus.edu.sg/handle/10635/314766"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["safety","reliable machine learning","trustworthy machine learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["https://scholarbank.nus.edu.sg/bitstreams/96e66d35-85d2-41d1-97c4-6b868cc3e7de/download"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://scholarbank.nus.edu.sg/bitstreams/afd22f91-b5b4-4eb3-8db8-cca544ffd959/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Trustworthy machine learning is critical for safe deployment of AI systems in high-stakes domains. Despite strong performance, models remain prone to reliability issues such as overconfidence, hallucinations, and modality bias. This thesis addresses these challenges through post-hoc methods and targeted fine-tuning strategies. First, we mitigate overconfidence in discriminative models by introducing a neighborhood agreement metric that leverages foundation model representations, offering more reliable confidence estimates than softmax scores. We further explore self-supervised probing tasks to capture semantic signals of correctness, improving calibration and failure detection on in- and out-of-distribution data. Second, we tackle hallucinations in large vision-language models (LVLMs) by proposing CLIP-guided decoding, which evaluates candidates during beam search to produce more faithful and visually grounded outputs. Finally, we study modality bias, revealing “blind faith in text” across LVLMs and showing that supervised fine-tuning with text augmentation reduces this imbalance. Together, these studies advance trustworthy AI."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["9f3c6f2aa8f96291ef77b0a18484521d","120e9bb9546b07a856cd8aaaa786223a","990fe93d016d525078026f5d20b6f8c4"]},{"key":"dc:title","label":"Title","values":["EFFORTS TOWARD TRUSTWORTHY MACHINE LEARNING: MITIGATING OVERCONFIDENCE, HALLUCINATION, AND MODALITY BIAS"]}]}],"canonical_facts":{"dc:creator":["DENG AILIN"],"dc:date.issued":["2025-05-26"],"dc:description.abstract":["Trustworthy machine learning is critical for safe deployment of AI systems in high-stakes domains. Despite strong performance, models remain prone to reliability issues such as overconfidence, hallucinations, and modality bias. This thesis addresses these challenges through post-hoc methods and targeted fine-tuning strategies. First, we mitigate overconfidence in discriminative models by introducing a neighborhood agreement metric that leverages foundation model representations, offering more reliable confidence estimates than softmax scores. We further explore self-supervised probing tasks to capture semantic signals of correctness, improving calibration and failure detection on in- and out-of-distribution data. Second, we tackle hallucinations in large vision-language models (LVLMs) by proposing CLIP-guided decoding, which evaluates candidates during beam search to produce more faithful and visually grounded outputs. Finally, we study modality bias, revealing “blind faith in text” across LVLMs and showing that supervised fine-tuning with text augmentation reduces this imbalance. Together, these studies advance trustworthy AI."],"dc:format.checksum.md5":["9f3c6f2aa8f96291ef77b0a18484521d","120e9bb9546b07a856cd8aaaa786223a","990fe93d016d525078026f5d20b6f8c4"],"dc:identifier.uri":["https://scholarbank.nus.edu.sg/bitstreams/afd22f91-b5b4-4eb3-8db8-cca544ffd959/download"],"dc:relation.isreferencedby":["https://scholarbank.nus.edu.sg/handle/10635/314766"],"dc:rights":["https://scholarbank.nus.edu.sg/bitstreams/96e66d35-85d2-41d1-97c4-6b868cc3e7de/download"],"dc:subject":["safety","reliable machine learning","trustworthy machine learning"],"dc:title":["EFFORTS TOWARD TRUSTWORTHY MACHINE LEARNING: MITIGATING OVERCONFIDENCE, HALLUCINATION, AND MODALITY BIAS"],"dc:type":["Thesis"]},"updated_at":"2026-07-24T03:30:34Z"}