{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/127404"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/127404","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Multi-modal learning for image and beyond","abstract":"Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2026-12-01","abstract_html":"Submission published under a 24 month embargo labeled &#x27;U of I Access&#x27;, the embargo will last until 2026-12-01","abstract_has_math":false,"creators":["Jiang, Qian"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Do, Minh N","Do, Minh N.","Chen, Deming","Schwing, Alexander","Zhao, Han","Yeh, Raymond A."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-12-06","date_published":"2024-12-06","updated_at":"2026-07-22T22:25:04Z","subjects":["Machine Learning","Multimodal"],"languages":["en","eng"],"rights":["Copyright 2024 Qian Jiang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/127404","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Do, Minh N","Do, Minh N.","Chen, Deming","Schwing, Alexander","Zhao, Han","Yeh, Raymond A."]},{"key":"dc:creator","label":"Author","values":["Jiang, Qian"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2024-12-06","2024-12"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Machine Learning","Multimodal"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2024 Qian Jiang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/127404"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2026-12-01","The student, Qian Jiang, accepted the attached license on 2024-12-06 at 03:06.","The student, Qian Jiang, submitted this Dissertation for approval on 2024-12-06 at 03:16.","This Dissertation was approved for publication on 2024-12-06 at 12:16.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21523 on 2025-03-28 at 14:44:46","The rapid advancement of technology has propelled the field of machine learning into new frontiers, with a particular emphasis on multi-modal learning approaches. This thesis explores the integration of multiple modalities with visual data, investigating three distinct yet complementary directions: text integration for enhanced understanding, hardware optimization for efficient deployment, and medical image representation learning for robust diagnostics. The study begins by delving into the integration of image and text modality. We first prove that exact modality alignment is sub-optimal in general for downstream prediction tasks. Thus we propose three general approaches to construct latent modality structures. We test our model on a variety of tasks including zero/few-shot image classification, image-text retrieval, visual question answering, visual reasoning, and visual entailment. Our method achieves consistent improvements over existing methods demonstrating the effectiveness and generalizability. We then extend the focus to image and hardware modality. We propose End-to-end Hardware-aware Differentiable Neural Architecture Search (EH-DNAS), a seamless integration of end-to-end hardware benchmarking, and fully automated DNAS to deliver hardware-efficient deep neural networks on various platforms such as mobile devices and dedicated AI accelerators. Experiments on CIFAR10 and ImageNet show that EH-DNAS improves the hardware performance by an average of $1.5\\times$ on customized accelerators and existing hardware processors than the state-of-the-art hardware-efficient networks while maintaining the classification accuracy. Finally, we address the challenges of multimodal learning in medical image representation learning, focusing on the B-mode and M-mode Optical Coherence Tomography (OCT) images. We proposed a novel triplet-based learning framework specifically designed for medical image applications with label noise. Our approach demonstrates superior performance on OCT disease classification, achieving up to 98.44\\% accuracy on M-mode and 94.12\\% on B-mode, achieving effective multimodal learning in scenarios with imperfect alignment. Collectively, this work advances multimodal learning across three dimensions: theoretical insights about modality alignment and methods for construct modality structures for image-text tasks, practical solutions for hardware-aware deployment, and robust frameworks for medical applications. Our comprehensive study, spanning from theoretical foundations to real-world implementations, provides both fundamental understanding and practical solutions for the evolving challenges in multimodal learning."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Multi-modal learning for image and beyond"]}]}],"canonical_facts":{"dc:contributor":["Do, Minh N","Do, Minh N.","Chen, Deming","Schwing, Alexander","Zhao, Han","Yeh, Raymond A."],"dc:creator":["Jiang, Qian"],"dc:date":["2024-12-06","2024-12"],"dc:description":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2026-12-01","The student, Qian Jiang, accepted the attached license on 2024-12-06 at 03:06.","The student, Qian Jiang, submitted this Dissertation for approval on 2024-12-06 at 03:16.","This Dissertation was approved for publication on 2024-12-06 at 12:16.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21523 on 2025-03-28 at 14:44:46","The rapid advancement of technology has propelled the field of machine learning into new frontiers, with a particular emphasis on multi-modal learning approaches. This thesis explores the integration of multiple modalities with visual data, investigating three distinct yet complementary directions: text integration for enhanced understanding, hardware optimization for efficient deployment, and medical image representation learning for robust diagnostics. The study begins by delving into the integration of image and text modality. We first prove that exact modality alignment is sub-optimal in general for downstream prediction tasks. Thus we propose three general approaches to construct latent modality structures. We test our model on a variety of tasks including zero/few-shot image classification, image-text retrieval, visual question answering, visual reasoning, and visual entailment. Our method achieves consistent improvements over existing methods demonstrating the effectiveness and generalizability. We then extend the focus to image and hardware modality. We propose End-to-end Hardware-aware Differentiable Neural Architecture Search (EH-DNAS), a seamless integration of end-to-end hardware benchmarking, and fully automated DNAS to deliver hardware-efficient deep neural networks on various platforms such as mobile devices and dedicated AI accelerators. Experiments on CIFAR10 and ImageNet show that EH-DNAS improves the hardware performance by an average of $1.5\\times$ on customized accelerators and existing hardware processors than the state-of-the-art hardware-efficient networks while maintaining the classification accuracy. Finally, we address the challenges of multimodal learning in medical image representation learning, focusing on the B-mode and M-mode Optical Coherence Tomography (OCT) images. We proposed a novel triplet-based learning framework specifically designed for medical image applications with label noise. Our approach demonstrates superior performance on OCT disease classification, achieving up to 98.44\\% accuracy on M-mode and 94.12\\% on B-mode, achieving effective multimodal learning in scenarios with imperfect alignment. Collectively, this work advances multimodal learning across three dimensions: theoretical insights about modality alignment and methods for construct modality structures for image-text tasks, practical solutions for hardware-aware deployment, and robust frameworks for medical applications. Our comprehensive study, spanning from theoretical foundations to real-world implementations, provides both fundamental understanding and practical solutions for the evolving challenges in multimodal learning."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/127404"],"dc:language":["en","eng"],"dc:rights":["Copyright 2024 Qian Jiang"],"dc:subject":["Machine Learning","Multimodal"],"dc:title":["Multi-modal learning for image and beyond"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:04Z"}