{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/124514"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/124514","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Dynamic multimodal learning: Empowering ai to interpret the temporally dynamic world through vision, language, audio, and video","abstract":"Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2026-05-01","abstract_html":"Submission published under a 24 month embargo labeled &#x27;U of I Access&#x27;, the embargo will last until 2026-05-01","abstract_has_math":false,"creators":["Khosla, Savya"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Hoiem, Derek W"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-05","date_published":"2024-05","updated_at":"2026-07-22T22:25:02Z","subjects":["Multimodal Learning","Video Representation Learning","Large Multimodal Models"],"languages":["en","eng"],"rights":["Copyright 2024 Savya Khosla"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/124514","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Hoiem, Derek W"]},{"key":"dc:creator","label":"Author","values":["Khosla, Savya"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2024-05","2024-04-12"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Multimodal Learning","Video Representation Learning","Large Multimodal Models"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2024 Savya Khosla"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/124514"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2026-05-01","The student, Savya Khosla, accepted the attached license on 2024-04-10 at 13:45.","The student, Savya Khosla, submitted this Thesis for approval on 2024-04-10 at 13:54.","This Thesis was approved for publication on 2024-04-12 at 11:03.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20348 on 2024-09-16 at 00:43:20","As humans, we perceive and comprehend our surroundings through various sensory inputs. Multimodal learning aims to empower machines to do the same -- learn by leveraging multiple modalities, such as vision, language, and audio, to develop a more holistic understanding of the world. This learning approach not only enhances the capabilities of Artificial Intelligence (AI) systems but also enables them to better navigate and interact with real-world scenarios. To further augment AI systems' understanding of the real world, it is essential to equip them with the ability to comprehend its dynamic nature. So, with the goal of enhancing AI systems' interpretation of multiple modalities and understanding of the temporally dynamic world, this work focuses on two key areas of investigation: The first area of investigation focuses on building general-purpose systems that are capable of performing tasks requiring several different modalities. To this end, we propose the first autoregressive multimodal model that is capable of parsing images, texts, audio, and videos as input, and generating images, texts, and audio as output. In particular, we discuss techniques to represent the multiple modalities into a shared semantic space, process them with a single encoder-decoder transformer model, stabilize model training, and evaluate its performance on a broad array of over 120 multimodal tasks. The second area of investigation explores multimodal learning within the dynamic context of the world, as represented in videos. Our focus lies on training a memory-augmented video encoder by jointly supervising various modalities present in video data. We showcase the proposed encoder's proficiency in modeling long-form videos while capturing both nuanced and overarching details of the video content. Additionally, we demonstrate the generalizability of the learned representations by adapting them to a challenging downstream task without any task-specific bells and whistles."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Dynamic multimodal learning: Empowering ai to interpret the temporally dynamic world through vision, language, audio, and video"]}]}],"canonical_facts":{"dc:contributor":["Hoiem, Derek W"],"dc:creator":["Khosla, Savya"],"dc:date":["2024-05","2024-04-12"],"dc:description":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2026-05-01","The student, Savya Khosla, accepted the attached license on 2024-04-10 at 13:45.","The student, Savya Khosla, submitted this Thesis for approval on 2024-04-10 at 13:54.","This Thesis was approved for publication on 2024-04-12 at 11:03.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20348 on 2024-09-16 at 00:43:20","As humans, we perceive and comprehend our surroundings through various sensory inputs. Multimodal learning aims to empower machines to do the same -- learn by leveraging multiple modalities, such as vision, language, and audio, to develop a more holistic understanding of the world. This learning approach not only enhances the capabilities of Artificial Intelligence (AI) systems but also enables them to better navigate and interact with real-world scenarios. To further augment AI systems' understanding of the real world, it is essential to equip them with the ability to comprehend its dynamic nature. So, with the goal of enhancing AI systems' interpretation of multiple modalities and understanding of the temporally dynamic world, this work focuses on two key areas of investigation: The first area of investigation focuses on building general-purpose systems that are capable of performing tasks requiring several different modalities. To this end, we propose the first autoregressive multimodal model that is capable of parsing images, texts, audio, and videos as input, and generating images, texts, and audio as output. In particular, we discuss techniques to represent the multiple modalities into a shared semantic space, process them with a single encoder-decoder transformer model, stabilize model training, and evaluate its performance on a broad array of over 120 multimodal tasks. The second area of investigation explores multimodal learning within the dynamic context of the world, as represented in videos. Our focus lies on training a memory-augmented video encoder by jointly supervising various modalities present in video data. We showcase the proposed encoder's proficiency in modeling long-form videos while capturing both nuanced and overarching details of the video content. Additionally, we demonstrate the generalizability of the learned representations by adapting them to a challenging downstream task without any task-specific bells and whistles."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/124514"],"dc:language":["en","eng"],"dc:rights":["Copyright 2024 Savya Khosla"],"dc:subject":["Multimodal Learning","Video Representation Learning","Large Multimodal Models"],"dc:title":["Dynamic multimodal learning: Empowering ai to interpret the temporally dynamic world through vision, language, audio, and video"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:02Z"}