Back to search

University of Illinois at Urbana-Champaign

Dynamic multimodal learning: Empowering ai to interpret the temporally dynamic world through vision, language, audio, and video

Abstract

dc:description

As humans, we perceive and comprehend our surroundings through various sensory inputs. Multimodal learning aims to empower machines to do the same -- learn by leveraging multiple modalities, such as vision, language, and audio, to develop a more holistic understanding of the world. This learning approach not only enhances the capabilities of Artificial Intelligence (AI) systems but also enables them to better navigate and interact with real-world scenarios. To further augment AI systems' understanding of the real world, it is essential to equip them with the ability to comprehend its dynamic nature. So, with the goal of enhancing AI systems' interpretation of multiple modalities and understanding of the temporally dynamic world, this work focuses on two key areas of investigation: The first area of investigation focuses on building general-purpose systems that are capable of performing tasks requiring several different modalities. To this end, we propose the first autoregressive multimodal model that is capable of parsing images, texts, audio, and videos as input, and generating images, texts, and audio as output. In particular, we discuss techniques to represent the multiple modalities into a shared semantic space, process them with a single encoder-decoder transformer model, stabilize model training, and evaluate its performance on a broad array of over 120 multimodal tasks. The second area of investigation explores multimodal learning within the dynamic context of the world, as represented in videos. Our focus lies on training a memory-augmented video encoder by jointly supervising various modalities present in video data. We showcase the proposed encoder's proficiency in modeling long-form videos while capturing both nuanced and overarching details of the video content. Additionally, we demonstrate the generalizability of the learned representations by adapting them to a challenging downstream task without any task-specific bells and whistles.

Degree

thesis:*
Name thesis:degree_name
M.S.
Level thesis:degree_level
Thesis
Discipline thesis:degree_discipline
Computer Science
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2024

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Khosla, Savya
Contributors dc:contributor
  • Hoiem, Derek W

Subjects

dc:subject × 3

Rights

dc:rights
Statement dc:rights
  • Copyright 2024 Savya Khosla
Language dc:language
en, eng

Identifiers

dc:identifier.*
Handle dc:identifier
https://hdl.handle.net/2142/124514

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Khosla, Savya. Dynamic multimodal learning: Empowering ai to interpret the temporally dynamic world through vision, language, audio, and video. Thesis thesis, University of Illinois at Urbana-Champaign, 2024. https://hdl.handle.net/2142/124514