Back to results

Georgia Institute of Technology

Multimodal Human Behavior Modeling: From Understanding to Generation

Abstract

dc:description.abstract

Humans are learning and changing the world through a variety of behaviors in our daily activities. Thus, human behavior modeling is a critical step to develop AI agents that are able to assist us in various tasks. In contrast with learning objects, scenes and textures, human behaviors are inherently purposeful, guided by underlying intentions and goals. Additionally, human behaviors involve precise and adaptive interactions with the environment, characterized by fine-grained and nuanced control. The two key differences require innovative approaches for AI models to understand our intentions in the behaviors and capture the nuance of our actions in different tasks. In my dissertation, I elaborate my research on leveraging multimodal inputs to capture the underlying intentions and enable precise controllability on human actions in both understanding and generation problems. First of all, I develop the first audio-visual egocentric gaze anticipation model that forecasts gaze behaviors by fusing audio-visual streams in temporal and spatial dimensions separately. Second, I collect a multimodal social interaction dataset with detailed annotations, and analyze the contribution of visual signals to social scenario understanding. Third, I introduce a novel egocentric action frame generation task for efficient skill learning, and an innovative method to enhance action generation performance by bridging the gap of large language models and diffusion models in the feature space. Finally, I propose a unified text-image-to-video (TI2V) generation problem that includes all existing TI2V settings, and introduce a novel training-free method to condition pre-trained text-to-video foundation models on any number of given images. In conclusion, the ultimate goal of my research is to enable AI models to better understand and interact with people, paving the way towards human-centric artificial intelligence.

Degree

thesis:*
Name thesis:degree_name
Machine Learning, PhD
Grantor
Georgia Institute of Technology
Year dc:date.issued
2026

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Lai, Bolin
Advisor dc:contributor.advisor
  • Kira, Zsolt
Committee members dc:contributor.committeemember
  • Rehg, James
  • Hays, James
  • Hoffman, Judy
  • Shi, Humphrey

Rights

Language dc:language.iso
English

Identifiers

dc:identifier.*
Handle dc:identifier.uri
https://hdl.handle.net/1853/81784
OAI identifier oai:identifier
oai:repository.gatech.edu:1853/81784

Chain of custody

source
Harvested from
Georgia Tech
Base URL
repository.gatech.edu/server/oai/request
Last updated
2026-07-27
Source record
OAI-PMH GetRecord
related terms
citation

Lai, Bolin. Multimodal Human Behavior Modeling: From Understanding to Generation. Georgia Institute of Technology, 2026. https://hdl.handle.net/1853/81784