{"id":{"repo_id":"gatech","oai_identifier":"oai:repository.gatech.edu:1853/81784"},"canonical_url":"https://search.dev.ndltd.org/etd/gatech/oai:repository.gatech.edu:1853/81784","repository":{"repo_id":"gatech","name":"Georgia Tech","base_url":"https://repository.gatech.edu/server/oai/request"},"display":{"title":"Multimodal Human Behavior Modeling: From Understanding to Generation","abstract":"Humans are learning and changing the world through a variety of behaviors in our daily activities. Thus, human behavior modeling is a critical step to develop AI agents that are able to assist us in various tasks. In contrast with learning objects, scenes and textures, human behaviors are inherently purposeful, guided by underlying intentions and goals. Additionally, human behaviors involve precise and adaptive interactions with the environment, characterized by fine-grained and nuanced control. The two key differences require innovative approaches for AI models to understand our intentions in the behaviors and capture the nuance of our actions in different tasks. In my dissertation, I elaborate my research on leveraging multimodal inputs to capture the underlying intentions and enable precise controllability on human actions in both understanding and generation problems. First of all, I develop the first audio-visual egocentric gaze anticipation model that forecasts gaze behaviors by fusing audio-visual streams in temporal and spatial dimensions separately. Second, I collect a multimodal social interaction dataset with detailed annotations, and analyze the contribution of visual signals to social scenario understanding. Third, I introduce a novel egocentric action frame generation task for efficient skill learning, and an innovative method to enhance action generation performance by bridging the gap of large language models and diffusion models in the feature space. Finally, I propose a unified text-image-to-video (TI2V) generation problem that includes all existing TI2V settings, and introduce a novel training-free method to condition pre-trained text-to-video foundation models on any number of given images. In conclusion, the ultimate goal of my research is to enable AI models to better understand and interact with people, paving the way towards human-centric artificial intelligence.","abstract_html":"Humans are learning and changing the world through a variety of behaviors in our daily activities. Thus, human behavior modeling is a critical step to develop AI agents that are able to assist us in various tasks. In contrast with learning objects, scenes and textures, human behaviors are inherently purposeful, guided by underlying intentions and goals. Additionally, human behaviors involve precise and adaptive interactions with the environment, characterized by fine-grained and nuanced control. The two key differences require innovative approaches for AI models to understand our intentions in the behaviors and capture the nuance of our actions in different tasks. In my dissertation, I elaborate my research on leveraging multimodal inputs to capture the underlying intentions and enable precise controllability on human actions in both understanding and generation problems. First of all, I develop the first audio-visual egocentric gaze anticipation model that forecasts gaze behaviors by fusing audio-visual streams in temporal and spatial dimensions separately. Second, I collect a multimodal social interaction dataset with detailed annotations, and analyze the contribution of visual signals to social scenario understanding. Third, I introduce a novel egocentric action frame generation task for efficient skill learning, and an innovative method to enhance action generation performance by bridging the gap of large language models and diffusion models in the feature space. Finally, I propose a unified text-image-to-video (TI2V) generation problem that includes all existing TI2V settings, and introduce a novel training-free method to condition pre-trained text-to-video foundation models on any number of given images. In conclusion, the ultimate goal of my research is to enable AI models to better understand and interact with people, paving the way towards human-centric artificial intelligence.","abstract_has_math":false,"creators":["Lai, Bolin"],"institution":"Georgia Institute of Technology","degree_name":"Machine Learning, PhD","degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Kira, Zsolt"],"committee_chairs":[],"committee_members":["Rehg, James","Hays, James","Hoffman, Judy","Shi, Humphrey"],"year":2026,"date_issued":"2026-05","date_published":"2026-05","updated_at":"2026-07-27T19:50:24Z","subjects":[],"languages":["English"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/1853/81784","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Kira, Zsolt"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Rehg, James","Hays, James","Hoffman, Judy","Shi, Humphrey"]},{"key":"dc:creator","label":"Author","values":["Lai, Bolin"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2026-05-29T14:23:06Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2026-05-29T14:23:06Z"]},{"key":"dc:date.issued","label":"Date","values":["2026-05"]},{"key":"dc:type","label":"Dc Type","values":["Text"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Machine Learning, PhD"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Georgia Institute of Technology"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["English"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/1853/81784"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Humans are learning and changing the world through a variety of behaviors in our daily activities. Thus, human behavior modeling is a critical step to develop AI agents that are able to assist us in various tasks. In contrast with learning objects, scenes and textures, human behaviors are inherently purposeful, guided by underlying intentions and goals. Additionally, human behaviors involve precise and adaptive interactions with the environment, characterized by fine-grained and nuanced control. The two key differences require innovative approaches for AI models to understand our intentions in the behaviors and capture the nuance of our actions in different tasks. In my dissertation, I elaborate my research on leveraging multimodal inputs to capture the underlying intentions and enable precise controllability on human actions in both understanding and generation problems. First of all, I develop the first audio-visual egocentric gaze anticipation model that forecasts gaze behaviors by fusing audio-visual streams in temporal and spatial dimensions separately. Second, I collect a multimodal social interaction dataset with detailed annotations, and analyze the contribution of visual signals to social scenario understanding. Third, I introduce a novel egocentric action frame generation task for efficient skill learning, and an innovative method to enhance action generation performance by bridging the gap of large language models and diffusion models in the feature space. Finally, I propose a unified text-image-to-video (TI2V) generation problem that includes all existing TI2V settings, and introduce a novel training-free method to condition pre-trained text-to-video foundation models on any number of given images. In conclusion, the ultimate goal of my research is to enable AI models to better understand and interact with people, paving the way towards human-centric artificial intelligence."]},{"key":"dc:format.mimetype","label":"Dc Format Mimetype","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Multimodal Human Behavior Modeling: From Understanding to Generation"]}]}],"canonical_facts":{"dc:contributor.advisor":["Kira, Zsolt"],"dc:contributor.committeemember":["Rehg, James","Hays, James","Hoffman, Judy","Shi, Humphrey"],"dc:creator":["Lai, Bolin"],"dc:date.accessioned":["2026-05-29T14:23:06Z"],"dc:date.available":["2026-05-29T14:23:06Z"],"dc:date.issued":["2026-05"],"dc:description.abstract":["Humans are learning and changing the world through a variety of behaviors in our daily activities. Thus, human behavior modeling is a critical step to develop AI agents that are able to assist us in various tasks. In contrast with learning objects, scenes and textures, human behaviors are inherently purposeful, guided by underlying intentions and goals. Additionally, human behaviors involve precise and adaptive interactions with the environment, characterized by fine-grained and nuanced control. The two key differences require innovative approaches for AI models to understand our intentions in the behaviors and capture the nuance of our actions in different tasks. In my dissertation, I elaborate my research on leveraging multimodal inputs to capture the underlying intentions and enable precise controllability on human actions in both understanding and generation problems. First of all, I develop the first audio-visual egocentric gaze anticipation model that forecasts gaze behaviors by fusing audio-visual streams in temporal and spatial dimensions separately. Second, I collect a multimodal social interaction dataset with detailed annotations, and analyze the contribution of visual signals to social scenario understanding. Third, I introduce a novel egocentric action frame generation task for efficient skill learning, and an innovative method to enhance action generation performance by bridging the gap of large language models and diffusion models in the feature space. Finally, I propose a unified text-image-to-video (TI2V) generation problem that includes all existing TI2V settings, and introduce a novel training-free method to condition pre-trained text-to-video foundation models on any number of given images. In conclusion, the ultimate goal of my research is to enable AI models to better understand and interact with people, paving the way towards human-centric artificial intelligence."],"dc:format.mimetype":["application/pdf"],"dc:identifier.uri":["https://hdl.handle.net/1853/81784"],"dc:language.iso":["English"],"dc:title":["Multimodal Human Behavior Modeling: From Understanding to Generation"],"dc:type":["Text"],"thesis:degree_name":["Machine Learning, PhD"],"thesis:institution_name":["Georgia Institute of Technology"]},"updated_at":"2026-07-27T19:50:24Z"}