{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/129510"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/129510","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Generative modeling of interactive and reactive digital humans","abstract":"Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2027-05-01","abstract_html":"Submission published under a 24 month embargo labeled &#x27;U of I Access&#x27;, the embargo will last until 2027-05-01","abstract_has_math":false,"creators":["Xu, Xiyan"],"institution":"University of Illinois Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Gui, Liangyan","Wang, Yuxiong"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-04-22","date_published":"2025-04-22","updated_at":"2026-07-22T22:25:05Z","subjects":["Human Motion Generation"],"languages":["en","eng"],"rights":["Copyright 2025 Xiyan Xu"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/129510","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Gui, Liangyan","Wang, Yuxiong"]},{"key":"dc:creator","label":"Author","values":["Xu, Xiyan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-04-22","2025-05"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Human Motion Generation"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Xiyan Xu"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/129510"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2027-05-01","The student, Xiyan Xu, accepted the attached license on 2025-04-22 at 01:38.","The student, Xiyan Xu, submitted this Thesis for approval on 2025-04-22 at 01:50.","This Thesis was approved for publication on 2025-04-22 at 15:26.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21703 on 2025-10-19 at 19:14:27","Modeling interactive and reactive human behaviors is essential for building intelligent virtual agents, enhancing immersive experiences in VR/AR, and advancing robot learning. Among the diverse forms of human activity, three key types of interactions—human-human, human-object, and hand-hand—play fundamental roles in shaping social communication, environmental engagement, and fine-grained physical manipulation. Accurately modeling these interactions is crucial for constructing holistic digital humans capable of situational awareness, social intelligence, and physical competence. Although recent advances in generative modeling, particularly diffusion-based approaches, have significantly improved the realism and diversity of synthesized motions, generating semantically meaningful, physically plausible, and generalizable interactive behaviors remains a formidable challenge across all three interaction types. In the domain of human-human interactions, existing methods often struggle to produce reactions that are simultaneously physically coherent and semantically aligned with contextual cues. To address these limitations, we propose MoReact, a two-stage diffusion framework addresses text-conditioned reaction generation task, guided by an interaction-aware loss designed to enhance both physical plausibility and semantic fidelity. For human-object interactions, prior works frequently rely on narrow assumptions, such as restricting interactions to hand contacts or focusing on specific object categories, which restricts generalization. We introduce AuxMoDiff, a unified and flexible diffusion model that incorporates auxiliary spatial cues to better capture dynamic human-object relationships, improving contact accuracy and adaptability to unseen objects and tasks. For hand-hand interactions, research progress has been hampered by the lack of highquality datasets with semantic annotations, making meaningful bimanual motion generation difficult. To fill this gap, we present TextHand, the first large-scale dataset of close two-hand interactions paired with rich natural language descriptions. Building upon this resource, we develop TextHHI, a text-driven diffusion model capable of synthesizing realistic, expressive, and semantically aligned bimanual interactions from textual prompts. Together, these contributions advance the generative modeling of interactive and reactive digital humans across multiple scenarios. By addressing the challenges in human-human, human-object, and hand-hand interaction modeling, this thesis takes an important step toward building intelligent virtual agents that can seamlessly coordinate social behaviors, object manipulations, and self-movements within complex, dynamic environments."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Generative modeling of interactive and reactive digital humans"]}]}],"canonical_facts":{"dc:contributor":["Gui, Liangyan","Wang, Yuxiong"],"dc:creator":["Xu, Xiyan"],"dc:date":["2025-04-22","2025-05"],"dc:description":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2027-05-01","The student, Xiyan Xu, accepted the attached license on 2025-04-22 at 01:38.","The student, Xiyan Xu, submitted this Thesis for approval on 2025-04-22 at 01:50.","This Thesis was approved for publication on 2025-04-22 at 15:26.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21703 on 2025-10-19 at 19:14:27","Modeling interactive and reactive human behaviors is essential for building intelligent virtual agents, enhancing immersive experiences in VR/AR, and advancing robot learning. Among the diverse forms of human activity, three key types of interactions—human-human, human-object, and hand-hand—play fundamental roles in shaping social communication, environmental engagement, and fine-grained physical manipulation. Accurately modeling these interactions is crucial for constructing holistic digital humans capable of situational awareness, social intelligence, and physical competence. Although recent advances in generative modeling, particularly diffusion-based approaches, have significantly improved the realism and diversity of synthesized motions, generating semantically meaningful, physically plausible, and generalizable interactive behaviors remains a formidable challenge across all three interaction types. In the domain of human-human interactions, existing methods often struggle to produce reactions that are simultaneously physically coherent and semantically aligned with contextual cues. To address these limitations, we propose MoReact, a two-stage diffusion framework addresses text-conditioned reaction generation task, guided by an interaction-aware loss designed to enhance both physical plausibility and semantic fidelity. For human-object interactions, prior works frequently rely on narrow assumptions, such as restricting interactions to hand contacts or focusing on specific object categories, which restricts generalization. We introduce AuxMoDiff, a unified and flexible diffusion model that incorporates auxiliary spatial cues to better capture dynamic human-object relationships, improving contact accuracy and adaptability to unseen objects and tasks. For hand-hand interactions, research progress has been hampered by the lack of highquality datasets with semantic annotations, making meaningful bimanual motion generation difficult. To fill this gap, we present TextHand, the first large-scale dataset of close two-hand interactions paired with rich natural language descriptions. Building upon this resource, we develop TextHHI, a text-driven diffusion model capable of synthesizing realistic, expressive, and semantically aligned bimanual interactions from textual prompts. Together, these contributions advance the generative modeling of interactive and reactive digital humans across multiple scenarios. By addressing the challenges in human-human, human-object, and hand-hand interaction modeling, this thesis takes an important step toward building intelligent virtual agents that can seamlessly coordinate social behaviors, object manipulations, and self-movements within complex, dynamic environments."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/129510"],"dc:language":["en","eng"],"dc:rights":["Copyright 2025 Xiyan Xu"],"dc:subject":["Human Motion Generation"],"dc:title":["Generative modeling of interactive and reactive digital humans"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:05Z"}