{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/120448"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/120448","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Towards versatile 3D human motion prediction in the wild","abstract":"Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2025-05-01","abstract_html":"Submission published under a 24 month embargo labeled &#x27;U of I Access&#x27;, the embargo will last until 2025-05-01","abstract_has_math":false,"creators":["Xu, Sirui"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Wang, Yuxiong","Gui, Liangyan"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2023,"date_issued":"2023-05","date_published":"2023-05","updated_at":"2026-07-22T22:24:57Z","subjects":["Human Motion Prediction"],"languages":["en","eng"],"rights":["Copyright 2023 Sirui Xu"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/120448","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Wang, Yuxiong","Gui, Liangyan"]},{"key":"dc:creator","label":"Author","values":["Xu, Sirui"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2023-05","2023-05-02"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Human Motion Prediction"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2023 Sirui Xu"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/120448"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2025-05-01","The student, Sirui Xu, accepted the attached license on 2023-05-01 at 13:18.","The student, Sirui Xu, submitted this Thesis for approval on 2023-05-01 at 13:47.","This Thesis was approved for publication on 2023-05-02 at 16:56.","DSpace SAF Submission Ingestion Package generated from Vireo submission #19279 on 2023-09-01 at 17:15:42","Being able to “look into the future” is a remarkable cognitive hallmark of humans. For example, humans can naturally anticipate how people move or act in the near future, based on their historical movements, even in a complex real-world scenario in the wild, which poses a critical challenge for machines to replicate. On the contrary, the state-of-the-art human motion forecasting method often focuses on simplified scenarios, e.g., predicting future motion of a single person in a deterministic way. This thesis endeavors to develop novel techniques that enable machines to anticipate human motion while considering real-world complexities. Our work is grounded in two fundamental insights: First, human motion prediction inherently involves uncertainty and multi-modality, especially in long-term forecasting. Second, such uncertainty does not suggest complete randomness in human movements; instead, they are highly dependent on the environment and its changes. To tackle these challenges, we integrate diverse generation and environment-aware prediction into various scenarios. We commence by investigating the prediction of diverse single-person motion. Our key insight is that future human motions are not completely random or independent, but rather exhibit deterministic properties consistent with physical laws and constraints. Based on this observation, we propose anchor-based representations that encode human motion in the latent space using deterministic and learnable components. These anchors have been trained to specialize and diversify for different modes of future motion, enabling us to generate more diverse and accurate predictions with only a few additional parameters. We then move on to introduce a task that considers the impact of social interactions on the diversity of future human poses, simultaneously considering the social aspects of multi-person interaction, the realism, and the diversity of human motion. The cumulative difficulties inherent in this task motivate us to adopt a divide-and-conquer strategy that we instantiate as a dual-level generative modeling framework. We demonstrate that various multi-person predictors and generative models can operationalize this general framework, leading to consistent improvements in both accuracy and diversity. Humans interact not only with other humans but also with the surrounding environment. With this in mind, we develop a novel task of anticipating 3D human-object interactions (HOIs), considering the dynamics of a general object. Our proposed framework injects prior to the interaction that it follows a simple pattern at reference contact points. We show that the framework can model objects with various shapes and ensure physically valid interactions."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Towards versatile 3D human motion prediction in the wild"]}]}],"canonical_facts":{"dc:contributor":["Wang, Yuxiong","Gui, Liangyan"],"dc:creator":["Xu, Sirui"],"dc:date":["2023-05","2023-05-02"],"dc:description":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2025-05-01","The student, Sirui Xu, accepted the attached license on 2023-05-01 at 13:18.","The student, Sirui Xu, submitted this Thesis for approval on 2023-05-01 at 13:47.","This Thesis was approved for publication on 2023-05-02 at 16:56.","DSpace SAF Submission Ingestion Package generated from Vireo submission #19279 on 2023-09-01 at 17:15:42","Being able to “look into the future” is a remarkable cognitive hallmark of humans. For example, humans can naturally anticipate how people move or act in the near future, based on their historical movements, even in a complex real-world scenario in the wild, which poses a critical challenge for machines to replicate. On the contrary, the state-of-the-art human motion forecasting method often focuses on simplified scenarios, e.g., predicting future motion of a single person in a deterministic way. This thesis endeavors to develop novel techniques that enable machines to anticipate human motion while considering real-world complexities. Our work is grounded in two fundamental insights: First, human motion prediction inherently involves uncertainty and multi-modality, especially in long-term forecasting. Second, such uncertainty does not suggest complete randomness in human movements; instead, they are highly dependent on the environment and its changes. To tackle these challenges, we integrate diverse generation and environment-aware prediction into various scenarios. We commence by investigating the prediction of diverse single-person motion. Our key insight is that future human motions are not completely random or independent, but rather exhibit deterministic properties consistent with physical laws and constraints. Based on this observation, we propose anchor-based representations that encode human motion in the latent space using deterministic and learnable components. These anchors have been trained to specialize and diversify for different modes of future motion, enabling us to generate more diverse and accurate predictions with only a few additional parameters. We then move on to introduce a task that considers the impact of social interactions on the diversity of future human poses, simultaneously considering the social aspects of multi-person interaction, the realism, and the diversity of human motion. The cumulative difficulties inherent in this task motivate us to adopt a divide-and-conquer strategy that we instantiate as a dual-level generative modeling framework. We demonstrate that various multi-person predictors and generative models can operationalize this general framework, leading to consistent improvements in both accuracy and diversity. Humans interact not only with other humans but also with the surrounding environment. With this in mind, we develop a novel task of anticipating 3D human-object interactions (HOIs), considering the dynamics of a general object. Our proposed framework injects prior to the interaction that it follows a simple pattern at reference contact points. We show that the framework can model objects with various shapes and ensure physically valid interactions."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/120448"],"dc:language":["en","eng"],"dc:rights":["Copyright 2023 Sirui Xu"],"dc:subject":["Human Motion Prediction"],"dc:title":["Towards versatile 3D human motion prediction in the wild"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:57Z"}