University of Illinois at Urbana-Champaign
Towards versatile 3D human motion prediction in the wild
Abstract
dc:descriptionBeing able to “look into the future” is a remarkable cognitive hallmark of humans. For example, humans can naturally anticipate how people move or act in the near future, based on their historical movements, even in a complex real-world scenario in the wild, which poses a critical challenge for machines to replicate. On the contrary, the state-of-the-art human motion forecasting method often focuses on simplified scenarios, e.g., predicting future motion of a single person in a deterministic way. This thesis endeavors to develop novel techniques that enable machines to anticipate human motion while considering real-world complexities. Our work is grounded in two fundamental insights: First, human motion prediction inherently involves uncertainty and multi-modality, especially in long-term forecasting. Second, such uncertainty does not suggest complete randomness in human movements; instead, they are highly dependent on the environment and its changes. To tackle these challenges, we integrate diverse generation and environment-aware prediction into various scenarios. We commence by investigating the prediction of diverse single-person motion. Our key insight is that future human motions are not completely random or independent, but rather exhibit deterministic properties consistent with physical laws and constraints. Based on this observation, we propose anchor-based representations that encode human motion in the latent space using deterministic and learnable components. These anchors have been trained to specialize and diversify for different modes of future motion, enabling us to generate more diverse and accurate predictions with only a few additional parameters. We then move on to introduce a task that considers the impact of social interactions on the diversity of future human poses, simultaneously considering the social aspects of multi-person interaction, the realism, and the diversity of human motion. The cumulative difficulties inherent in this task motivate us to adopt a divide-and-conquer strategy that we instantiate as a dual-level generative modeling framework. We demonstrate that various multi-person predictors and generative models can operationalize this general framework, leading to consistent improvements in both accuracy and diversity. Humans interact not only with other humans but also with the surrounding environment. With this in mind, we develop a novel task of anticipating 3D human-object interactions (HOIs), considering the dynamics of a general object. Our proposed framework injects prior to the interaction that it follows a simple pattern at reference contact points. We show that the framework can model objects with various shapes and ensure physically valid interactions.
Degree
thesis:*- Name thesis:degree_name
- M.S.
- Level thesis:degree_level
- Thesis
- Discipline thesis:degree_discipline
- Computer Science
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2023
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Xu, Sirui
- Contributors dc:contributor
-
- Wang, Yuxiong
- Gui, Liangyan
Subjects
dc:subject × 1Rights
dc:rights- Statement dc:rights
-
- Copyright 2023 Sirui Xu
- Language dc:language
- en, eng
Identifiers
dc:identifier.*- Handle dc:identifier
- https://hdl.handle.net/2142/120448