Back to results

University of Illinois Urbana-Champaign

Hands in action: from 4D reconstruction to animation and robotics

Abstract

dc:description

Hands are central to human interactions in the real world, enabling us to perform several tasks with different objects on a daily basis. Imagine if we can develop AI agents, either digital (in the form of animation sequences) or physical (such as real-world robots), that allow us to interact with the environment in the same way as our hands. Digital agents can enable coaching applications using mixed reality devices by overlaying animations with the objects in the 3D space. For example, they can teach us how to move our hands & fingers, to switch between notes while playing a piano or to screw a bolt when repairing a machine. We can even have household robots to perform everyday tasks like loading the dishwasher, doing laundry, folding clothes, organizing shelves or cooking. To develop such AI agents, it is crucial to study different aspects of hand-object interactions, including the 3D structure of hands, objects, contact regions, and motion sequences. A central challenge in training these AI agents is acquiring 3D data to train machine learning models for 3D prediction. Unlike 2D labels that can be obtained through human labelers, collecting 3D annotations requires extensive lab setups. These setups often constrain the interactions depending on the capture settings, thereby hindering the performance of the models (trained on lab data) when transferred to everyday scenarios. Fortunately, there is an abundance of egocentric videos showing hand-object interactions in the wild with a large diversity of objects and actions. However, unlike 2D annotations, it is not possible to obtain 3D labels for these videos through human labelers. This limits the utility of these videos for training learning-based methods for 3D prediction. In this work, we propose methods for 4D reconstruction of hands, objects and contacts from large-scale everyday videos, extending beyond controlled lab settings. This involves two main insights: (1) 2D information, such as segmentation masks and keypoints, can be estimated effectively from videos using large vision models, (2) 3D shape priors can be extracted from lab & synthetic sources and combined with 2D masks & keypoints for 3D reconstruction. Using the data generated from these reconstruction models, we develop techniques for 4D motion forecasting (for generating animations) and robot learning from human videos. Overall, this dissertation lays the groundwork for how to scale up learning-based methods for 4D reconstruction of hand-object interactions from everyday videos for applications in animation and robotics. We hope this serves as a stepping-stone towards developing both digital and physical AI agents to perform tasks similar to human hands.

Degree

thesis:*
Name thesis:degree_name
Ph.D.
Level thesis:degree_level
Dissertation
Discipline thesis:degree_discipline
Computer Science
Grantor
University of Illinois Urbana-Champaign
Year dc:date
2025

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Aditya Prakash, -
Contributors dc:contributor
  • Gupta, Saurabh
  • Forsyth, David
  • Lazebnik, Svetlana
  • Damen, Dima

Subjects

dc:subject × 2

Rights

dc:rights
Statement dc:rights
  • Copyright 2025 - Aditya Prakash
Language dc:language
en

Identifiers

dc:identifier.*
Handle dc:identifier
https://hdl.handle.net/2142/132522
OAI identifier oai:identifier
oai:www.ideals.illinois.edu:2142/132522

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Aditya Prakash, -. Hands in action: from 4D reconstruction to animation and robotics. Dissertation thesis, University of Illinois Urbana-Champaign, 2025. https://hdl.handle.net/2142/132522