Back to results

University of Minnesota

Leveraging Geometric Priors for In-the-wild Pose Estimation and 3D Reconstruction

Abstract

dc:description.abstract

The ability to observe and understand the physical world is essential for robots to operate effectively in unstructured environments. Visual perception provides important cues for recognizing objects, reasoning about their spatial relationships, and inferring 3D geometry or motion. However, interpreting visual data is challenging due to occlusions, limited viewpoints, and the partial nature of sensor observations, which make many real-world perception problems highly ill-posed. While deep learning methods have been applied to extract meaningful 3D information by learning underlying priors, they often suffer from domain gaps and require large labeled datasets. In contrast, optimization-based approaches, which do not require labeled data, can adapt to individual instances but are sensitive to initialization. To address these challenges in solving real-world perception problems, we design algorithms that leverage geometric priors paired with observation-driven optimization. By incorporating learned or explicitly defined priors to guide optimization, our methods are less sensitive to domain gaps and more robust to unseen scenarios than existing methods. This dissertation is divided into two main parts. In the first part, we study two problems related to pose estimation for static scenes. In the second part, we extend the setting from static scene to dynamic scene analysis and modeling. The first problem we consider is to estimate the camera pose of an image that captures an object from a known category. Specifically, we are interested in solving the camera pose with respect to the object frame of a template shape. The template shape may be significantly different from the observed object, therefore building one-to-one correspondences in a single shot is error prone, even with the help of deep learning models. To handle the difference between the template and the observed object, we propose an optimization method that retains all possible correspondences and gradually update them to find the optimal matching. Experiments on both real and synthetic data show that our method outperforms previous methods even when the template shape is significantly different from the observed object. The second pose estimation problem focuses on registering a video of an indoor environment to the 2D LiDAR scan of the entire environment. The main challenge lies in the fact that the video captures only part of the environment, therefore, there might not be sufficient information in the video for registration. To solve this problem, we propose to use a pre-trained text-to-image inpainting model for filling in the missing information. Experiments show that the performance of pose registration is significantly improved by leveraging the learned prior of generative models. In the second half of this dissertation, we extend from static scene analysis to dynamic object modeling. We investigate the problem of part segmentation and motion estimation for point clouds of articulated objects. These are fundamental problems in articulated object modeling, where most existing methods focus on point clouds that completely cover the object surface. However, in real-world scenarios, observations at different time steps may only partially cover the object. Such partial observations can happen when the object undergoes occlusions, or when the object is captured by multiple sensors asynchronously. To handle these incomplete data, we propose to represent rigid parts and their poses with dynamic 3D Gaussians. Compared to methods that solely rely on tracking point correspondences across time, our method that models rigid parts as point distributions achieves a 13% improvement in segmentation accuracy. We further extend the articulated object modeling problem to full 4D reconstruction from sparse and partial observations. Our goal is to reconstruct the continuous motion, geometry, and appearance of a dynamic target from partial observations captured at sparse time steps. To address this highly ill-posed problem, we incorporate additional geometric priors in the form of a skeleton structure and an initial static reconstruction to guide the motion and deformation estimation. We propose a skeleton-driven deformation model that enables smooth motion interpolation while preserving fine-grained details when supervised only on sparse observations. Finally, we show that the need for an initial reconstruction can be relaxed by replacing it with a diffusion-based generative prior, improving applicability to real-world scenarios. In summary, this dissertation develops robust algorithms for real-world perception problems under partial and sparse observations. By combining geometric priors with observation-driven optimization, we address a range of challenging problems, including camera pose estimation, articulated object modeling, and dynamic 4D reconstruction. This research aims to contribute to the deployment of robots to complex real-world environments by showing the robustness and generalization in in-the-wild settings.

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Chao, Jun-Jee

Subjects

dc:subject × 2

Rights

Language dc:language.iso
en

Identifiers

dc:identifier.*
Handle dc:identifier.uri
https://hdl.handle.net/11299/280363
OAI identifier oai:identifier
oai:conservancy.umn.edu:11299/280363

Chain of custody

source
Harvested from
University of Minnesota
Base URL
conservancy.umn.edu/server/oai/request
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Chao, Jun-Jee. Leveraging Geometric Priors for In-the-wild Pose Estimation and 3D Reconstruction. 2026. https://hdl.handle.net/11299/280363