Back to results

University of Washington

Towards Multi-Person 3D Pose Estimation in Natural Videos

Abstract

dc:description.abstract

Despite the increasing need of analyzing human poses on the street and in the wild, multi-person 3D pose estimation using static or moving monocular camera in real-world scenarios remains a challenge, requiring large-scale training data or high computation complexity due to the high degrees of freedom in 3D human poses. To address these challenges, a novel scheme, Hierarchical 3D Human Pose Estimation (H3DHPE), is proposed to effectively track and hierarchically estimate 3D human poses in natural videos in an efficient fashion. Torso estimation is formulated as a Perspective-N-Point (PNP) problem, limb pose estimation is solved as an optimization problem, and the high dimensional pose estimation is hierarchically addressed efficiently. As an extension to Hierarchical 3D Human Pose Estimation (H3DHPE), Universal Hierarchical 3D Human Pose Estimation (UH3DHPE) is proposed to handle the case of an occluded or inaccurate 2D torso keypoints, which makes torso-first estimation in H3DHPE unreliable. An effective method to directly estimate limb poses without building upon the estimated torso pose is proposed, and the torso pose can then be further refined to form the hierarchy in a bottom-up fashion. An adaptive merging strategy is proposed to determine the best hierarchy. The advantages of the proposed unsupervised methods are validated on various datasets including a lot of natural real-world scenes. For better evaluation and future research, a unique dataset called Moving camera Multi-Human interactions (MMHuman) is collected, with accurate MoCap ground truth, for multi-person interaction scenarios recorded by a monocular moving camera. Superior performance is shown on the newly collected MMHuman compared to state-of-the-art methods, including supervised methods, proving that our unsupervised solution generalize better to natural videos. To further tackle the problem of long term occlusions, a deep neutral network (DNN) solution is explored for trajectory recovery. To our best knowledge, it’s the first to use temporal gated convolutions to recover missing poses and address the occlusion issues in the pose estimation. A simple yet effective approach is proposed to transform normalized poses to the global trajectory into the camera coordinate.

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • GU, Renshu
Advisor dc:contributor.advisor
  • Hwang, Jenq-Neng

Subjects

dc:subject × 6

Rights

dc:rights
Statement dc:rights
  • none
Language dc:language.iso
en_US

Identifiers

dc:identifier.*
Handle dc:identifier.uri
http://hdl.handle.net/1773/45769
OAI identifier oai:identifier
oai:digital.lib.washington.edu:1773/45769

Chain of custody

source
Harvested from
University of Washington
Base URL
digital.lib.washington.edu/server/oai/request
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

GU, Renshu. Towards Multi-Person 3D Pose Estimation in Natural Videos. 2020. http://hdl.handle.net/1773/45769