Back to results

University of Adelaide

Self-supervised Learning of Monocular Depth from Video

Abstract

dc:description.abstract

Image-based depth estimation as a fundamental problem in computer vision allows for understanding the scene geometry using only cameras. This thesis addresses the specific problem of monocular depth estimation via self-supervised learning from RGB-only videos. Although existing work has shown partial excellent results in benchmark datasets, there remain several vital challenges that limit the use of these algorithms in general scenarios. To summarize, my identified challenges and contributions include: (i) Previous methods predict inconsistent depths over a video, which limits their uses in visual localization and mapping. To this end, I propose a geometry consistency loss that penalizes the multi-view depth misalignment in training, which enables scale-consistent depth estimation at inference time; (ii) Previous methods often diverge or show low-accuracy results when training on handheld camera captured videos. To address the challenge, I analyze the effect of camera motion on depth network gradients, and I propose an auto-rectify network to remove the relative rotation in training image pairs for robust learning; (iii) Previous methods fail to learn reasonable depths from highly dynamic scenes due to the non-rigidity. In this scenario, I propose a novel method, which constrains dynamic regions using an external well-trained depth estimation network and supervises static regions via multi-view losses. Comprehensive quantitative results and rich qualitative results are provided to demonstrate the advantages of the proposed methods over existing alternatives. The codes and pre-trained models have been released at https://github.com/JiawangBian

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Bian, Jiawang
Advisors dc:contributor.advisor
  • Reid, Ian
  • Shen, Chunhua (Zhejiang University)

Subjects

dc:subject × 1

Rights

Language dc:language.iso
en

Identifiers

dc:identifier.*
Handle dc:identifier.uri
https://hdl.handle.net/2440/136692
OAI identifier oai:identifier
oai:digital.library.adelaide.edu.au:2440/136692

Chain of custody

source
Harvested from
University of Adelaide
Base URL
digital.library.adelaide.edu.au/server/oai/request
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Bian, Jiawang. Self-supervised Learning of Monocular Depth from Video. 2022. https://hdl.handle.net/2440/136692