Back to results

University of Technology Sydney

Towards Comprehensive Visual Understanding via Deep Neural Networks

Abstract

dc:description.abstract

Deep neural networks (DNNs) have made significant advancements in visual scene understanding, demonstrating great potential for applications in downstream tasks such as autonomous driving, robotic navigation, and human-computer interaction. Despite these successes, generalization ability remains a major obstacle on the path to comprehensive visual understanding, particularly when dealing with i) diverse scenes, as well as ii) diverse semantic structures within those scenes. Existing work typically requires extensive annotation for different scenes (domains) and separates the understanding of semantic targets into distinct tasks, designing meticulous networks and corresponding optimization for each. This poses challenges from two perspectives: i) generalizing from one domain to another, and ii) generalizing from one task to another. To adapt an existing model to various domains (challenge i)), this thesis proposes a self-supervised learning framework to learn generalizable structural representations, and a multi-task learning framework to extract transferable knowledge from multi-modalities. To enhance a model’s ability to process various semantic structures (challenge ii)), this thesis introduces a holistic disentanglement and modeling for segment targets under an identical framework. Extensive experiments are conducted to verify the effectiveness of the proposed methods on scene understanding tasks, including Unsupervised Domain Adaptation (UDA), Exemplar-guided Video Segmentation (EVS), Video Instance Segmentation (VIS), Video Semantic Segmentation (VSS), Video Panoptic Segmentation (VPS), and Human-Object Interaction Detection (HOI Detection).

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Chen, Mu

Rights

dc:rights
Statement dc:rights
  • info:eu-repo/semantics/openAccess
  • The author owns the copyright in this thesis including all reproduction and reuse rights for the work. The work may not be altered without the permission of the copyright owner. Attribution is essential when quoting or paraphrasing from this thesis.
  • © 2025 Mu Chen
  • au.edu.uts.lib/cph
Language dc:language.iso
en_US

Identifiers

dc:identifier.*
Handle dc:identifier.uri
http://hdl.handle.net/10453/188843
OAI identifier oai:identifier
oai:opus.lib.uts.edu.au:10453/188843

Chain of custody

source
Harvested from
University of Technology Sydney
Base URL
opus.lib.uts.edu.au/oai/request
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
related terms
citation

Chen, Mu. Towards Comprehensive Visual Understanding via Deep Neural Networks. 2025. http://hdl.handle.net/10453/188843