University of Technology Sydney
Towards Comprehensive Visual Understanding via Deep Neural Networks
Abstract
dc:description.abstractDeep neural networks (DNNs) have made significant advancements in visual scene understanding, demonstrating great potential for applications in downstream tasks such as autonomous driving, robotic navigation, and human-computer interaction. Despite these successes, generalization ability remains a major obstacle on the path to comprehensive visual understanding, particularly when dealing with i) diverse scenes, as well as ii) diverse semantic structures within those scenes. Existing work typically requires extensive annotation for different scenes (domains) and separates the understanding of semantic targets into distinct tasks, designing meticulous networks and corresponding optimization for each. This poses challenges from two perspectives: i) generalizing from one domain to another, and ii) generalizing from one task to another. To adapt an existing model to various domains (challenge i)), this thesis proposes a self-supervised learning framework to learn generalizable structural representations, and a multi-task learning framework to extract transferable knowledge from multi-modalities. To enhance a model’s ability to process various semantic structures (challenge ii)), this thesis introduces a holistic disentanglement and modeling for segment targets under an identical framework. Extensive experiments are conducted to verify the effectiveness of the proposed methods on scene understanding tasks, including Unsupervised Domain Adaptation (UDA), Exemplar-guided Video Segmentation (EVS), Video Instance Segmentation (VIS), Video Semantic Segmentation (VSS), Video Panoptic Segmentation (VPS), and Human-Object Interaction Detection (HOI Detection).
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Chen, Mu
Rights
dc:rights- Statement dc:rights
-
- info:eu-repo/semantics/openAccess
- The author owns the copyright in this thesis including all reproduction and reuse rights for the work. The work may not be altered without the permission of the copyright owner. Attribution is essential when quoting or paraphrasing from this thesis.
- © 2025 Mu Chen
- au.edu.uts.lib/cph
- Language dc:language.iso
- en_US
Identifiers
dc:identifier.*- Handle dc:identifier.uri
- http://hdl.handle.net/10453/188843
- OAI identifier oai:identifier
- oai:opus.lib.uts.edu.au:10453/188843