Back to results

University of Technology Sydney

Learning Object Detection with Weak Supervision

Abstract

dc:description.abstract

Deep learning technique has achieved astonishing success in many computer vision applications. However, training deep models typically requires large-scale datasets with elaborate annotations. Collecting and annotating large-scale datasets are laborious, especially for object detection --- a challenging vision task. A promising solution for reducing costs is to train models with weak supervision, which provides a good trade-off between model performance and annotation efficiency. This thesis dedicates to weakly supervised learning in two object-centered application scenarios, i.e., general object detection and RGB-D salient object detection. The first task is to predict the category of an object and its location in the given image with image-level weak supervision. A pyramidal multiple instance detection network is first introduced to reduce the exposure of local discriminative proposal regions, alleviating the local optimum issue in training detectors with only image-level annotations. Besides learning detectors with only image-level supervision, two more practical scenarios in weakly supervised object detection are considered. With a well-annotated object detection dataset, this thesis further investigates how to scale detectors to novel domains or categories using weak supervision. Concretely, a holistic and hierarchical feature alignment R-CNN is presented to perform coarse-to-fine alignments in pace with the detection pipeline and effectively reduce the discrepancy between different domains with weak supervision. A cyclic self-training framework with a proposal weight modulation module is introduced to compensate for the instance-level supervision of novel classes and adaptively adjust loss weights for the training samples. The second task is to predict pixel-level masks for the foreground objects in paired RGB-D inputs (i.e., images and depth maps) with scribble-based weak supervision. This thesis explores annotator-friendly scribble annotations for training models. A dual-modal edge-guided network and a prediction consistency training method are developed to fully take advantage of the complementary information from both modalities and exploit the information residing in the unlabeled pixels, respectively. Extensive experiments are conducted and analyzed to evaluate the effectiveness of the proposed approaches with weak supervision. Competitive performance on commonly used benchmarks verifies the effectiveness and universality.

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Xu, Yunqiu

Rights

dc:rights
Statement dc:rights
  • info:eu-repo/semantics/openAccess
  • The author owns the copyright in this thesis including all reproduction and reuse rights for the work. The work may not be altered without the permission of the copyright owner. Attribution is essential when quoting or paraphrasing from this thesis.
  • @ 2023 Yunqiu Xu
  • au.edu.uts.lib/cph
Language dc:language.iso
en_US

Identifiers

dc:identifier.*
Handle dc:identifier.uri
http://hdl.handle.net/10453/170728
OAI identifier oai:identifier
oai:opus.lib.uts.edu.au:10453/170728

Chain of custody

source
Harvested from
University of Technology Sydney
Base URL
opus.lib.uts.edu.au/oai/request
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
related terms
citation

Xu, Yunqiu. Learning Object Detection with Weak Supervision. 2023. http://hdl.handle.net/10453/170728