University of Illinois at Urbana-Champaign
From pixels to regions: Toward universal image segmentation
Abstract
dc:descriptionImage segmentation is about grouping pixels with different semantics, e.g., category or instance membership, where each choice of semantics defines a task. While only the semantics of each task differ, current research focuses on designing specialized architectures for each task: semantic segmentation is usually formulated as per-pixel classification and mask classification dominates instance-level segmentation tasks. In this dissertation, we demonstrate how to build a single unified architecture that can address any image segmentation task. We first introduce an effort in unifying image segmentation with either per-pixel classification (Panoptic-DeepLab) or mask classification (MaskFormer). We observe mask classification is sufficiently general to solve both semantic- and instance-level segmentation tasks. Based on this observation we propose Mask2Former, which outperforms even the best specialized architectures by a significant margin on four popular datasets for three image segmentation tasks (panoptic, instance and semantic). Then we discuss how to evaluate image segmentation models with a new Boundary IoU metric. Finally, we conclude this dissertation with promising future directions to explore.
Degree
thesis:*- Name thesis:degree_name
- Ph.D.
- Level thesis:degree_level
- Dissertation
- Discipline thesis:degree_discipline
- Electrical & Computer Engr
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2022
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Cheng, Bowen
- Contributors dc:contributor
-
- Schwing, Alexander
- Shi, Humphrey
- Hasegawa-Johnson, Mark
- Darrell, Trevor
- Liang, Zhi-Pei
Subjects
dc:subject × 5Rights
dc:rights- Statement dc:rights
-
- Copyright 2022 Bowen Cheng
- Language dc:language
- en, eng
Identifiers
dc:identifier.*- Handle dc:identifier
- https://hdl.handle.net/2142/116182