Abstract
dc:descriptionWe explore approaches to improve over existing amodal prediction models for the task of semantic amodal instance level video object segmentation, i.e., the task to delineate objects and their occluded parts in video data. We propose Amodal-Net with three improvements: First, we leverage temporal information. Specifically, we employ 3D convolutions and a flow alignment module which permits to aggregate the objects’ features across frames. Second, we develop a cascaded box-head with soft-non-maximum-suppression to address the challenge that amodal segmentations overlap significantly. Third, we address the challenge that occlusions require observation information to be propagated over larger distances by developing an attention-based mask-head. Then we also study reprojection, another way of using temporal information which also uses 3D information. We evaluate our approach on amodal segmentation for video data, SAILVOS.
Degree
thesis:*- Name thesis:degree_name
- M.S.
- Level thesis:degree_level
- Thesis
- Discipline thesis:degree_discipline
- Computer Science
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2022
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Sun, Mingxi
- Contributors dc:contributor
-
- Schwing, Alexander
Subjects
dc:subject × 4Rights
dc:rights- Statement dc:rights
-
- Copyright 2021 Mingxi Sun
- Language dc:language
- en
Identifiers
dc:identifier.*- Handle dc:identifier
- http://hdl.handle.net/2142/113229
- OAI identifier oai:identifier
- oai:www.ideals.illinois.edu:2142/113229