{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/113229"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/113229","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Amodal video instance segmentation","abstract":"We explore approaches to improve over existing amodal prediction models for the task of semantic amodal instance level video object segmentation, i.e., the task to delineate objects and their occluded parts in video data. We propose Amodal-Net with three improvements: First, we leverage temporal information. Specifically, we employ 3D convolutions and a flow alignment module which permits to aggregate the objects’ features across frames. Second, we develop a cascaded box-head with soft-non-maximum-suppression to address the challenge that amodal segmentations overlap significantly. Third, we address the challenge that occlusions require observation information to be propagated over larger distances by developing an attention-based mask-head. Then we also study reprojection, another way of using temporal information which also uses 3D information. We evaluate our approach on amodal segmentation for video data, SAILVOS.","abstract_html":"We explore approaches to improve over existing amodal prediction models for the task of semantic amodal instance level video object segmentation, i.e., the task to delineate objects and their occluded parts in video data. We propose Amodal-Net with three improvements: First, we leverage temporal information. Specifically, we employ 3D convolutions and a flow alignment module which permits to aggregate the objects’ features across frames. Second, we develop a cascaded box-head with soft-non-maximum-suppression to address the challenge that amodal segmentations overlap significantly. Third, we address the challenge that occlusions require observation information to be propagated over larger distances by developing an attention-based mask-head. Then we also study reprojection, another way of using temporal information which also uses 3D information. We evaluate our approach on amodal segmentation for video data, SAILVOS.","abstract_has_math":false,"creators":["Sun, Mingxi"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Schwing, Alexander"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-01-12T22:35:26Z","date_published":"2022-01-12T22:35:26Z","updated_at":"2026-07-22T22:24:53Z","subjects":["machine learning","computer vision","instance segmentation","amodal segmentation"],"languages":["en"],"rights":["Copyright 2021 Mingxi Sun"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/113229","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Schwing, Alexander"]},{"key":"dc:creator","label":"Author","values":["Sun, Mingxi"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022-01-12T22:35:26Z","2024-01-12T22:35:30Z","2021-07-23","2021-08"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["machine learning","computer vision","instance segmentation","amodal segmentation"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2021 Mingxi Sun"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/113229"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["We explore approaches to improve over existing amodal prediction models for the task of semantic amodal instance level video object segmentation, i.e., the task to delineate objects and their occluded parts in video data. We propose Amodal-Net with three improvements: First, we leverage temporal information. Specifically, we employ 3D convolutions and a flow alignment module which permits to aggregate the objects’ features across frames. Second, we develop a cascaded box-head with soft-non-maximum-suppression to address the challenge that amodal segmentations overlap significantly. Third, we address the challenge that occlusions require observation information to be propagated over larger distances by developing an attention-based mask-head. Then we also study reprojection, another way of using temporal information which also uses 3D information. We evaluate our approach on amodal segmentation for video data, SAILVOS.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2023-08-01","The student, Mingxi Sun, accepted the attached license on 2021-07-23 at 12:35.","The student, Mingxi Sun, submitted this Thesis for approval on 2021-07-23 at 12:45.","This Thesis was approved for publication on 2021-07-23 at 13:07.","DSpace SAF Submission Ingestion Package generated from Vireo submission #17087 on 2022-01-12 at 12:55:40","Made available in DSpace on 2022-01-12T22:35:26Z (GMT). No. of bitstreams: 2 SUN-THESIS-2021.pdf: 58659078 bytes, checksum: 43dadfc2463b133778b27d595a7f9d79 (MD5) LICENSE.txt: 4205 bytes, checksum: cc8808e79fc95466674f60cd0c160f41 (MD5) Previous issue date: 2021-07-23","Embargo set by: Seth Robbins for item 121155 Lift date: 2024-01-12T22:35:30Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Amodal video instance segmentation"]}]}],"canonical_facts":{"dc:contributor":["Schwing, Alexander"],"dc:creator":["Sun, Mingxi"],"dc:date":["2022-01-12T22:35:26Z","2024-01-12T22:35:30Z","2021-07-23","2021-08"],"dc:description":["We explore approaches to improve over existing amodal prediction models for the task of semantic amodal instance level video object segmentation, i.e., the task to delineate objects and their occluded parts in video data. We propose Amodal-Net with three improvements: First, we leverage temporal information. Specifically, we employ 3D convolutions and a flow alignment module which permits to aggregate the objects’ features across frames. Second, we develop a cascaded box-head with soft-non-maximum-suppression to address the challenge that amodal segmentations overlap significantly. Third, we address the challenge that occlusions require observation information to be propagated over larger distances by developing an attention-based mask-head. Then we also study reprojection, another way of using temporal information which also uses 3D information. We evaluate our approach on amodal segmentation for video data, SAILVOS.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2023-08-01","The student, Mingxi Sun, accepted the attached license on 2021-07-23 at 12:35.","The student, Mingxi Sun, submitted this Thesis for approval on 2021-07-23 at 12:45.","This Thesis was approved for publication on 2021-07-23 at 13:07.","DSpace SAF Submission Ingestion Package generated from Vireo submission #17087 on 2022-01-12 at 12:55:40","Made available in DSpace on 2022-01-12T22:35:26Z (GMT). No. of bitstreams: 2 SUN-THESIS-2021.pdf: 58659078 bytes, checksum: 43dadfc2463b133778b27d595a7f9d79 (MD5) LICENSE.txt: 4205 bytes, checksum: cc8808e79fc95466674f60cd0c160f41 (MD5) Previous issue date: 2021-07-23","Embargo set by: Seth Robbins for item 121155 Lift date: 2024-01-12T22:35:30Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/113229"],"dc:language":["en"],"dc:rights":["Copyright 2021 Mingxi Sun"],"dc:subject":["machine learning","computer vision","instance segmentation","amodal segmentation"],"dc:title":["Amodal video instance segmentation"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:53Z"}