{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/121199"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/121199","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Robust online video instance segmentation with track queries","abstract":"Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2025-08-01","abstract_html":"Submission published under a 24 month embargo labeled &#x27;Closed Access&#x27;, the embargo will last until 2025-08-01","abstract_has_math":false,"creators":["Zhan, Zitong"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Lazebnik, Svetlana"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2023,"date_issued":"2023-08","date_published":"2023-08","updated_at":"2026-07-22T22:24:57Z","subjects":["Computer Vision","Video Tracking"],"languages":["en","eng"],"rights":["Copyright 2023 Zitong Zhan"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/121199","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Lazebnik, Svetlana"]},{"key":"dc:creator","label":"Author","values":["Zhan, Zitong"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2023-08","2023-07-17"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer Vision","Video Tracking"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2023 Zitong Zhan"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/121199"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2025-08-01","The student, Zitong Zhan, accepted the attached license on 2023-07-11 at 14:26.","The student, Zitong Zhan, submitted this Thesis for approval on 2023-07-11 at 14:33.","This Thesis was approved for publication on 2023-07-17 at 12:04.","DSpace SAF Submission Ingestion Package generated from Vireo submission #19397 on 2023-12-04 at 17:30:25","With support of data and maturity of model, deep learning has been successful in high level understanding on videos. Recently, transformer-based methods have achieved impressive results on Video Instance Segmentation (VIS). However, most of these top-performing methods run in an offline manner by processing the entire video clip at once to predict instance mask volumes. This makes them incapable of handling the long videos that appear in challenging new video instance segmentation datasets like Unidentified Video Objects (UVO) and Occluded Video Instance Segmentation (OVIS). We propose a fully online transformer-based video instance segmentation model that performs comparably to top offline methods on the YouTube-VIS 2019 benchmark and considerably outperforms them on UVO and OVIS. This method, called Robust Online Video Segmentation (ROVIS), augments the Mask2Former image instance segmentation model with track queries, a lightweight mechanism for carrying track information from frame to frame, originally introduced by the TrackFormer method for multi-object tracking. We show that, when combined with a strong enough image segmentation architecture, track queries can exhibit impressive accuracy while not being constrained to short videos. Lastly, we extend ROVIS to the Refer-VOS task which is highly relevant to VIS, such that ROVIS can make segmentation predictions based on language input, and present our preliminary results."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Robust online video instance segmentation with track queries"]}]}],"canonical_facts":{"dc:contributor":["Lazebnik, Svetlana"],"dc:creator":["Zhan, Zitong"],"dc:date":["2023-08","2023-07-17"],"dc:description":["Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2025-08-01","The student, Zitong Zhan, accepted the attached license on 2023-07-11 at 14:26.","The student, Zitong Zhan, submitted this Thesis for approval on 2023-07-11 at 14:33.","This Thesis was approved for publication on 2023-07-17 at 12:04.","DSpace SAF Submission Ingestion Package generated from Vireo submission #19397 on 2023-12-04 at 17:30:25","With support of data and maturity of model, deep learning has been successful in high level understanding on videos. Recently, transformer-based methods have achieved impressive results on Video Instance Segmentation (VIS). However, most of these top-performing methods run in an offline manner by processing the entire video clip at once to predict instance mask volumes. This makes them incapable of handling the long videos that appear in challenging new video instance segmentation datasets like Unidentified Video Objects (UVO) and Occluded Video Instance Segmentation (OVIS). We propose a fully online transformer-based video instance segmentation model that performs comparably to top offline methods on the YouTube-VIS 2019 benchmark and considerably outperforms them on UVO and OVIS. This method, called Robust Online Video Segmentation (ROVIS), augments the Mask2Former image instance segmentation model with track queries, a lightweight mechanism for carrying track information from frame to frame, originally introduced by the TrackFormer method for multi-object tracking. We show that, when combined with a strong enough image segmentation architecture, track queries can exhibit impressive accuracy while not being constrained to short videos. Lastly, we extend ROVIS to the Refer-VOS task which is highly relevant to VIS, such that ROVIS can make segmentation predictions based on language input, and present our preliminary results."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/121199"],"dc:language":["en","eng"],"dc:rights":["Copyright 2023 Zitong Zhan"],"dc:subject":["Computer Vision","Video Tracking"],"dc:title":["Robust online video instance segmentation with track queries"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:57Z"}