{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/129355"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/129355","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Object understanding in generalized video segmentation","abstract":"Spatio-temporal object representation is a key component in video understanding and is important for tasks like generalized video segmentation. Current methods often fail to represent objects spatio-temporally, by failing to capture the motion and interactions of objects over time, leading to identity switches, especially when objects are diverse and never seen before. This limits the generalizability of current methods. In this dissertation, we develop approaches to overcome this. We first develop AS-MOTS, a globally optimal tracker based on ‘assignment-space’ that improves tracking accuracy and enriches spatial object proposals temporally. We then develop CAROQ, an approach to generate spatio-temporally rich object proposals to unify closed-world video segmentation and seamlessly address tracking. Finally, we present OW-VISCapTor, a generalized approach that simultaneously detects, segments, and generates object-centric captions for both seen and previously unseen objects in a video.","abstract_html":"Spatio-temporal object representation is a key component in video understanding and is important for tasks like generalized video segmentation. Current methods often fail to represent objects spatio-temporally, by failing to capture the motion and interactions of objects over time, leading to identity switches, especially when objects are diverse and never seen before. This limits the generalizability of current methods. In this dissertation, we develop approaches to overcome this. We first develop AS-MOTS, a globally optimal tracker based on ‘assignment-space’ that improves tracking accuracy and enriches spatial object proposals temporally. We then develop CAROQ, an approach to generate spatio-temporally rich object proposals to unify closed-world video segmentation and seamlessly address tracking. Finally, we present OW-VISCapTor, a generalized approach that simultaneously detects, segments, and generates object-centric captions for both seen and previously unseen objects in a video.","abstract_has_math":false,"creators":["Choudhuri, Anwesa"],"institution":"University of Illinois Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Chowdhary, Girish","Schwing, Alexander Gerhard","Driggs-Campbell, Katherine","Wang, Yuxiong"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-08","date_published":"2024-08","updated_at":"2026-07-22T22:25:05Z","subjects":["Generalized video segmentation, video understanding"],"languages":["en"],"rights":["Copyright 2024 Anwesa Choudhuri"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/129355","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Chowdhary, Girish","Schwing, Alexander Gerhard","Driggs-Campbell, Katherine","Wang, Yuxiong"]},{"key":"dc:creator","label":"Author","values":["Choudhuri, Anwesa"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2024-08","2024-07-10"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Generalized video segmentation, video understanding"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2024 Anwesa Choudhuri"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/129355"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Spatio-temporal object representation is a key component in video understanding and is important for tasks like generalized video segmentation. Current methods often fail to represent objects spatio-temporally, by failing to capture the motion and interactions of objects over time, leading to identity switches, especially when objects are diverse and never seen before. This limits the generalizability of current methods. In this dissertation, we develop approaches to overcome this. We first develop AS-MOTS, a globally optimal tracker based on ‘assignment-space’ that improves tracking accuracy and enriches spatial object proposals temporally. We then develop CAROQ, an approach to generate spatio-temporally rich object proposals to unify closed-world video segmentation and seamlessly address tracking. Finally, we present OW-VISCapTor, a generalized approach that simultaneously detects, segments, and generates object-centric captions for both seen and previously unseen objects in a video.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","The student, Anwesa Choudhuri, accepted the attached license on 2024-07-09 at 09:04.","The student, Anwesa Choudhuri, submitted this Dissertation for approval on 2024-07-09 at 09:05.","This Dissertation was approved for publication on 2024-07-10 at 15:26.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20998 on 2025-10-19 at 18:16:56"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Object understanding in generalized video segmentation"]}]}],"canonical_facts":{"dc:contributor":["Chowdhary, Girish","Schwing, Alexander Gerhard","Driggs-Campbell, Katherine","Wang, Yuxiong"],"dc:creator":["Choudhuri, Anwesa"],"dc:date":["2024-08","2024-07-10"],"dc:description":["Spatio-temporal object representation is a key component in video understanding and is important for tasks like generalized video segmentation. Current methods often fail to represent objects spatio-temporally, by failing to capture the motion and interactions of objects over time, leading to identity switches, especially when objects are diverse and never seen before. This limits the generalizability of current methods. In this dissertation, we develop approaches to overcome this. We first develop AS-MOTS, a globally optimal tracker based on ‘assignment-space’ that improves tracking accuracy and enriches spatial object proposals temporally. We then develop CAROQ, an approach to generate spatio-temporally rich object proposals to unify closed-world video segmentation and seamlessly address tracking. Finally, we present OW-VISCapTor, a generalized approach that simultaneously detects, segments, and generates object-centric captions for both seen and previously unseen objects in a video.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","The student, Anwesa Choudhuri, accepted the attached license on 2024-07-09 at 09:04.","The student, Anwesa Choudhuri, submitted this Dissertation for approval on 2024-07-09 at 09:05.","This Dissertation was approved for publication on 2024-07-10 at 15:26.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20998 on 2025-10-19 at 18:16:56"],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/129355"],"dc:language":["en"],"dc:rights":["Copyright 2024 Anwesa Choudhuri"],"dc:subject":["Generalized video segmentation, video understanding"],"dc:title":["Object understanding in generalized video segmentation"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:05Z"}