University of Illinois Urbana-Champaign
Object understanding in generalized video segmentation
Abstract
dc:descriptionSpatio-temporal object representation is a key component in video understanding and is important for tasks like generalized video segmentation. Current methods often fail to represent objects spatio-temporally, by failing to capture the motion and interactions of objects over time, leading to identity switches, especially when objects are diverse and never seen before. This limits the generalizability of current methods. In this dissertation, we develop approaches to overcome this. We first develop AS-MOTS, a globally optimal tracker based on ‘assignment-space’ that improves tracking accuracy and enriches spatial object proposals temporally. We then develop CAROQ, an approach to generate spatio-temporally rich object proposals to unify closed-world video segmentation and seamlessly address tracking. Finally, we present OW-VISCapTor, a generalized approach that simultaneously detects, segments, and generates object-centric captions for both seen and previously unseen objects in a video.
Degree
thesis:*- Name thesis:degree_name
- Ph.D.
- Level thesis:degree_level
- Dissertation
- Discipline thesis:degree_discipline
- Electrical & Computer Engr
- Grantor
- University of Illinois Urbana-Champaign
- Year dc:date
- 2024
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Choudhuri, Anwesa
- Contributors dc:contributor
-
- Chowdhary, Girish
- Schwing, Alexander Gerhard
- Driggs-Campbell, Katherine
- Wang, Yuxiong
Subjects
dc:subject × 1Rights
dc:rights- Statement dc:rights
-
- Copyright 2024 Anwesa Choudhuri
- Language dc:language
- en
Identifiers
dc:identifier.*- Handle dc:identifier
- https://hdl.handle.net/2142/129355
- OAI identifier oai:identifier
- oai:www.ideals.illinois.edu:2142/129355