University of Illinois at Urbana-Champaign
Tracking objects and distinguishing their states by watching egocentric videos
Abstract
dc:descriptionInteractive object understanding, or what we can do to objects and how, is a long-standing goal of computer vision. However, the inherent ambiguity of this task makes it difficult to annotate, and very few large-scale datasets exist. We realize that videos, especially egocentric ones, naturally contain this information through objects undergoing constant state changes, but learning from this data is nontrivial. Furthermore, objects are difficult to track in egocentric settings due to occlusion, drastic pose changes, and viewpoint changes. In this thesis, we propose solutions for these two challenges by (1) taking advantage of existing sparse annotations and self-supervision to achieve state-of-the-art tracking performance on TREK-150 and (2) observing human hands and their interactions with objects to learn object state-sensitive features in a self-supervised manner.
Degree
thesis:*- Name thesis:degree_name
- M.S.
- Level thesis:degree_level
- Thesis
- Discipline thesis:degree_discipline
- Computer Science
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2022
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Modi, Sahil Ketan
- Contributors dc:contributor
-
- Gupta, Saurabh
Subjects
dc:subject × 3Rights
dc:rights- Statement dc:rights
-
- Copyright 2022 Sahil Modi
- Language dc:language
- en, eng
Identifiers
dc:identifier.*- Handle dc:identifier
- https://hdl.handle.net/2142/115402