{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/115402"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/115402","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Tracking objects and distinguishing their states by watching egocentric videos","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-11 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2022-11-11 without embargo terms","abstract_has_math":false,"creators":["Modi, Sahil Ketan"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Gupta, Saurabh"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-05","date_published":"2022-05","updated_at":"2026-07-22T22:24:54Z","subjects":["computer vision","tracking","egocentric"],"languages":["en","eng"],"rights":["Copyright 2022 Sahil Modi"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/115402","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Gupta, Saurabh"]},{"key":"dc:creator","label":"Author","values":["Modi, Sahil Ketan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022-05","2022-04-26"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["computer vision","tracking","egocentric"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2022 Sahil Modi"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/115402"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-11 without embargo terms","The student, Sahil Modi, accepted the attached license on 2022-04-20 at 10:51.","The student, Sahil Modi, submitted this Thesis for approval on 2022-04-20 at 10:57.","This Thesis was approved for publication on 2022-04-26 at 15:01.","DSpace SAF Submission Ingestion Package generated from Vireo submission #17835 on 2022-11-11 at 13:42:25","Interactive object understanding, or what we can do to objects and how, is a long-standing goal of computer vision. However, the inherent ambiguity of this task makes it difficult to annotate, and very few large-scale datasets exist. We realize that videos, especially egocentric ones, naturally contain this information through objects undergoing constant state changes, but learning from this data is nontrivial. Furthermore, objects are difficult to track in egocentric settings due to occlusion, drastic pose changes, and viewpoint changes. In this thesis, we propose solutions for these two challenges by (1) taking advantage of existing sparse annotations and self-supervision to achieve state-of-the-art tracking performance on TREK-150 and (2) observing human hands and their interactions with objects to learn object state-sensitive features in a self-supervised manner."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Tracking objects and distinguishing their states by watching egocentric videos"]}]}],"canonical_facts":{"dc:contributor":["Gupta, Saurabh"],"dc:creator":["Modi, Sahil Ketan"],"dc:date":["2022-05","2022-04-26"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-11 without embargo terms","The student, Sahil Modi, accepted the attached license on 2022-04-20 at 10:51.","The student, Sahil Modi, submitted this Thesis for approval on 2022-04-20 at 10:57.","This Thesis was approved for publication on 2022-04-26 at 15:01.","DSpace SAF Submission Ingestion Package generated from Vireo submission #17835 on 2022-11-11 at 13:42:25","Interactive object understanding, or what we can do to objects and how, is a long-standing goal of computer vision. However, the inherent ambiguity of this task makes it difficult to annotate, and very few large-scale datasets exist. We realize that videos, especially egocentric ones, naturally contain this information through objects undergoing constant state changes, but learning from this data is nontrivial. Furthermore, objects are difficult to track in egocentric settings due to occlusion, drastic pose changes, and viewpoint changes. In this thesis, we propose solutions for these two challenges by (1) taking advantage of existing sparse annotations and self-supervision to achieve state-of-the-art tracking performance on TREK-150 and (2) observing human hands and their interactions with objects to learn object state-sensitive features in a self-supervised manner."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/115402"],"dc:language":["en","eng"],"dc:rights":["Copyright 2022 Sahil Modi"],"dc:subject":["computer vision","tracking","egocentric"],"dc:title":["Tracking objects and distinguishing their states by watching egocentric videos"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:54Z"}