{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/116107"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/116107","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Towards holistic scene understanding from monocular video","abstract":"Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2024-08-01","abstract_html":"Submission published under a 24 month embargo labeled &#x27;U of I Access&#x27;, the embargo will last until 2024-08-01","abstract_has_math":false,"creators":["Hu, Yuan-Ting"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Schwing, Alexander Gerhard","Forsyth, David","Hoiem, Derek","Patel, Sanjay","Huang, Jia-Bin"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-08","date_published":"2022-08","updated_at":"2026-07-22T22:24:55Z","subjects":["Scene understanding","Video understanding","Video segmentation","Object segmentation","Amodal segmentation","3D reconstruction"],"languages":["en","eng"],"rights":["Copyright 2022 Yuan-Ting Hu"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/116107","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Schwing, Alexander Gerhard","Forsyth, David","Hoiem, Derek","Patel, Sanjay","Huang, Jia-Bin"]},{"key":"dc:creator","label":"Author","values":["Hu, Yuan-Ting"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022-08","2022-07-15"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Scene understanding","Video understanding","Video segmentation","Object segmentation","Amodal segmentation","3D reconstruction"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2022 Yuan-Ting Hu"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/116107"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2024-08-01","The student, Yuan-Ting Hu, accepted the attached license on 2022-07-14 at 15:33.","The student, Yuan-Ting Hu, submitted this Dissertation for approval on 2022-07-14 at 17:56.","This Dissertation was approved for publication on 2022-07-15 at 09:58.","DSpace SAF Submission Ingestion Package generated from Vireo submission #18316 on 2022-11-15 at 21:40:22","Humans have the remarkable ability to vividly envision future scenarios as they are capable of understanding scenes in a holistic manner. They can extrapolate scene information such as object shapes and interactions from the observed scene content and its dynamics. Importantly, they can reason about unseen information, e.g., when objects are partially observed. In contrast, while computer vision and machine learning systems can successfully explain observations, it remains challenging to develop autonomous agents that can infer the unseen and have a holistic understanding of the environment. In this dissertation, we discuss techniques that tackle research problems related to holistic scene understanding from monocular video data. To study holistic scene understanding from monocular video, we first present models for human pose understanding from video. Second, we study the research problem of track moving objects under challenging conditions such as occlusion and appearance change. Third, we then consider a challenging task, amodal understanding of objects in a scene from video, aiming to infer the entirety of objects even if they are only partially observed. To enable data-driven approaches towards video amodal perception, we present a large-scale video dataset where more than 1.8 million objects are annotated with amodal labels. With the proposed dataset, we study and present video algorithms that infer the unseen and understand the scene dynamics as well as 3D shapes with partially occluded data. Last, we present a method to show how geometric cues predicted from 2D can improve 3D understanding of objects in the scene. We then conclude and discuss future directions towards holistic video scene understanding."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Towards holistic scene understanding from monocular video"]}]}],"canonical_facts":{"dc:contributor":["Schwing, Alexander Gerhard","Forsyth, David","Hoiem, Derek","Patel, Sanjay","Huang, Jia-Bin"],"dc:creator":["Hu, Yuan-Ting"],"dc:date":["2022-08","2022-07-15"],"dc:description":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2024-08-01","The student, Yuan-Ting Hu, accepted the attached license on 2022-07-14 at 15:33.","The student, Yuan-Ting Hu, submitted this Dissertation for approval on 2022-07-14 at 17:56.","This Dissertation was approved for publication on 2022-07-15 at 09:58.","DSpace SAF Submission Ingestion Package generated from Vireo submission #18316 on 2022-11-15 at 21:40:22","Humans have the remarkable ability to vividly envision future scenarios as they are capable of understanding scenes in a holistic manner. They can extrapolate scene information such as object shapes and interactions from the observed scene content and its dynamics. Importantly, they can reason about unseen information, e.g., when objects are partially observed. In contrast, while computer vision and machine learning systems can successfully explain observations, it remains challenging to develop autonomous agents that can infer the unseen and have a holistic understanding of the environment. In this dissertation, we discuss techniques that tackle research problems related to holistic scene understanding from monocular video data. To study holistic scene understanding from monocular video, we first present models for human pose understanding from video. Second, we study the research problem of track moving objects under challenging conditions such as occlusion and appearance change. Third, we then consider a challenging task, amodal understanding of objects in a scene from video, aiming to infer the entirety of objects even if they are only partially observed. To enable data-driven approaches towards video amodal perception, we present a large-scale video dataset where more than 1.8 million objects are annotated with amodal labels. With the proposed dataset, we study and present video algorithms that infer the unseen and understand the scene dynamics as well as 3D shapes with partially occluded data. Last, we present a method to show how geometric cues predicted from 2D can improve 3D understanding of objects in the scene. We then conclude and discuss future directions towards holistic video scene understanding."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/116107"],"dc:language":["en","eng"],"dc:rights":["Copyright 2022 Yuan-Ting Hu"],"dc:subject":["Scene understanding","Video understanding","Video segmentation","Object segmentation","Amodal segmentation","3D reconstruction"],"dc:title":["Towards holistic scene understanding from monocular video"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:55Z"}