{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/29773"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/29773","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"3D spatial layout and geometric constraints for scene understanding","abstract":"An image is nothing but a projection of the physical world around us, where objects do not occur randomly but follow certain spatial rules. Many existing computer vision approaches tend to ignore this aspect of understanding images. In this work, we build representations and propose strategies for exploiting such constraints towards extracting a 3D understanding of a scene from its single image. We model a scene in terms of its spatial layout abstracted as a box, object cuboids, camera viewpoint, and interactions between them. We take a supervised approach towards estimation, and learn models from training data that is fully annotated with the 3D spatial extent of objects, walls, and floor. We assume the world is populated with axis aligned objects and surfaces, and exploit constrained appearance models which use geometric cues from the scene. Our methods are tailored towards indoor scenes that are highly structured and require careful spatial reasoning. We show that our box layout representation is able to capture the full spatial extent of a 3D scene, which we can successfully estimate even for heavily cluttered rooms. Similarly, by exploiting the geometric constraints offered by the scene, we can approximate the extent of the objects as cuboids in 3D. The box layout provides rich contextual information for detecting objects. We show that modeling the 3D interactions between object cuboids and scene layout improves object detection. Finally, we show how to use our 3D spatial layout models together with object cuboid models to predict the free space in the scene.","abstract_html":"An image is nothing but a projection of the physical world around us, where objects do not occur randomly but follow certain spatial rules. Many existing computer vision approaches tend to ignore this aspect of understanding images. In this work, we build representations and propose strategies for exploiting such constraints towards extracting a 3D understanding of a scene from its single image. We model a scene in terms of its spatial layout abstracted as a box, object cuboids, camera viewpoint, and interactions between them. We take a supervised approach towards estimation, and learn models from training data that is fully annotated with the 3D spatial extent of objects, walls, and floor. We assume the world is populated with axis aligned objects and surfaces, and exploit constrained appearance models which use geometric cues from the scene. Our methods are tailored towards indoor scenes that are highly structured and require careful spatial reasoning. We show that our box layout representation is able to capture the full spatial extent of a 3D scene, which we can successfully estimate even for heavily cluttered rooms. Similarly, by exploiting the geometric constraints offered by the scene, we can approximate the extent of the objects as cuboids in 3D. The box layout provides rich contextual information for detecting objects. We show that modeling the 3D interactions between object cuboids and scene layout improves object detection. Finally, we show how to use our 3D spatial layout models together with object cuboid models to predict the free space in the scene.","abstract_has_math":false,"creators":["Hedau, Varsha"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Forsyth, David A.","Hoiem, Derek W.","Ma, Yi","Huang, Thomas S.","Hasegawa-Johnson, Mark A."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2012,"date_issued":"2012-02-06T20:15:35Z","date_published":"2012-02-06T20:15:35Z","updated_at":"2026-07-22T22:25:29Z","subjects":["Computer vision","Image Processing","Scene understanding","Object recognition","Single view 3D modeling","Vanishing points"],"languages":["en"],"rights":["copyright 2011 Varsha Chandrashekhar Hedau"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/29773","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Forsyth, David A.","Hoiem, Derek W.","Ma, Yi","Huang, Thomas S.","Hasegawa-Johnson, Mark A."]},{"key":"dc:creator","label":"Author","values":["Hedau, Varsha"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2012-02-06T20:15:35Z","2011-12"]},{"key":"dc:type","label":"Dc Type","values":["Dissertation / Thesis","text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer vision","Image Processing","Scene understanding","Object recognition","Single view 3D modeling","Vanishing points"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["copyright 2011 Varsha Chandrashekhar Hedau"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/29773"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["An image is nothing but a projection of the physical world around us, where objects do not occur randomly but follow certain spatial rules. Many existing computer vision approaches tend to ignore this aspect of understanding images. In this work, we build representations and propose strategies for exploiting such constraints towards extracting a 3D understanding of a scene from its single image. We model a scene in terms of its spatial layout abstracted as a box, object cuboids, camera viewpoint, and interactions between them. We take a supervised approach towards estimation, and learn models from training data that is fully annotated with the 3D spatial extent of objects, walls, and floor. We assume the world is populated with axis aligned objects and surfaces, and exploit constrained appearance models which use geometric cues from the scene. Our methods are tailored towards indoor scenes that are highly structured and require careful spatial reasoning. We show that our box layout representation is able to capture the full spatial extent of a 3D scene, which we can successfully estimate even for heavily cluttered rooms. Similarly, by exploiting the geometric constraints offered by the scene, we can approximate the extent of the objects as cuboids in 3D. The box layout provides rich contextual information for detecting objects. We show that modeling the 3D interactions between object cuboids and scene layout improves object detection. Finally, we show how to use our 3D spatial layout models together with object cuboid models to predict the free space in the scene.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2011-11-22T21:02:43Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Hedau_Varsha.pdf: 8407684 bytes, checksum: 1255b10542309447c35bd2c19c70fd09 (MD5)","Made available in DSpace on 2012-02-06T20:15:35Z (GMT). No. of bitstreams: 2 Hedau_Varsha.pdf: 8408202 bytes, checksum: 2729be9a90291d2763f1efd4d0aa627b (MD5) license.txt: 4061 bytes, checksum: 631f1b9217a18917c749d33b59f24405 (MD5)"]},{"key":"dc:title","label":"Title","values":["3D spatial layout and geometric constraints for scene understanding"]}]}],"canonical_facts":{"dc:contributor":["Forsyth, David A.","Hoiem, Derek W.","Ma, Yi","Huang, Thomas S.","Hasegawa-Johnson, Mark A."],"dc:creator":["Hedau, Varsha"],"dc:date":["2012-02-06T20:15:35Z","2011-12"],"dc:description":["An image is nothing but a projection of the physical world around us, where objects do not occur randomly but follow certain spatial rules. Many existing computer vision approaches tend to ignore this aspect of understanding images. In this work, we build representations and propose strategies for exploiting such constraints towards extracting a 3D understanding of a scene from its single image. We model a scene in terms of its spatial layout abstracted as a box, object cuboids, camera viewpoint, and interactions between them. We take a supervised approach towards estimation, and learn models from training data that is fully annotated with the 3D spatial extent of objects, walls, and floor. We assume the world is populated with axis aligned objects and surfaces, and exploit constrained appearance models which use geometric cues from the scene. Our methods are tailored towards indoor scenes that are highly structured and require careful spatial reasoning. We show that our box layout representation is able to capture the full spatial extent of a 3D scene, which we can successfully estimate even for heavily cluttered rooms. Similarly, by exploiting the geometric constraints offered by the scene, we can approximate the extent of the objects as cuboids in 3D. The box layout provides rich contextual information for detecting objects. We show that modeling the 3D interactions between object cuboids and scene layout improves object detection. Finally, we show how to use our 3D spatial layout models together with object cuboid models to predict the free space in the scene.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2011-11-22T21:02:43Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Hedau_Varsha.pdf: 8407684 bytes, checksum: 1255b10542309447c35bd2c19c70fd09 (MD5)","Made available in DSpace on 2012-02-06T20:15:35Z (GMT). No. of bitstreams: 2 Hedau_Varsha.pdf: 8408202 bytes, checksum: 2729be9a90291d2763f1efd4d0aa627b (MD5) license.txt: 4061 bytes, checksum: 631f1b9217a18917c749d33b59f24405 (MD5)"],"dc:identifier":["http://hdl.handle.net/2142/29773"],"dc:language":["en"],"dc:rights":["copyright 2011 Varsha Chandrashekhar Hedau"],"dc:subject":["Computer vision","Image Processing","Scene understanding","Object recognition","Single view 3D modeling","Vanishing points"],"dc:title":["3D spatial layout and geometric constraints for scene understanding"],"dc:type":["Dissertation / Thesis","text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:29Z"}