{"id":{"repo_id":"mit","oai_identifier":"oai:dspace.mit.edu:1721.1/79237"},"canonical_url":"https://search.dev.ndltd.org/etd/mit/oai:dspace.mit.edu:1721.1/79237","repository":{"repo_id":"mit","name":"MIT","base_url":"https://dspace.mit.edu/oai/request"},"display":{"title":"Combining recognition and geometry for data-driven 3D reconstruction","abstract":"Today's multi-view 3D reconstruction techniques rely almost exclusively on depth cues that come from multiple view geometry. While these cues can be used to produce highly accurate reconstructions, the resulting point clouds are often noisy and incomplete. Due to these issues, it may also be difficult to answer higher-level questions about the geometry, such as whether two surfaces meet at a right angle or whether a surface is planar. Furthermore, state-of-the-art reconstruction techniques generally cannot learn from training data, so having the ground-truth geometry for one scene does not aid in reconstructing similar scenes. In this work, we make two contributions toward data-driven 3D reconstruction. First, we present a dataset containing hundreds of RGBD videos that can be used as a source of training data for reconstruction algorithms. Second, we introduce the concept of the Shape Anchor, a region for which the combination of recognition and multiple view geometry allows us to accurately predict the latent, dense point cloud. We propose a technique to detect these regions and to predict their shapes, and we demonstrate it on our dataset.","abstract_html":"Today&#x27;s multi-view 3D reconstruction techniques rely almost exclusively on depth cues that come from multiple view geometry. While these cues can be used to produce highly accurate reconstructions, the resulting point clouds are often noisy and incomplete. Due to these issues, it may also be difficult to answer higher-level questions about the geometry, such as whether two surfaces meet at a right angle or whether a surface is planar. Furthermore, state-of-the-art reconstruction techniques generally cannot learn from training data, so having the ground-truth geometry for one scene does not aid in reconstructing similar scenes. In this work, we make two contributions toward data-driven 3D reconstruction. First, we present a dataset containing hundreds of RGBD videos that can be used as a source of training data for reconstruction algorithms. Second, we introduce the concept of the Shape Anchor, a region for which the combination of recognition and multiple view geometry allows us to accurately predict the latent, dense point cloud. We propose a technique to detect these regions and to predict their shapes, and we demonstrate it on our dataset.","abstract_has_math":false,"creators":["Owens, Andrew (Andrew Hale)"],"institution":"Massachusetts Institute of Technology","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":"Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science.","school":null,"contributors":[],"advisors":["William T. Freeman and Antonio Torralba."],"committee_chairs":[],"committee_members":[],"year":2013,"date_issued":"2013","date_published":"2013","updated_at":"2026-07-22T22:21:03Z","subjects":["Electrical Engineering and Computer Science."],"languages":["eng"],"rights":["M.I.T. theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. See provided URL for inquiries about permission."],"rights_urls":["http://dspace.mit.edu/handle/1721.1/7582"],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/1721.1/79237","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["William T. Freeman and Antonio Torralba."]},{"key":"dc:contributor.department","label":"Department","values":["Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science."]},{"key":"dc:contributor.other","label":"Dc Contributor Other","values":["Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science."]},{"key":"dc:creator","label":"Author","values":["Owens, Andrew (Andrew Hale)"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2013-06-17T19:49:43Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2013-06-17T19:49:43Z"]},{"key":"dc:date.issued","label":"Date","values":["2013"]},{"key":"dc:publisher","label":"Institution","values":["Massachusetts Institute of Technology"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Electrical Engineering and Computer Science."]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["M.I.T. theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. See provided URL for inquiries about permission."]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://dspace.mit.edu/handle/1721.1/7582"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["http://hdl.handle.net/1721.1/79237"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Thesis (S.M.)--Massachusetts Institute of Technology, Dept. of Electrical Engineering and Computer Science, 2013.","Cataloged from PDF version of thesis.","Includes bibliographical references (p. 47-50)."]},{"key":"dc:description.abstract","label":"Abstract","values":["Today's multi-view 3D reconstruction techniques rely almost exclusively on depth cues that come from multiple view geometry. While these cues can be used to produce highly accurate reconstructions, the resulting point clouds are often noisy and incomplete. Due to these issues, it may also be difficult to answer higher-level questions about the geometry, such as whether two surfaces meet at a right angle or whether a surface is planar. Furthermore, state-of-the-art reconstruction techniques generally cannot learn from training data, so having the ground-truth geometry for one scene does not aid in reconstructing similar scenes. In this work, we make two contributions toward data-driven 3D reconstruction. First, we present a dataset containing hundreds of RGBD videos that can be used as a source of training data for reconstruction algorithms. Second, we introduce the concept of the Shape Anchor, a region for which the combination of recognition and multiple view geometry allows us to accurately predict the latent, dense point cloud. We propose a technique to detect these regions and to predict their shapes, and we demonstrate it on our dataset."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["S.M."]},{"key":"dc:title","label":"Title","values":["Combining recognition and geometry for data-driven 3D reconstruction"]}]}],"canonical_facts":{"dc:contributor.advisor":["William T. Freeman and Antonio Torralba."],"dc:contributor.department":["Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science."],"dc:contributor.other":["Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science."],"dc:creator":["Owens, Andrew (Andrew Hale)"],"dc:date.accessioned":["2013-06-17T19:49:43Z"],"dc:date.available":["2013-06-17T19:49:43Z"],"dc:date.issued":["2013"],"dc:description":["Thesis (S.M.)--Massachusetts Institute of Technology, Dept. of Electrical Engineering and Computer Science, 2013.","Cataloged from PDF version of thesis.","Includes bibliographical references (p. 47-50)."],"dc:description.abstract":["Today's multi-view 3D reconstruction techniques rely almost exclusively on depth cues that come from multiple view geometry. While these cues can be used to produce highly accurate reconstructions, the resulting point clouds are often noisy and incomplete. Due to these issues, it may also be difficult to answer higher-level questions about the geometry, such as whether two surfaces meet at a right angle or whether a surface is planar. Furthermore, state-of-the-art reconstruction techniques generally cannot learn from training data, so having the ground-truth geometry for one scene does not aid in reconstructing similar scenes. In this work, we make two contributions toward data-driven 3D reconstruction. First, we present a dataset containing hundreds of RGBD videos that can be used as a source of training data for reconstruction algorithms. Second, we introduce the concept of the Shape Anchor, a region for which the combination of recognition and multiple view geometry allows us to accurately predict the latent, dense point cloud. We propose a technique to detect these regions and to predict their shapes, and we demonstrate it on our dataset."],"dc:description.degree":["S.M."],"dc:identifier.uri":["http://hdl.handle.net/1721.1/79237"],"dc:language.iso":["eng"],"dc:publisher":["Massachusetts Institute of Technology"],"dc:rights":["M.I.T. theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. See provided URL for inquiries about permission."],"dc:rights.uri":["http://dspace.mit.edu/handle/1721.1/7582"],"dc:subject":["Electrical Engineering and Computer Science."],"dc:title":["Combining recognition and geometry for data-driven 3D reconstruction"],"dc:type":["Thesis"]},"updated_at":"2026-07-22T22:21:03Z"}