{"id":{"repo_id":"mit","oai_identifier":"oai:dspace.mit.edu:1721.1/128089"},"canonical_url":"https://search.dev.ndltd.org/etd/mit/oai:dspace.mit.edu:1721.1/128089","repository":{"repo_id":"mit","name":"MIT","base_url":"https://dspace.mit.edu/oai/request"},"display":{"title":"Deep learning-based approaches for depth and 6-DoF pose estimation","abstract":"In this thesis, we investigated two important geometric vision problems, namely, depth estimation from a single RGB image, and 6-DoF object pose estimation from a partial point cloud. Geometric vision problems are concerned with extracting information (e.g. depth, agent trajectory, 3D structure, 6-DoF pose of objects) of the scene from noisy sensor data (e.g. RGB images, LiDAR) by exploiting geometric constraints (e.g. epipolar constraint, rigid motion of objects). Deep learning framework has achieved impressive progress in many computer vision tasks such as image recognition and segmentation. However, applying deep learning-based approaches to geometric vision problems, which are particularly important in safety-critical robotics applications, remains an open problem. The main challenge lies in the fact that it is not straightforward to incorporate geometric constraints, arising from image formation process and physical properties, to optimization problems. To this end, we explore possibilities of enforcing such constraints either by decomposing a problem into two sub-problems each respecting desired constraints, or designing an estimator establishing relationship between intermediate representations and predicted outputs. We propose a deep learning-based approach for -each problem. Through extensive experiments, we show that our proposed approaches produce results comparable with state of the art on public datasets.","abstract_html":"In this thesis, we investigated two important geometric vision problems, namely, depth estimation from a single RGB image, and 6-DoF object pose estimation from a partial point cloud. Geometric vision problems are concerned with extracting information (e.g. depth, agent trajectory, 3D structure, 6-DoF pose of objects) of the scene from noisy sensor data (e.g. RGB images, LiDAR) by exploiting geometric constraints (e.g. epipolar constraint, rigid motion of objects). Deep learning framework has achieved impressive progress in many computer vision tasks such as image recognition and segmentation. However, applying deep learning-based approaches to geometric vision problems, which are particularly important in safety-critical robotics applications, remains an open problem. The main challenge lies in the fact that it is not straightforward to incorporate geometric constraints, arising from image formation process and physical properties, to optimization problems. To this end, we explore possibilities of enforcing such constraints either by decomposing a problem into two sub-problems each respecting desired constraints, or designing an estimator establishing relationship between intermediate representations and predicted outputs. We propose a deep learning-based approach for -each problem. Through extensive experiments, we show that our proposed approaches produce results comparable with state of the art on public datasets.","abstract_has_math":false,"creators":["Lin, Muyuan(Scientist in mechanical engineering)Massachusetts Institute of Technology."],"institution":"Massachusetts Institute of Technology","degree_name":"Master","degree_level":null,"degree_discipline":null,"degree_department":"Massachusetts Institute of Technology. Department of Mechanical Engineering","school":null,"contributors":[],"advisors":["Sertac Karaman."],"committee_chairs":[],"committee_members":[],"year":2020,"date_issued":"2020","date_published":"2020","updated_at":"2026-07-22T22:22:26Z","subjects":["Mechanical Engineering."],"languages":["eng"],"rights":["MIT theses may be protected by copyright. Please reuse MIT thesis content according to the MIT Libraries Permissions Policy, which is available through the URL provided."],"rights_urls":["http://dspace.mit.edu/handle/1721.1/7582"],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/1721.1/128089","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Sertac Karaman."]},{"key":"dc:contributor.department","label":"Department","values":["Massachusetts Institute of Technology. Department of Mechanical Engineering","MechE"]},{"key":"dc:contributor.other","label":"Dc Contributor Other","values":["Massachusetts Institute of Technology. Department of Mechanical Engineering."]},{"key":"dc:creator","label":"Author","values":["Lin, Muyuan(Scientist in mechanical engineering)Massachusetts Institute of Technology."]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2020-10-19T00:42:27Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2020-10-19T00:42:27Z"]},{"key":"dc:date.issued","label":"Date","values":["2020"]},{"key":"dc:publisher","label":"Institution","values":["Massachusetts Institute of Technology"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Mechanical Engineering."]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["MIT theses may be protected by copyright. Please reuse MIT thesis content according to the MIT Libraries Permissions Policy, which is available through the URL provided."]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://dspace.mit.edu/handle/1721.1/7582"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/1721.1/128089"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Thesis: S.M., Massachusetts Institute of Technology, Department of Mechanical Engineering, 2020","Cataloged from PDF of thesis.","Includes bibliographical references (pages 67-79)."]},{"key":"dc:description.abstract","label":"Abstract","values":["In this thesis, we investigated two important geometric vision problems, namely, depth estimation from a single RGB image, and 6-DoF object pose estimation from a partial point cloud. Geometric vision problems are concerned with extracting information (e.g. depth, agent trajectory, 3D structure, 6-DoF pose of objects) of the scene from noisy sensor data (e.g. RGB images, LiDAR) by exploiting geometric constraints (e.g. epipolar constraint, rigid motion of objects). Deep learning framework has achieved impressive progress in many computer vision tasks such as image recognition and segmentation. However, applying deep learning-based approaches to geometric vision problems, which are particularly important in safety-critical robotics applications, remains an open problem. The main challenge lies in the fact that it is not straightforward to incorporate geometric constraints, arising from image formation process and physical properties, to optimization problems. To this end, we explore possibilities of enforcing such constraints either by decomposing a problem into two sub-problems each respecting desired constraints, or designing an estimator establishing relationship between intermediate representations and predicted outputs. We propose a deep learning-based approach for -each problem. Through extensive experiments, we show that our proposed approaches produce results comparable with state of the art on public datasets."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["S.M."]},{"key":"dc:title","label":"Title","values":["Deep learning-based approaches for depth and 6-DoF pose estimation"]}]}],"canonical_facts":{"dc:contributor.advisor":["Sertac Karaman."],"dc:contributor.department":["Massachusetts Institute of Technology. Department of Mechanical Engineering","MechE"],"dc:contributor.other":["Massachusetts Institute of Technology. Department of Mechanical Engineering."],"dc:creator":["Lin, Muyuan(Scientist in mechanical engineering)Massachusetts Institute of Technology."],"dc:date.accessioned":["2020-10-19T00:42:27Z"],"dc:date.available":["2020-10-19T00:42:27Z"],"dc:date.issued":["2020"],"dc:description":["Thesis: S.M., Massachusetts Institute of Technology, Department of Mechanical Engineering, 2020","Cataloged from PDF of thesis.","Includes bibliographical references (pages 67-79)."],"dc:description.abstract":["In this thesis, we investigated two important geometric vision problems, namely, depth estimation from a single RGB image, and 6-DoF object pose estimation from a partial point cloud. Geometric vision problems are concerned with extracting information (e.g. depth, agent trajectory, 3D structure, 6-DoF pose of objects) of the scene from noisy sensor data (e.g. RGB images, LiDAR) by exploiting geometric constraints (e.g. epipolar constraint, rigid motion of objects). Deep learning framework has achieved impressive progress in many computer vision tasks such as image recognition and segmentation. However, applying deep learning-based approaches to geometric vision problems, which are particularly important in safety-critical robotics applications, remains an open problem. The main challenge lies in the fact that it is not straightforward to incorporate geometric constraints, arising from image formation process and physical properties, to optimization problems. To this end, we explore possibilities of enforcing such constraints either by decomposing a problem into two sub-problems each respecting desired constraints, or designing an estimator establishing relationship between intermediate representations and predicted outputs. We propose a deep learning-based approach for -each problem. Through extensive experiments, we show that our proposed approaches produce results comparable with state of the art on public datasets."],"dc:description.degree":["S.M."],"dc:identifier.uri":["https://hdl.handle.net/1721.1/128089"],"dc:language.iso":["eng"],"dc:publisher":["Massachusetts Institute of Technology"],"dc:rights":["MIT theses may be protected by copyright. Please reuse MIT thesis content according to the MIT Libraries Permissions Policy, which is available through the URL provided."],"dc:rights.uri":["http://dspace.mit.edu/handle/1721.1/7582"],"dc:subject":["Mechanical Engineering."],"dc:title":["Deep learning-based approaches for depth and 6-DoF pose estimation"],"dc:type":["Thesis"],"thesis:degree_name":["Master"]},"updated_at":"2026-07-22T22:22:26Z"}