{"id":{"repo_id":"mit","oai_identifier":"oai:dspace.mit.edu:1721.1/139136"},"canonical_url":"https://search.dev.ndltd.org/etd/mit/oai:dspace.mit.edu:1721.1/139136","repository":{"repo_id":"mit","name":"MIT","base_url":"https://dspace.mit.edu/oai/request"},"display":{"title":"Consistent Depth Estimation in Data-Driven Simulation for Autonomous Driving","abstract":"In this work we propose consistent depth estimation for viewpoint reconstruction in data-driven simulation, combining aspects of learning-based monocular depth prediction and structure-from-motion to increase temporal video depth accuracy. We demonstrate efficacy in VISTA, an end-to-end autonomous vehicle simulation engine capable of training robust control policies directly applicable to the real-world. Taking advantage of geometrically consistent depth map estimations, we see a several order of magnitude improvement in whole-frame depth accuracy averaged over the course of input traces compared to VISTA’s current depth method, and a 39% reduction in intra-frame depth variance compared to current state of the art methods (i.e. Monodepth2) while maintaining similar error. Better depth enables more accurate viewpoint reconstruction thus improving the training of reinforcement learning (RL) control policies in simulation, increasing RL-based control’s practicality. We train several end-to-end policy gradient models in varying versions of VISTA, each utilizing a different depth method, and see that end-to-end models trained in the consistent depth version of VISTA deviate least from the human driven center line.","abstract_html":"In this work we propose consistent depth estimation for viewpoint reconstruction in data-driven simulation, combining aspects of learning-based monocular depth prediction and structure-from-motion to increase temporal video depth accuracy. We demonstrate efficacy in VISTA, an end-to-end autonomous vehicle simulation engine capable of training robust control policies directly applicable to the real-world. Taking advantage of geometrically consistent depth map estimations, we see a several order of magnitude improvement in whole-frame depth accuracy averaged over the course of input traces compared to VISTA’s current depth method, and a 39% reduction in intra-frame depth variance compared to current state of the art methods (i.e. Monodepth2) while maintaining similar error. Better depth enables more accurate viewpoint reconstruction thus improving the training of reinforcement learning (RL) control policies in simulation, increasing RL-based control’s practicality. We train several end-to-end policy gradient models in varying versions of VISTA, each utilizing a different depth method, and see that end-to-end models trained in the consistent depth version of VISTA deviate least from the human driven center line.","abstract_has_math":false,"creators":["Beveridge, Matthew"],"institution":"Massachusetts Institute of Technology","degree_name":"Master","degree_level":null,"degree_discipline":null,"degree_department":"Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science","school":null,"contributors":[],"advisors":["Rus, Daniela"],"committee_chairs":[],"committee_members":[],"year":2021,"date_issued":"2021-06","date_published":"2021-06","updated_at":"2026-07-22T22:22:14Z","subjects":[],"languages":[],"rights":["In Copyright - Educational Use Permitted","Copyright MIT"],"rights_urls":["http://rightsstatements.org/page/InC-EDU/1.0/"],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/1721.1/139136","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Rus, Daniela"]},{"key":"dc:contributor.department","label":"Department","values":["Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science"]},{"key":"dc:creator","label":"Author","values":["Beveridge, Matthew"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2022-01-14T14:52:03Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2022-01-14T14:52:03Z"]},{"key":"dc:date.issued","label":"Date","values":["2021-06"]},{"key":"dc:publisher","label":"Institution","values":["Massachusetts Institute of Technology"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master","Master of Engineering in Electrical Engineering and Computer Science"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["In Copyright - Educational Use Permitted","Copyright MIT"]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://rightsstatements.org/page/InC-EDU/1.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/1721.1/139136"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["In this work we propose consistent depth estimation for viewpoint reconstruction in data-driven simulation, combining aspects of learning-based monocular depth prediction and structure-from-motion to increase temporal video depth accuracy. We demonstrate efficacy in VISTA, an end-to-end autonomous vehicle simulation engine capable of training robust control policies directly applicable to the real-world. Taking advantage of geometrically consistent depth map estimations, we see a several order of magnitude improvement in whole-frame depth accuracy averaged over the course of input traces compared to VISTA’s current depth method, and a 39% reduction in intra-frame depth variance compared to current state of the art methods (i.e. Monodepth2) while maintaining similar error. Better depth enables more accurate viewpoint reconstruction thus improving the training of reinforcement learning (RL) control policies in simulation, increasing RL-based control’s practicality. We train several end-to-end policy gradient models in varying versions of VISTA, each utilizing a different depth method, and see that end-to-end models trained in the consistent depth version of VISTA deviate least from the human driven center line."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["M.Eng."]},{"key":"dc:title","label":"Title","values":["Consistent Depth Estimation in Data-Driven Simulation for Autonomous Driving"]}]}],"canonical_facts":{"dc:contributor.advisor":["Rus, Daniela"],"dc:contributor.department":["Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science"],"dc:creator":["Beveridge, Matthew"],"dc:date.accessioned":["2022-01-14T14:52:03Z"],"dc:date.available":["2022-01-14T14:52:03Z"],"dc:date.issued":["2021-06"],"dc:description.abstract":["In this work we propose consistent depth estimation for viewpoint reconstruction in data-driven simulation, combining aspects of learning-based monocular depth prediction and structure-from-motion to increase temporal video depth accuracy. We demonstrate efficacy in VISTA, an end-to-end autonomous vehicle simulation engine capable of training robust control policies directly applicable to the real-world. Taking advantage of geometrically consistent depth map estimations, we see a several order of magnitude improvement in whole-frame depth accuracy averaged over the course of input traces compared to VISTA’s current depth method, and a 39% reduction in intra-frame depth variance compared to current state of the art methods (i.e. Monodepth2) while maintaining similar error. Better depth enables more accurate viewpoint reconstruction thus improving the training of reinforcement learning (RL) control policies in simulation, increasing RL-based control’s practicality. We train several end-to-end policy gradient models in varying versions of VISTA, each utilizing a different depth method, and see that end-to-end models trained in the consistent depth version of VISTA deviate least from the human driven center line."],"dc:description.degree":["M.Eng."],"dc:identifier.uri":["https://hdl.handle.net/1721.1/139136"],"dc:publisher":["Massachusetts Institute of Technology"],"dc:rights":["In Copyright - Educational Use Permitted","Copyright MIT"],"dc:rights.uri":["http://rightsstatements.org/page/InC-EDU/1.0/"],"dc:title":["Consistent Depth Estimation in Data-Driven Simulation for Autonomous Driving"],"dc:type":["Thesis"],"thesis:degree_name":["Master","Master of Engineering in Electrical Engineering and Computer Science"]},"updated_at":"2026-07-22T22:22:14Z"}