{"id":{"repo_id":"mit","oai_identifier":"oai:dspace.mit.edu:1721.1/147476"},"canonical_url":"https://search.dev.ndltd.org/etd/mit/oai:dspace.mit.edu:1721.1/147476","repository":{"repo_id":"mit","name":"MIT","base_url":"https://dspace.mit.edu/oai/request"},"display":{"title":"Diverse Behavior Prediction through Deep Hybrid Models","abstract":"Predicting future motions of agents is a crucial task for autonomous vehicles. This task is challenging due to multi-modal traffic agent behaviors such as maneuvers. In addition, predictors are often required to generate a limited number of prediction samples to cover diverse behaviors, due to the time complexity of processing these samples for downstream tasks. Existing model-based prediction methods leverage hybrid reasoning techniques to predict qualitatively representative agent motions from a large prediction space, yet they often assume simple agent dynamics and fail to account for scene context. Recently, learning-based approaches have demonstrated great success in learning complicated agent dynamics and scene context through deep neural networks to produce accurate trajectories in complex traffic scenes. On the other hand, they often leverage a black box deep neural network and fail the explore the structure of the problem. In this thesis, we propose deep hybrid models by unifying the power of model-based hybrid reasoning algorithms and learning-based models to predict a small set of accurate trajectory samples that cover qualitatively representative agent maneuvers. Our approach offers several advantages compared to existing model-based and learning-based predictors. First, it handles evolving intent over time by learning an accurate hybrid model representing evolving discrete maneuvers and continuous trajectories, and sampling a set of hybrid trajectory sequences through importance sampling based on learned proposal distributions. Second, it handles ambiguous maneuvers by learning a latent space of qualitative maneuvers that mimics human concepts of qualitatively representative maneuvers. Third, it generates samples that support multiple downstream tasks, including autonomous planning and driver warning, by adding a task-informed loss that leverages the specification of the task when the additional task information is given. We train and validate our models on large-scale public driving benchmarks, including Argoverse forecasting dataset and Waymo open motion dataset. We perform extensive qualitative and quantitative experimental results to demonstrate the advantage of our predictor over state-of-the-art model-based and learning-based baselines, in terms of accuracy, diversity, and task performance.","abstract_html":"Predicting future motions of agents is a crucial task for autonomous vehicles. This task is challenging due to multi-modal traffic agent behaviors such as maneuvers. In addition, predictors are often required to generate a limited number of prediction samples to cover diverse behaviors, due to the time complexity of processing these samples for downstream tasks. Existing model-based prediction methods leverage hybrid reasoning techniques to predict qualitatively representative agent motions from a large prediction space, yet they often assume simple agent dynamics and fail to account for scene context. Recently, learning-based approaches have demonstrated great success in learning complicated agent dynamics and scene context through deep neural networks to produce accurate trajectories in complex traffic scenes. On the other hand, they often leverage a black box deep neural network and fail the explore the structure of the problem. In this thesis, we propose deep hybrid models by unifying the power of model-based hybrid reasoning algorithms and learning-based models to predict a small set of accurate trajectory samples that cover qualitatively representative agent maneuvers. Our approach offers several advantages compared to existing model-based and learning-based predictors. First, it handles evolving intent over time by learning an accurate hybrid model representing evolving discrete maneuvers and continuous trajectories, and sampling a set of hybrid trajectory sequences through importance sampling based on learned proposal distributions. Second, it handles ambiguous maneuvers by learning a latent space of qualitative maneuvers that mimics human concepts of qualitatively representative maneuvers. Third, it generates samples that support multiple downstream tasks, including autonomous planning and driver warning, by adding a task-informed loss that leverages the specification of the task when the additional task information is given. We train and validate our models on large-scale public driving benchmarks, including Argoverse forecasting dataset and Waymo open motion dataset. We perform extensive qualitative and quantitative experimental results to demonstrate the advantage of our predictor over state-of-the-art model-based and learning-based baselines, in terms of accuracy, diversity, and task performance.","abstract_has_math":false,"creators":["Huang, Xin"],"institution":"Massachusetts Institute of Technology","degree_name":"Doctoral","degree_level":null,"degree_discipline":null,"degree_department":"Massachusetts Institute of Technology. Department of Aeronautics and Astronautics","school":null,"contributors":[],"advisors":["Williams, Brian C."],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-09","date_published":"2022-09","updated_at":"2026-07-22T22:22:22Z","subjects":[],"languages":[],"rights":["In Copyright - Educational Use Permitted","Copyright MIT"],"rights_urls":["http://rightsstatements.org/page/InC-EDU/1.0/"],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/1721.1/147476","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Williams, Brian C."]},{"key":"dc:contributor.department","label":"Department","values":["Massachusetts Institute of Technology. Department of Aeronautics and Astronautics"]},{"key":"dc:creator","label":"Author","values":["Huang, Xin"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2023-01-19T19:52:58Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2023-01-19T19:52:58Z"]},{"key":"dc:date.issued","label":"Date","values":["2022-09"]},{"key":"dc:publisher","label":"Institution","values":["Massachusetts Institute of Technology"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctoral","Doctor of Philosophy"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["In Copyright - Educational Use Permitted","Copyright MIT"]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://rightsstatements.org/page/InC-EDU/1.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/1721.1/147476"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Predicting future motions of agents is a crucial task for autonomous vehicles. This task is challenging due to multi-modal traffic agent behaviors such as maneuvers. In addition, predictors are often required to generate a limited number of prediction samples to cover diverse behaviors, due to the time complexity of processing these samples for downstream tasks. Existing model-based prediction methods leverage hybrid reasoning techniques to predict qualitatively representative agent motions from a large prediction space, yet they often assume simple agent dynamics and fail to account for scene context. Recently, learning-based approaches have demonstrated great success in learning complicated agent dynamics and scene context through deep neural networks to produce accurate trajectories in complex traffic scenes. On the other hand, they often leverage a black box deep neural network and fail the explore the structure of the problem. In this thesis, we propose deep hybrid models by unifying the power of model-based hybrid reasoning algorithms and learning-based models to predict a small set of accurate trajectory samples that cover qualitatively representative agent maneuvers. Our approach offers several advantages compared to existing model-based and learning-based predictors. First, it handles evolving intent over time by learning an accurate hybrid model representing evolving discrete maneuvers and continuous trajectories, and sampling a set of hybrid trajectory sequences through importance sampling based on learned proposal distributions. Second, it handles ambiguous maneuvers by learning a latent space of qualitative maneuvers that mimics human concepts of qualitatively representative maneuvers. Third, it generates samples that support multiple downstream tasks, including autonomous planning and driver warning, by adding a task-informed loss that leverages the specification of the task when the additional task information is given. We train and validate our models on large-scale public driving benchmarks, including Argoverse forecasting dataset and Waymo open motion dataset. We perform extensive qualitative and quantitative experimental results to demonstrate the advantage of our predictor over state-of-the-art model-based and learning-based baselines, in terms of accuracy, diversity, and task performance."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["Ph.D."]},{"key":"dc:title","label":"Title","values":["Diverse Behavior Prediction through Deep Hybrid Models"]}]}],"canonical_facts":{"dc:contributor.advisor":["Williams, Brian C."],"dc:contributor.department":["Massachusetts Institute of Technology. Department of Aeronautics and Astronautics"],"dc:creator":["Huang, Xin"],"dc:date.accessioned":["2023-01-19T19:52:58Z"],"dc:date.available":["2023-01-19T19:52:58Z"],"dc:date.issued":["2022-09"],"dc:description.abstract":["Predicting future motions of agents is a crucial task for autonomous vehicles. This task is challenging due to multi-modal traffic agent behaviors such as maneuvers. In addition, predictors are often required to generate a limited number of prediction samples to cover diverse behaviors, due to the time complexity of processing these samples for downstream tasks. Existing model-based prediction methods leverage hybrid reasoning techniques to predict qualitatively representative agent motions from a large prediction space, yet they often assume simple agent dynamics and fail to account for scene context. Recently, learning-based approaches have demonstrated great success in learning complicated agent dynamics and scene context through deep neural networks to produce accurate trajectories in complex traffic scenes. On the other hand, they often leverage a black box deep neural network and fail the explore the structure of the problem. In this thesis, we propose deep hybrid models by unifying the power of model-based hybrid reasoning algorithms and learning-based models to predict a small set of accurate trajectory samples that cover qualitatively representative agent maneuvers. Our approach offers several advantages compared to existing model-based and learning-based predictors. First, it handles evolving intent over time by learning an accurate hybrid model representing evolving discrete maneuvers and continuous trajectories, and sampling a set of hybrid trajectory sequences through importance sampling based on learned proposal distributions. Second, it handles ambiguous maneuvers by learning a latent space of qualitative maneuvers that mimics human concepts of qualitatively representative maneuvers. Third, it generates samples that support multiple downstream tasks, including autonomous planning and driver warning, by adding a task-informed loss that leverages the specification of the task when the additional task information is given. We train and validate our models on large-scale public driving benchmarks, including Argoverse forecasting dataset and Waymo open motion dataset. We perform extensive qualitative and quantitative experimental results to demonstrate the advantage of our predictor over state-of-the-art model-based and learning-based baselines, in terms of accuracy, diversity, and task performance."],"dc:description.degree":["Ph.D."],"dc:identifier.uri":["https://hdl.handle.net/1721.1/147476"],"dc:publisher":["Massachusetts Institute of Technology"],"dc:rights":["In Copyright - Educational Use Permitted","Copyright MIT"],"dc:rights.uri":["http://rightsstatements.org/page/InC-EDU/1.0/"],"dc:title":["Diverse Behavior Prediction through Deep Hybrid Models"],"dc:type":["Thesis"],"thesis:degree_name":["Doctoral","Doctor of Philosophy"]},"updated_at":"2026-07-22T22:22:22Z"}