{"id":{"repo_id":"gmu","oai_identifier":"oai:MARS:1920/14741"},"canonical_url":"https://search.dev.ndltd.org/etd/gmu/oai:MARS:1920/14741","repository":{"repo_id":"gmu","name":"George Mason University","base_url":"https://mars.gmu.edu/server/oai/request"},"display":{"title":"MULTI-MODAL SCENE UNDERSTANDING","abstract":"Over the past few years, advancements in scene understanding have played a crucial rolein enabling autonomous systems to better perceive and interact with their surroundings. These improvements have been driven by progress in deep learning, the creation of both real and synthetic datasets, and increased computational power. However, deploying these models in novel, real-world environments remains challenging due to visual domain gaps, including varying lighting conditions, differences in object scale, appearance, and orientation, as well as occlusion, and clutter. This dissertation proposes enhancing existing models by exploiting geometric cues and constraints for learning and optimization. The developed methods demonstrate that leveraging geometric cues can lead to more data-efficient models, reduce the need for large amounts of training data, and enable strong performance in diverse environments. First, we tackle the problem of 6D object pose estimation by incorporating geometric features, along with a refinement stage to improve accuracy in data-scarce scenarios. Second, we propose an end-to-end multi-view panoramic global pose estimation framework using a graph neural network (GNN), which leverages geometry-aware objective functions for more accurate and efficient localization. Third, we develop a structured approach that utilizes the 3D positions, scales, and poses of objects to reason about spatial relations with high accuracy, resulting in improved spatial relation detection. Finally, we propose a zero-shot3D semantic segmentation framework by fusing multi-view 2D predictions. This novel fusion strategy leverages 2D semantic segmentation with associated uncertainties, depth measurements, and camera poses to enhance adaptability to novel environments without requiring any training or data labeling in 3D. Together, these contributions advance the state of the art in scene understanding and support the design and development of more robust and adaptable embodied AI systems capable of operating in challenging real-world environments.","abstract_html":"Over the past few years, advancements in scene understanding have played a crucial rolein enabling autonomous systems to better perceive and interact with their surroundings. These improvements have been driven by progress in deep learning, the creation of both real and synthetic datasets, and increased computational power. However, deploying these models in novel, real-world environments remains challenging due to visual domain gaps, including varying lighting conditions, differences in object scale, appearance, and orientation, as well as occlusion, and clutter. This dissertation proposes enhancing existing models by exploiting geometric cues and constraints for learning and optimization. The developed methods demonstrate that leveraging geometric cues can lead to more data-efficient models, reduce the need for large amounts of training data, and enable strong performance in diverse environments. First, we tackle the problem of 6D object pose estimation by incorporating geometric features, along with a refinement stage to improve accuracy in data-scarce scenarios. Second, we propose an end-to-end multi-view panoramic global pose estimation framework using a graph neural network (GNN), which leverages geometry-aware objective functions for more accurate and efficient localization. Third, we develop a structured approach that utilizes the 3D positions, scales, and poses of objects to reason about spatial relations with high accuracy, resulting in improved spatial relation detection. Finally, we propose a zero-shot3D semantic segmentation framework by fusing multi-view 2D predictions. This novel fusion strategy leverages 2D semantic segmentation with associated uncertainties, depth measurements, and camera poses to enhance adaptability to novel environments without requiring any training or data labeling in 3D. Together, these contributions advance the state of the art in scene understanding and support the design and development of more robust and adaptable embodied AI systems capable of operating in challenging real-world environments.","abstract_has_math":false,"creators":["Nejatishahidin, Negar"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025","date_published":"2025","updated_at":"2026-07-27T19:51:46Z","subjects":["Camera Pose Estimation","Computer Vision","Object Pose Estimation","Scene Underestanding","Semantic Segmentation","Spatial Reasoning"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["hdl:1920/14741"],"render_values":[{"text":"hdl:1920/14741","href":null,"code":true}]}]},"links":{"outbound_url":null,"outbound_label":null,"outbound_source":null},"metadata_groups":[{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2025"]},{"key":"dc:type","label":"Dc Type","values":["Dissertation"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Camera Pose Estimation","Computer Vision","Object Pose Estimation","Scene Underestanding","Semantic Segmentation","Spatial Reasoning"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["hdl:1920/14741"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.other","label":"Dc Description Other","values":["Over the past few years, advancements in scene understanding have played a crucial rolein enabling autonomous systems to better perceive and interact with their surroundings. These improvements have been driven by progress in deep learning, the creation of both real and synthetic datasets, and increased computational power. However, deploying these models in novel, real-world environments remains challenging due to visual domain gaps, including varying lighting conditions, differences in object scale, appearance, and orientation, as well as occlusion, and clutter. This dissertation proposes enhancing existing models by exploiting geometric cues and constraints for learning and optimization. The developed methods demonstrate that leveraging geometric cues can lead to more data-efficient models, reduce the need for large amounts of training data, and enable strong performance in diverse environments. First, we tackle the problem of 6D object pose estimation by incorporating geometric features, along with a refinement stage to improve accuracy in data-scarce scenarios. Second, we propose an end-to-end multi-view panoramic global pose estimation framework using a graph neural network (GNN), which leverages geometry-aware objective functions for more accurate and efficient localization. Third, we develop a structured approach that utilizes the 3D positions, scales, and poses of objects to reason about spatial relations with high accuracy, resulting in improved spatial relation detection. Finally, we propose a zero-shot3D semantic segmentation framework by fusing multi-view 2D predictions. This novel fusion strategy leverages 2D semantic segmentation with associated uncertainties, depth measurements, and camera poses to enhance adaptability to novel environments without requiring any training or data labeling in 3D. Together, these contributions advance the state of the art in scene understanding and support the design and development of more robust and adaptable embodied AI systems capable of operating in challenging real-world environments."]},{"key":"dc:title","label":"Title","values":["MULTI-MODAL SCENE UNDERSTANDING"]}]}],"canonical_facts":{"dc:date.issued":["2025"],"dc:description.other":["Over the past few years, advancements in scene understanding have played a crucial rolein enabling autonomous systems to better perceive and interact with their surroundings. These improvements have been driven by progress in deep learning, the creation of both real and synthetic datasets, and increased computational power. However, deploying these models in novel, real-world environments remains challenging due to visual domain gaps, including varying lighting conditions, differences in object scale, appearance, and orientation, as well as occlusion, and clutter. This dissertation proposes enhancing existing models by exploiting geometric cues and constraints for learning and optimization. The developed methods demonstrate that leveraging geometric cues can lead to more data-efficient models, reduce the need for large amounts of training data, and enable strong performance in diverse environments. First, we tackle the problem of 6D object pose estimation by incorporating geometric features, along with a refinement stage to improve accuracy in data-scarce scenarios. Second, we propose an end-to-end multi-view panoramic global pose estimation framework using a graph neural network (GNN), which leverages geometry-aware objective functions for more accurate and efficient localization. Third, we develop a structured approach that utilizes the 3D positions, scales, and poses of objects to reason about spatial relations with high accuracy, resulting in improved spatial relation detection. Finally, we propose a zero-shot3D semantic segmentation framework by fusing multi-view 2D predictions. This novel fusion strategy leverages 2D semantic segmentation with associated uncertainties, depth measurements, and camera poses to enhance adaptability to novel environments without requiring any training or data labeling in 3D. Together, these contributions advance the state of the art in scene understanding and support the design and development of more robust and adaptable embodied AI systems capable of operating in challenging real-world environments."],"dc:identifier":["hdl:1920/14741"],"dc:subject":["Camera Pose Estimation","Computer Vision","Object Pose Estimation","Scene Underestanding","Semantic Segmentation","Spatial Reasoning"],"dc:title":["MULTI-MODAL SCENE UNDERSTANDING"],"dc:type":["Dissertation"]},"updated_at":"2026-07-27T19:51:46Z"}