Back to results

George Mason University

MULTI-MODAL SCENE UNDERSTANDING

Abstract

Over the past few years, advancements in scene understanding have played a crucial rolein enabling autonomous systems to better perceive and interact with their surroundings. These improvements have been driven by progress in deep learning, the creation of both real and synthetic datasets, and increased computational power. However, deploying these models in novel, real-world environments remains challenging due to visual domain gaps, including varying lighting conditions, differences in object scale, appearance, and orientation, as well as occlusion, and clutter. This dissertation proposes enhancing existing models by exploiting geometric cues and constraints for learning and optimization. The developed methods demonstrate that leveraging geometric cues can lead to more data-efficient models, reduce the need for large amounts of training data, and enable strong performance in diverse environments. First, we tackle the problem of 6D object pose estimation by incorporating geometric features, along with a refinement stage to improve accuracy in data-scarce scenarios. Second, we propose an end-to-end multi-view panoramic global pose estimation framework using a graph neural network (GNN), which leverages geometry-aware objective functions for more accurate and efficient localization. Third, we develop a structured approach that utilizes the 3D positions, scales, and poses of objects to reason about spatial relations with high accuracy, resulting in improved spatial relation detection. Finally, we propose a zero-shot3D semantic segmentation framework by fusing multi-view 2D predictions. This novel fusion strategy leverages 2D semantic segmentation with associated uncertainties, depth measurements, and camera poses to enhance adaptability to novel environments without requiring any training or data labeling in 3D. Together, these contributions advance the state of the art in scene understanding and support the design and development of more robust and adaptable embodied AI systems capable of operating in challenging real-world environments.

Author and committee

dc:creator, dc:contributor.*
Author
  • Nejatishahidin, Negar

Subjects

dc:subject × 6

Identifiers

dc:identifier.*
Identifier
hdl:1920/14741
OAI identifier oai:identifier
oai:MARS:1920/14741

Chain of custody

source
Harvested from
George Mason University
Base URL
mars.gmu.edu/server/oai/request
Last updated
2026-07-27
Source record
OAI-PMH GetRecord
citation

Nejatishahidin, Negar. MULTI-MODAL SCENE UNDERSTANDING. 2025.