National University of Singapore
COMPOSITIONAL OBJECT-CENTRIC REPRESENTATIONS FOR ROBUST VISUAL PERCEPTION
Abstract
dc:description.abstractDespite significant advancements, real-world deployment of modern vision models in critical applications remains limited by poor out-of-distribution generalization, failures under occlusion, reliance on large high-quality datasets, and limited interpretability. We posit that these limitations arise from fundamental differences between modern vision models and the human visual cortex. While current models mimic hierarchical processing, they treat images as holistic inputs. In contrast, human vision learns objects as compositional building blocks, enabling robust recognition, occlusion reasoning, generalization, and sample-efficient learning. To address this gap, this thesis proposes a unified object-centric representation framework with three stages: discovery, representation, and application. The proposed models improve robustness, sample efficiency, occlusion awareness, and generalization. OC-Net addresses limitations of pixel-based reconstruction using a feature connectivity algorithm and object-centric regularization optimized for downstream tasks. ECO-Net introduces graph-based representations of object parts, a co-part discovery algorithm, and a memory module, enabling improved performance on multi-part and occluded objects. Finally, DP-GAT integrates these ideas by combining OC-Net’s unsupervised object discovery with ECO-Net’s relational modeling, achieving strong results in disease progression prediction and 3D tumor segmentation. Together, these contributions support more reliable deployment of vision models in real-world applications.
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- ALEX FOO DA WENG