{"id":{"repo_id":"nus","oai_identifier":"oai:scholarbank.nus.edu.sg:10635/319125"},"canonical_url":"https://search.dev.ndltd.org/etd/nus/oai:scholarbank.nus.edu.sg:10635/319125","repository":{"repo_id":"nus","name":"National University of Singapore","base_url":"https://scholarbank.nus.edu.sg/oai/request"},"display":{"title":"COMPOSITIONAL OBJECT-CENTRIC REPRESENTATIONS FOR ROBUST VISUAL PERCEPTION","abstract":"Despite significant advancements, real-world deployment of modern vision models in critical applications remains limited by poor out-of-distribution generalization, failures under occlusion, reliance on large high-quality datasets, and limited interpretability. We posit that these limitations arise from fundamental differences between modern vision models and the human visual cortex. While current models mimic hierarchical processing, they treat images as holistic inputs. In contrast, human vision learns objects as compositional building blocks, enabling robust recognition, occlusion reasoning, generalization, and sample-efficient learning. To address this gap, this thesis proposes a unified object-centric representation framework with three stages: discovery, representation, and application. The proposed models improve robustness, sample efficiency, occlusion awareness, and generalization. OC-Net addresses limitations of pixel-based reconstruction using a feature connectivity algorithm and object-centric regularization optimized for downstream tasks. ECO-Net introduces graph-based representations of object parts, a co-part discovery algorithm, and a memory module, enabling improved performance on multi-part and occluded objects. Finally, DP-GAT integrates these ideas by combining OC-Net’s unsupervised object discovery with ECO-Net’s relational modeling, achieving strong results in disease progression prediction and 3D tumor segmentation. Together, these contributions support more reliable deployment of vision models in real-world applications.","abstract_html":"Despite significant advancements, real-world deployment of modern vision models in critical applications remains limited by poor out-of-distribution generalization, failures under occlusion, reliance on large high-quality datasets, and limited interpretability. We posit that these limitations arise from fundamental differences between modern vision models and the human visual cortex. While current models mimic hierarchical processing, they treat images as holistic inputs. In contrast, human vision learns objects as compositional building blocks, enabling robust recognition, occlusion reasoning, generalization, and sample-efficient learning. To address this gap, this thesis proposes a unified object-centric representation framework with three stages: discovery, representation, and application. The proposed models improve robustness, sample efficiency, occlusion awareness, and generalization. OC-Net addresses limitations of pixel-based reconstruction using a feature connectivity algorithm and object-centric regularization optimized for downstream tasks. ECO-Net introduces graph-based representations of object parts, a co-part discovery algorithm, and a memory module, enabling improved performance on multi-part and occluded objects. Finally, DP-GAT integrates these ideas by combining OC-Net’s unsupervised object discovery with ECO-Net’s relational modeling, achieving strong results in disease progression prediction and 3D tumor segmentation. Together, these contributions support more reliable deployment of vision models in real-world applications.","abstract_has_math":false,"creators":["ALEX FOO DA WENG"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-12-20","date_published":"2025-12-20","updated_at":"2026-07-24T03:31:26Z","subjects":["Occlusion-Aware Perception","Out-of-Distribution Generalization","Multi-Object Representation Learning","Object-Centric Learning","Vision Models"],"languages":[],"rights":[],"rights_urls":["https://scholarbank.nus.edu.sg/bitstreams/08c12eda-14b2-4bca-b936-6529a3af8be0/download"],"identifier_entries":[]},"links":{"outbound_url":null,"outbound_label":null,"outbound_source":null},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["ALEX FOO DA WENG"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2025-12-20"]},{"key":"dc:relation.isreferencedby","label":"Dc Relation Isreferencedby","values":["https://scholarbank.nus.edu.sg/handle/10635/319125"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Occlusion-Aware Perception","Out-of-Distribution Generalization","Multi-Object Representation Learning","Object-Centric Learning","Vision Models"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["https://scholarbank.nus.edu.sg/bitstreams/08c12eda-14b2-4bca-b936-6529a3af8be0/download"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://scholarbank.nus.edu.sg/bitstreams/1c011bd7-900c-48ec-9d24-5a1f50374d84/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Despite significant advancements, real-world deployment of modern vision models in critical applications remains limited by poor out-of-distribution generalization, failures under occlusion, reliance on large high-quality datasets, and limited interpretability. We posit that these limitations arise from fundamental differences between modern vision models and the human visual cortex. While current models mimic hierarchical processing, they treat images as holistic inputs. In contrast, human vision learns objects as compositional building blocks, enabling robust recognition, occlusion reasoning, generalization, and sample-efficient learning. To address this gap, this thesis proposes a unified object-centric representation framework with three stages: discovery, representation, and application. The proposed models improve robustness, sample efficiency, occlusion awareness, and generalization. OC-Net addresses limitations of pixel-based reconstruction using a feature connectivity algorithm and object-centric regularization optimized for downstream tasks. ECO-Net introduces graph-based representations of object parts, a co-part discovery algorithm, and a memory module, enabling improved performance on multi-part and occluded objects. Finally, DP-GAT integrates these ideas by combining OC-Net’s unsupervised object discovery with ECO-Net’s relational modeling, achieving strong results in disease progression prediction and 3D tumor segmentation. Together, these contributions support more reliable deployment of vision models in real-world applications."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["9f3c6f2aa8f96291ef77b0a18484521d","7b07434aaf16d9825f32cad56477a9f6","4f682e383baa25d8c70ee15fa01f7a7b"]},{"key":"dc:title","label":"Title","values":["COMPOSITIONAL OBJECT-CENTRIC REPRESENTATIONS FOR ROBUST VISUAL PERCEPTION"]}]}],"canonical_facts":{"dc:creator":["ALEX FOO DA WENG"],"dc:date.issued":["2025-12-20"],"dc:description.abstract":["Despite significant advancements, real-world deployment of modern vision models in critical applications remains limited by poor out-of-distribution generalization, failures under occlusion, reliance on large high-quality datasets, and limited interpretability. We posit that these limitations arise from fundamental differences between modern vision models and the human visual cortex. While current models mimic hierarchical processing, they treat images as holistic inputs. In contrast, human vision learns objects as compositional building blocks, enabling robust recognition, occlusion reasoning, generalization, and sample-efficient learning. To address this gap, this thesis proposes a unified object-centric representation framework with three stages: discovery, representation, and application. The proposed models improve robustness, sample efficiency, occlusion awareness, and generalization. OC-Net addresses limitations of pixel-based reconstruction using a feature connectivity algorithm and object-centric regularization optimized for downstream tasks. ECO-Net introduces graph-based representations of object parts, a co-part discovery algorithm, and a memory module, enabling improved performance on multi-part and occluded objects. Finally, DP-GAT integrates these ideas by combining OC-Net’s unsupervised object discovery with ECO-Net’s relational modeling, achieving strong results in disease progression prediction and 3D tumor segmentation. Together, these contributions support more reliable deployment of vision models in real-world applications."],"dc:format.checksum.md5":["9f3c6f2aa8f96291ef77b0a18484521d","7b07434aaf16d9825f32cad56477a9f6","4f682e383baa25d8c70ee15fa01f7a7b"],"dc:identifier.uri":["https://scholarbank.nus.edu.sg/bitstreams/1c011bd7-900c-48ec-9d24-5a1f50374d84/download"],"dc:relation.isreferencedby":["https://scholarbank.nus.edu.sg/handle/10635/319125"],"dc:rights":["https://scholarbank.nus.edu.sg/bitstreams/08c12eda-14b2-4bca-b936-6529a3af8be0/download"],"dc:subject":["Occlusion-Aware Perception","Out-of-Distribution Generalization","Multi-Object Representation Learning","Object-Centric Learning","Vision Models"],"dc:title":["COMPOSITIONAL OBJECT-CENTRIC REPRESENTATIONS FOR ROBUST VISUAL PERCEPTION"],"dc:type":["Thesis"]},"updated_at":"2026-07-24T03:31:26Z"}