{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/132655"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/132655","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Evaluating ventral visual stream contributions to human visual robustness with deep learning approaches","abstract":"The human visual system exhibits remarkable robustness, allowing resilient object recognition despite noisy and ambiguous sensory inputs. Such perceptual invariance has been proposed to stem from hierarchical computations along the ventral visual stream (VVS), a series of brain regions that progressively transform visual representation into increasingly abstract representations. In contrast, deep convolutional neural networks (DCNNs), despite being regarded as the closest approximator of the biological visual system, remain vulnerable to variations and adversarial perturbations that humans easily overcome. This discrepancy raises questions about the underlying mechanisms that confer robustness in human vision and whether they can be adequately captured by DCNNs. This thesis leverages deep convolutional neural networks (DCNNs) to investigate the role of VVS in achieving visual robustness, with a particular focus on how neural representational spaces evolve along the hierarchy. We start with aligning DCNNs to successive regions of the VVS and demonstrate hierarchical improvements in robustness that emerge as progressively higher-order representations are used, suggesting that robustness arises from the transformation of representations across the visual hierarchy rather than from any single brain region. We next examine what makes VVS representations uniquely suited to supporting robustness from two complementary perspectives. First, we assess whether the critical aspect of the representational spaces across the VVS is on the granularity of neural manifold characteristics and their separability, in accordance with a popular manifold disentangling framework. Second, inspired by recent efforts to characterize differences in spatial frequency preferences between humans and models, we test whether VVS representations implicitly bias DCNNs towards the frequency channels most relied on by human vision, thereby guiding models towards the spectral features critical for robust visual processing. The findings from these experiments are expected to contribute to not only the understanding of robust visual inference in the human brain, but also the development of more robust artificial vision systems that better emulate human perceptual capabilities.","abstract_html":"The human visual system exhibits remarkable robustness, allowing resilient object recognition despite noisy and ambiguous sensory inputs. Such perceptual invariance has been proposed to stem from hierarchical computations along the ventral visual stream (VVS), a series of brain regions that progressively transform visual representation into increasingly abstract representations. In contrast, deep convolutional neural networks (DCNNs), despite being regarded as the closest approximator of the biological visual system, remain vulnerable to variations and adversarial perturbations that humans easily overcome. This discrepancy raises questions about the underlying mechanisms that confer robustness in human vision and whether they can be adequately captured by DCNNs. This thesis leverages deep convolutional neural networks (DCNNs) to investigate the role of VVS in achieving visual robustness, with a particular focus on how neural representational spaces evolve along the hierarchy. We start with aligning DCNNs to successive regions of the VVS and demonstrate hierarchical improvements in robustness that emerge as progressively higher-order representations are used, suggesting that robustness arises from the transformation of representations across the visual hierarchy rather than from any single brain region. We next examine what makes VVS representations uniquely suited to supporting robustness from two complementary perspectives. First, we assess whether the critical aspect of the representational spaces across the VVS is on the granularity of neural manifold characteristics and their separability, in accordance with a popular manifold disentangling framework. Second, inspired by recent efforts to characterize differences in spatial frequency preferences between humans and models, we test whether VVS representations implicitly bias DCNNs towards the frequency channels most relied on by human vision, thereby guiding models towards the spectral features critical for robust visual processing. The findings from these experiments are expected to contribute to not only the understanding of robust visual inference in the human brain, but also the development of more robust artificial vision systems that better emulate human perceptual capabilities.","abstract_has_math":false,"creators":["Shao, Zhenan"],"institution":"University of Illinois Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Psychology","degree_department":null,"school":null,"contributors":["Beck, Diane M","Federmeier, Kara","Koyejo, Sanmi","Uddenburg, Stefan","Willits, Jon"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-12","date_published":"2025-12","updated_at":"2026-07-22T22:25:07Z","subjects":["Human visual robustness","Deep neural networks","Ventral visual stream","Representation learning","Adversarial robustness","Object recognition"],"languages":["en"],"rights":["Copyright 2025 Zhenan Shao"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/132655","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Beck, Diane M","Federmeier, Kara","Koyejo, Sanmi","Uddenburg, Stefan","Willits, Jon"]},{"key":"dc:creator","label":"Author","values":["Shao, Zhenan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-12","2025-11-25"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Psychology"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Human visual robustness","Deep neural networks","Ventral visual stream","Representation learning","Adversarial robustness","Object recognition"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Zhenan Shao"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/132655"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The human visual system exhibits remarkable robustness, allowing resilient object recognition despite noisy and ambiguous sensory inputs. Such perceptual invariance has been proposed to stem from hierarchical computations along the ventral visual stream (VVS), a series of brain regions that progressively transform visual representation into increasingly abstract representations. In contrast, deep convolutional neural networks (DCNNs), despite being regarded as the closest approximator of the biological visual system, remain vulnerable to variations and adversarial perturbations that humans easily overcome. This discrepancy raises questions about the underlying mechanisms that confer robustness in human vision and whether they can be adequately captured by DCNNs. This thesis leverages deep convolutional neural networks (DCNNs) to investigate the role of VVS in achieving visual robustness, with a particular focus on how neural representational spaces evolve along the hierarchy. We start with aligning DCNNs to successive regions of the VVS and demonstrate hierarchical improvements in robustness that emerge as progressively higher-order representations are used, suggesting that robustness arises from the transformation of representations across the visual hierarchy rather than from any single brain region. We next examine what makes VVS representations uniquely suited to supporting robustness from two complementary perspectives. First, we assess whether the critical aspect of the representational spaces across the VVS is on the granularity of neural manifold characteristics and their separability, in accordance with a popular manifold disentangling framework. Second, inspired by recent efforts to characterize differences in spatial frequency preferences between humans and models, we test whether VVS representations implicitly bias DCNNs towards the frequency channels most relied on by human vision, thereby guiding models towards the spectral features critical for robust visual processing. The findings from these experiments are expected to contribute to not only the understanding of robust visual inference in the human brain, but also the development of more robust artificial vision systems that better emulate human perceptual capabilities.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2027-12-01","The student, Zhenan Shao, accepted the attached license on 2025-11-23 at 15:43.","The student, Zhenan Shao, submitted this Dissertation for approval on 2025-11-23 at 15:50.","This Dissertation was approved for publication on 2025-11-25 at 11:35.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22933 on 2026-02-19 at 18:45:57"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Evaluating ventral visual stream contributions to human visual robustness with deep learning approaches"]}]}],"canonical_facts":{"dc:contributor":["Beck, Diane M","Federmeier, Kara","Koyejo, Sanmi","Uddenburg, Stefan","Willits, Jon"],"dc:creator":["Shao, Zhenan"],"dc:date":["2025-12","2025-11-25"],"dc:description":["The human visual system exhibits remarkable robustness, allowing resilient object recognition despite noisy and ambiguous sensory inputs. Such perceptual invariance has been proposed to stem from hierarchical computations along the ventral visual stream (VVS), a series of brain regions that progressively transform visual representation into increasingly abstract representations. In contrast, deep convolutional neural networks (DCNNs), despite being regarded as the closest approximator of the biological visual system, remain vulnerable to variations and adversarial perturbations that humans easily overcome. This discrepancy raises questions about the underlying mechanisms that confer robustness in human vision and whether they can be adequately captured by DCNNs. This thesis leverages deep convolutional neural networks (DCNNs) to investigate the role of VVS in achieving visual robustness, with a particular focus on how neural representational spaces evolve along the hierarchy. We start with aligning DCNNs to successive regions of the VVS and demonstrate hierarchical improvements in robustness that emerge as progressively higher-order representations are used, suggesting that robustness arises from the transformation of representations across the visual hierarchy rather than from any single brain region. We next examine what makes VVS representations uniquely suited to supporting robustness from two complementary perspectives. First, we assess whether the critical aspect of the representational spaces across the VVS is on the granularity of neural manifold characteristics and their separability, in accordance with a popular manifold disentangling framework. Second, inspired by recent efforts to characterize differences in spatial frequency preferences between humans and models, we test whether VVS representations implicitly bias DCNNs towards the frequency channels most relied on by human vision, thereby guiding models towards the spectral features critical for robust visual processing. The findings from these experiments are expected to contribute to not only the understanding of robust visual inference in the human brain, but also the development of more robust artificial vision systems that better emulate human perceptual capabilities.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2027-12-01","The student, Zhenan Shao, accepted the attached license on 2025-11-23 at 15:43.","The student, Zhenan Shao, submitted this Dissertation for approval on 2025-11-23 at 15:50.","This Dissertation was approved for publication on 2025-11-25 at 11:35.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22933 on 2026-02-19 at 18:45:57"],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/132655"],"dc:language":["en"],"dc:rights":["Copyright 2025 Zhenan Shao"],"dc:subject":["Human visual robustness","Deep neural networks","Ventral visual stream","Representation learning","Adversarial robustness","Object recognition"],"dc:title":["Evaluating ventral visual stream contributions to human visual robustness with deep learning approaches"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Psychology"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:07Z"}