{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/132654"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/132654","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Category information in real-world scenes: evaluation and reconstruction of human category spaces","abstract":"Categorization is fundamental to scene understanding, yet there is relatively little research into the structure of human scene categories. This thesis focuses on examining whether various image-based feature spaces can approximate human category representations for real-world scenes. Similarity was used to assess category representation, and the category structure was visualized and compared by constructing geometric representations. Several feature spaces were tested, ranging from low- and mid-level visual features, layer activations of convolutional neural networks (CNNs) trained on real-world scenes, and transformer models. Chapter 2 examined whether the layer activations of CNNs can capture human typicality effects. Chapter 3 compared various categorization tasks and laid the groundwork for building a reliable and valid human categorization space. Building on the results of Chapter 3, Chapter 4 further measured the correspondence between feature spaces and human categorization results by comparing the distance matrix correlations derived from ordinal multidimensional scaling (MDS) and the alignment of p-median clustering results. Among the tested features here, the fc7 layer, the last layer of CNN before the classification layer, showed the strongest typicality effect, the high correlations (r = 0.74) in the ordinal MDS distance matrix and the highest correspondence to human p-median clusters (ARI = 0.56). This suggests that the category information aggregated in the later layer of the CNNs has some agreement with the category information used in similarity tasks for humans. Moreover, GPT4-o models when prompted with the same task instructions as human participants, had the best correlations to human category spaces, implicating the importance of task alignment between the models and human task. More clustering methods that capture the flexible nature of similarity, however, are needed to support stronger claims. Overall, this thesis demonstrates and highlights the usefulness of DNN information in modeling human categorization of real-world scenes. Based on these findings, it may be possible, eventually, to use DNNs as a substitute for human participants when generating similarity measures for items.","abstract_html":"Categorization is fundamental to scene understanding, yet there is relatively little research into the structure of human scene categories. This thesis focuses on examining whether various image-based feature spaces can approximate human category representations for real-world scenes. Similarity was used to assess category representation, and the category structure was visualized and compared by constructing geometric representations. Several feature spaces were tested, ranging from low- and mid-level visual features, layer activations of convolutional neural networks (CNNs) trained on real-world scenes, and transformer models. Chapter 2 examined whether the layer activations of CNNs can capture human typicality effects. Chapter 3 compared various categorization tasks and laid the groundwork for building a reliable and valid human categorization space. Building on the results of Chapter 3, Chapter 4 further measured the correspondence between feature spaces and human categorization results by comparing the distance matrix correlations derived from ordinal multidimensional scaling (MDS) and the alignment of p-median clustering results. Among the tested features here, the fc7 layer, the last layer of CNN before the classification layer, showed the strongest typicality effect, the high correlations (r = 0.74) in the ordinal MDS distance matrix and the highest correspondence to human p-median clusters (ARI = 0.56). This suggests that the category information aggregated in the later layer of the CNNs has some agreement with the category information used in similarity tasks for humans. Moreover, GPT4-o models when prompted with the same task instructions as human participants, had the best correlations to human category spaces, implicating the importance of task alignment between the models and human task. More clustering methods that capture the flexible nature of similarity, however, are needed to support stronger claims. Overall, this thesis demonstrates and highlights the usefulness of DNN information in modeling human categorization of real-world scenes. Based on these findings, it may be possible, eventually, to use DNNs as a substitute for human participants when generating similarity measures for items.","abstract_has_math":false,"creators":["Yang, Pei-Ling"],"institution":"University of Illinois Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Psychology","degree_department":null,"school":null,"contributors":["Beck, Diane M","Koehn, Hans F","Hummel, John E","Simons, Daniel J","Federmeier, Kara D","Willits, Jon A"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-12","date_published":"2025-12","updated_at":"2026-07-22T22:25:07Z","subjects":["scene categorization, similarity tasks, convolutional neural network, multidimensional scaling"],"languages":["en"],"rights":["Copyright 2025 Pei-Ling Yang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/132654","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Beck, Diane M","Koehn, Hans F","Hummel, John E","Simons, Daniel J","Federmeier, Kara D","Willits, Jon A"]},{"key":"dc:creator","label":"Author","values":["Yang, Pei-Ling"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-12","2025-11-26"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Psychology"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["scene categorization, similarity tasks, convolutional neural network, multidimensional scaling"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Pei-Ling Yang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/132654"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Categorization is fundamental to scene understanding, yet there is relatively little research into the structure of human scene categories. This thesis focuses on examining whether various image-based feature spaces can approximate human category representations for real-world scenes. Similarity was used to assess category representation, and the category structure was visualized and compared by constructing geometric representations. Several feature spaces were tested, ranging from low- and mid-level visual features, layer activations of convolutional neural networks (CNNs) trained on real-world scenes, and transformer models. Chapter 2 examined whether the layer activations of CNNs can capture human typicality effects. Chapter 3 compared various categorization tasks and laid the groundwork for building a reliable and valid human categorization space. Building on the results of Chapter 3, Chapter 4 further measured the correspondence between feature spaces and human categorization results by comparing the distance matrix correlations derived from ordinal multidimensional scaling (MDS) and the alignment of p-median clustering results. Among the tested features here, the fc7 layer, the last layer of CNN before the classification layer, showed the strongest typicality effect, the high correlations (r = 0.74) in the ordinal MDS distance matrix and the highest correspondence to human p-median clusters (ARI = 0.56). This suggests that the category information aggregated in the later layer of the CNNs has some agreement with the category information used in similarity tasks for humans. Moreover, GPT4-o models when prompted with the same task instructions as human participants, had the best correlations to human category spaces, implicating the importance of task alignment between the models and human task. More clustering methods that capture the flexible nature of similarity, however, are needed to support stronger claims. Overall, this thesis demonstrates and highlights the usefulness of DNN information in modeling human categorization of real-world scenes. Based on these findings, it may be possible, eventually, to use DNNs as a substitute for human participants when generating similarity measures for items.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2027-12-01","The student, Pei-Ling Yang, accepted the attached license on 2025-11-24 at 08:49.","The student, Pei-Ling Yang, submitted this Dissertation for approval on 2025-11-24 at 08:49.","This Dissertation was approved for publication on 2025-11-26 at 13:37.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22928 on 2026-02-19 at 18:45:57"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Category information in real-world scenes: evaluation and reconstruction of human category spaces"]}]}],"canonical_facts":{"dc:contributor":["Beck, Diane M","Koehn, Hans F","Hummel, John E","Simons, Daniel J","Federmeier, Kara D","Willits, Jon A"],"dc:creator":["Yang, Pei-Ling"],"dc:date":["2025-12","2025-11-26"],"dc:description":["Categorization is fundamental to scene understanding, yet there is relatively little research into the structure of human scene categories. This thesis focuses on examining whether various image-based feature spaces can approximate human category representations for real-world scenes. Similarity was used to assess category representation, and the category structure was visualized and compared by constructing geometric representations. Several feature spaces were tested, ranging from low- and mid-level visual features, layer activations of convolutional neural networks (CNNs) trained on real-world scenes, and transformer models. Chapter 2 examined whether the layer activations of CNNs can capture human typicality effects. Chapter 3 compared various categorization tasks and laid the groundwork for building a reliable and valid human categorization space. Building on the results of Chapter 3, Chapter 4 further measured the correspondence between feature spaces and human categorization results by comparing the distance matrix correlations derived from ordinal multidimensional scaling (MDS) and the alignment of p-median clustering results. Among the tested features here, the fc7 layer, the last layer of CNN before the classification layer, showed the strongest typicality effect, the high correlations (r = 0.74) in the ordinal MDS distance matrix and the highest correspondence to human p-median clusters (ARI = 0.56). This suggests that the category information aggregated in the later layer of the CNNs has some agreement with the category information used in similarity tasks for humans. Moreover, GPT4-o models when prompted with the same task instructions as human participants, had the best correlations to human category spaces, implicating the importance of task alignment between the models and human task. More clustering methods that capture the flexible nature of similarity, however, are needed to support stronger claims. Overall, this thesis demonstrates and highlights the usefulness of DNN information in modeling human categorization of real-world scenes. Based on these findings, it may be possible, eventually, to use DNNs as a substitute for human participants when generating similarity measures for items.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2027-12-01","The student, Pei-Ling Yang, accepted the attached license on 2025-11-24 at 08:49.","The student, Pei-Ling Yang, submitted this Dissertation for approval on 2025-11-24 at 08:49.","This Dissertation was approved for publication on 2025-11-26 at 13:37.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22928 on 2026-02-19 at 18:45:57"],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/132654"],"dc:language":["en"],"dc:rights":["Copyright 2025 Pei-Ling Yang"],"dc:subject":["scene categorization, similarity tasks, convolutional neural network, multidimensional scaling"],"dc:title":["Category information in real-world scenes: evaluation and reconstruction of human category spaces"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Psychology"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:07Z"}