{"id":{"repo_id":"tu-berlin","oai_identifier":"oai:depositonce.tu-berlin.de:11303/24674"},"canonical_url":"https://search.dev.ndltd.org/etd/tu-berlin/oai:depositonce.tu-berlin.de:11303/24674","repository":{"repo_id":"tu-berlin","name":"Technische Universität Berlin","base_url":"https://api-depositonce.tu-berlin.de/server/oai/request"},"display":{"title":"Representational alignment of humans and machines for computer vision","abstract":"Recent progress in computer vision has led to deep neural networks that match or even surpass human-level performance across various machine learning tasks. However, we identify a significant difference between the representations learned by these models and human conceptual understanding. We introduce a novel approximate Bayesian method for learning object concept representations from human behavior in a triplet odd-one-out task. The method uses variational inference to produce sparse, non-negative representations with uncertainty estimates, improving the reproducibility of object dimensions and the consistency in predicting human behavior. These mental representations can downstream be used to unravel the differences between human and neural network representations of object concepts. Our findings reveal that contemporary computer vision models, while achieving human-level performance, fail to align with human mental representations. Factors such as training data and objective function appear to play a crucial role in alignment, while the model architecture and number of parameters have minimal impact. We pinpoint the reason for this misalignment: human conceptual knowledge is hierarchically organized, while traditional training methods embed similar images close together, thereby often ignoring the global organization of object concepts. To address this mismatch, we developed a novel method that improves the global structure of these representations by linearly aligning them with human similarity judgments, leading to improved performance in few-shot learning and anomaly detection tasks. However, linearly aligning model representation can only improve the representations of large image/text models that have seen billions of images during pretraining but it fails to achieve the same for smaller models trained on less data. By training a teacher model to imitate human judgments and subsequently distilling this human-like structure into pretrained student models, we can improve the representations of any vision foundation model—irrespective of its pretraining task or the number of images it has seen during pretraining. The human-aligned models can better approximate human conceptual knowledge and achieve substantially improved generalization performance and out-of-distribution robustness than the non-aligned base models. Our findings underscore the importance of incorporating the hierarchical nature of human conceptual knowledge into the representations of neural networks, paving the way for more robust, interpretable, and human-aligned artificial intelligence systems.","abstract_html":"Recent progress in computer vision has led to deep neural networks that match or even surpass human-level performance across various machine learning tasks. However, we identify a significant difference between the representations learned by these models and human conceptual understanding. We introduce a novel approximate Bayesian method for learning object concept representations from human behavior in a triplet odd-one-out task. The method uses variational inference to produce sparse, non-negative representations with uncertainty estimates, improving the reproducibility of object dimensions and the consistency in predicting human behavior. These mental representations can downstream be used to unravel the differences between human and neural network representations of object concepts. Our findings reveal that contemporary computer vision models, while achieving human-level performance, fail to align with human mental representations. Factors such as training data and objective function appear to play a crucial role in alignment, while the model architecture and number of parameters have minimal impact. We pinpoint the reason for this misalignment: human conceptual knowledge is hierarchically organized, while traditional training methods embed similar images close together, thereby often ignoring the global organization of object concepts. To address this mismatch, we developed a novel method that improves the global structure of these representations by linearly aligning them with human similarity judgments, leading to improved performance in few-shot learning and anomaly detection tasks. However, linearly aligning model representation can only improve the representations of large image/text models that have seen billions of images during pretraining but it fails to achieve the same for smaller models trained on less data. By training a teacher model to imitate human judgments and subsequently distilling this human-like structure into pretrained student models, we can improve the representations of any vision foundation model—irrespective of its pretraining task or the number of images it has seen during pretraining. The human-aligned models can better approximate human conceptual knowledge and achieve substantially improved generalization performance and out-of-distribution robustness than the non-aligned base models. Our findings underscore the importance of incorporating the hierarchical nature of human conceptual knowledge into the representations of neural networks, paving the way for more robust, interpretable, and human-aligned artificial intelligence systems.","abstract_has_math":false,"creators":["Muttenthaler, Lukas"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Müller, Klaus-Robert"],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025","date_published":"2025","updated_at":"2026-07-27T21:28:40Z","subjects":[],"languages":["en"],"rights":[],"rights_urls":["https://creativecommons.org/licenses/by/4.0/"],"identifier_entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://doi.org/10.14279/depositonce-23490"],"render_values":[{"text":"https://doi.org/10.14279/depositonce-23490","href":"https://doi.org/10.14279/depositonce-23490","code":true}]}]},"links":{"outbound_url":"https://depositonce.tu-berlin.de/handle/11303/24674","outbound_label":"Repository record","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Müller, Klaus-Robert"]},{"key":"dc:creator","label":"Author","values":["Muttenthaler, Lukas"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2025-06-04T10:04:06Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2025-06-04T10:04:06Z"]},{"key":"dc:date.issued","label":"Date","values":["2025"]},{"key":"dc:type","label":"Dc Type","values":["Doctoral Thesis"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]},{"key":"dc:rights.uri","label":"Rights URI","values":["https://creativecommons.org/licenses/by/4.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://depositonce.tu-berlin.de/handle/11303/24674","https://doi.org/10.14279/depositonce-23490"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Recent progress in computer vision has led to deep neural networks that match or even surpass human-level performance across various machine learning tasks. However, we identify a significant difference between the representations learned by these models and human conceptual understanding. We introduce a novel approximate Bayesian method for learning object concept representations from human behavior in a triplet odd-one-out task. The method uses variational inference to produce sparse, non-negative representations with uncertainty estimates, improving the reproducibility of object dimensions and the consistency in predicting human behavior. These mental representations can downstream be used to unravel the differences between human and neural network representations of object concepts. Our findings reveal that contemporary computer vision models, while achieving human-level performance, fail to align with human mental representations. Factors such as training data and objective function appear to play a crucial role in alignment, while the model architecture and number of parameters have minimal impact. We pinpoint the reason for this misalignment: human conceptual knowledge is hierarchically organized, while traditional training methods embed similar images close together, thereby often ignoring the global organization of object concepts. To address this mismatch, we developed a novel method that improves the global structure of these representations by linearly aligning them with human similarity judgments, leading to improved performance in few-shot learning and anomaly detection tasks. However, linearly aligning model representation can only improve the representations of large image/text models that have seen billions of images during pretraining but it fails to achieve the same for smaller models trained on less data. By training a teacher model to imitate human judgments and subsequently distilling this human-like structure into pretrained student models, we can improve the representations of any vision foundation model—irrespective of its pretraining task or the number of images it has seen during pretraining. The human-aligned models can better approximate human conceptual knowledge and achieve substantially improved generalization performance and out-of-distribution robustness than the non-aligned base models. Our findings underscore the importance of incorporating the hierarchical nature of human conceptual knowledge into the representations of neural networks, paving the way for more robust, interpretable, and human-aligned artificial intelligence systems.","Die jüngsten Fortschritte im Bereich des maschinellen Lernens haben zu tiefen neuronalen Netzen geführt, die bei verschiedenen visuellen Aufgaben die Leistung des Menschen erreichen oder sogar übertreffen. Wir finden jedoch einen erheblichen Unterschieds zwischen den von diesen Modellen erlernten Darstellungen und dem menschlichen Konzeptverständnis. Wir stellen eine approximative Bayes’sche Methode zum Lernen von Objektrepräsentationen aus menschlichem Verhalten vor. Die Methode verwendet Variationsinferenz, um spärliche, nicht-negative Repräsentationen mit Unsicherheitsschätzungen zu erzeugen, die die Reproduzierbarkeit von Objektdimensionen und die Konsistenz bei der Vorhersage menschlichen Verhaltens verbessern. Unsere Analysen zeigen, dass aktuelle Bildverarbeitungsmodelle nicht mit den mentalen Objektrepräsentationen des Menschen übereinstimmen. Faktoren wie Trainingsdaten und Zielfunktion scheinen eine entscheidende Rolle bei der Anpassung zu spielen, während die Modellarchitektur und die Anzahl der Parameter nur minimale Auswirkungen haben. Wir haben den Grund für diese Fehlanpassung ausgemacht: Das menschliche Begriffswissen ist hierarchisch organisiert, während herkömmliche Trainingsmethoden ähnliche Bilder nahe beieinander einbetten und dabei oft die globale Organisation von Objektkonzepten vernachlässigen. Um diese Diskrepanz zu beheben, haben wir eine Methode entwickelt, die die globale Struktur dieser Repräsentationen verbessert, indem sie linear mit menschlichen Ähnlichkeitsurteilen abgeglichen wird, was zu einer verbesserten Leistung beim Lernen mit wenigen Bildern und bei der Erkennung von Anomalien führt. Die lineare Ausrichtung der Modellrepräsentation kann jedoch nur die Repräsentationen großer Bild-/Text Modelle verbessern, die während des Trainings Milliarden von Bildern gesehen haben, während dies bei kleineren Modellen, die mit weniger Daten trainiert wurden, nicht möglich ist. Indem wir ein (grosses) Lehrermodell trainieren, um menschliche Urteile zu imitieren, und anschließend diese menschenähnliche Struktur in vortrainierte Schülermodelle destillieren, können wir die Repräsentationen jedes beliebigen Bildverarbeitungsmodells verbessern— unabhängig von dessen Trainingsaufgabe oder der Anzahl der Bilder, die es während des Trainings gesehen hat. Die neuen Modelle können die menschlichen Konzepte besser verstehen und erreichen eine wesentlich bessere Generalisierungsleistung und Robustheit gegenüber Verteilungsfehlern als die nicht-menschenähnlichen Basismodelle. Unsere Ergebnisse unterstreichen, wie wichtig es ist, die hierarchische Natur des menschlichen Konzeptwissens in die Darstellungen neuronaler Netze einzubeziehen, um denWeg für robustere, interpretierbare und auf den Menschen abgestimmte künstliche Intelligenzsysteme zu ebnen."]},{"key":"dc:title","label":"Title","values":["Representational alignment of humans and machines for computer vision"]}]}],"canonical_facts":{"dc:contributor.advisor":["Müller, Klaus-Robert"],"dc:creator":["Muttenthaler, Lukas"],"dc:date.accessioned":["2025-06-04T10:04:06Z"],"dc:date.available":["2025-06-04T10:04:06Z"],"dc:date.issued":["2025"],"dc:description.abstract":["Recent progress in computer vision has led to deep neural networks that match or even surpass human-level performance across various machine learning tasks. However, we identify a significant difference between the representations learned by these models and human conceptual understanding. We introduce a novel approximate Bayesian method for learning object concept representations from human behavior in a triplet odd-one-out task. The method uses variational inference to produce sparse, non-negative representations with uncertainty estimates, improving the reproducibility of object dimensions and the consistency in predicting human behavior. These mental representations can downstream be used to unravel the differences between human and neural network representations of object concepts. Our findings reveal that contemporary computer vision models, while achieving human-level performance, fail to align with human mental representations. Factors such as training data and objective function appear to play a crucial role in alignment, while the model architecture and number of parameters have minimal impact. We pinpoint the reason for this misalignment: human conceptual knowledge is hierarchically organized, while traditional training methods embed similar images close together, thereby often ignoring the global organization of object concepts. To address this mismatch, we developed a novel method that improves the global structure of these representations by linearly aligning them with human similarity judgments, leading to improved performance in few-shot learning and anomaly detection tasks. However, linearly aligning model representation can only improve the representations of large image/text models that have seen billions of images during pretraining but it fails to achieve the same for smaller models trained on less data. By training a teacher model to imitate human judgments and subsequently distilling this human-like structure into pretrained student models, we can improve the representations of any vision foundation model—irrespective of its pretraining task or the number of images it has seen during pretraining. The human-aligned models can better approximate human conceptual knowledge and achieve substantially improved generalization performance and out-of-distribution robustness than the non-aligned base models. Our findings underscore the importance of incorporating the hierarchical nature of human conceptual knowledge into the representations of neural networks, paving the way for more robust, interpretable, and human-aligned artificial intelligence systems.","Die jüngsten Fortschritte im Bereich des maschinellen Lernens haben zu tiefen neuronalen Netzen geführt, die bei verschiedenen visuellen Aufgaben die Leistung des Menschen erreichen oder sogar übertreffen. Wir finden jedoch einen erheblichen Unterschieds zwischen den von diesen Modellen erlernten Darstellungen und dem menschlichen Konzeptverständnis. Wir stellen eine approximative Bayes’sche Methode zum Lernen von Objektrepräsentationen aus menschlichem Verhalten vor. Die Methode verwendet Variationsinferenz, um spärliche, nicht-negative Repräsentationen mit Unsicherheitsschätzungen zu erzeugen, die die Reproduzierbarkeit von Objektdimensionen und die Konsistenz bei der Vorhersage menschlichen Verhaltens verbessern. Unsere Analysen zeigen, dass aktuelle Bildverarbeitungsmodelle nicht mit den mentalen Objektrepräsentationen des Menschen übereinstimmen. Faktoren wie Trainingsdaten und Zielfunktion scheinen eine entscheidende Rolle bei der Anpassung zu spielen, während die Modellarchitektur und die Anzahl der Parameter nur minimale Auswirkungen haben. Wir haben den Grund für diese Fehlanpassung ausgemacht: Das menschliche Begriffswissen ist hierarchisch organisiert, während herkömmliche Trainingsmethoden ähnliche Bilder nahe beieinander einbetten und dabei oft die globale Organisation von Objektkonzepten vernachlässigen. Um diese Diskrepanz zu beheben, haben wir eine Methode entwickelt, die die globale Struktur dieser Repräsentationen verbessert, indem sie linear mit menschlichen Ähnlichkeitsurteilen abgeglichen wird, was zu einer verbesserten Leistung beim Lernen mit wenigen Bildern und bei der Erkennung von Anomalien führt. Die lineare Ausrichtung der Modellrepräsentation kann jedoch nur die Repräsentationen großer Bild-/Text Modelle verbessern, die während des Trainings Milliarden von Bildern gesehen haben, während dies bei kleineren Modellen, die mit weniger Daten trainiert wurden, nicht möglich ist. Indem wir ein (grosses) Lehrermodell trainieren, um menschliche Urteile zu imitieren, und anschließend diese menschenähnliche Struktur in vortrainierte Schülermodelle destillieren, können wir die Repräsentationen jedes beliebigen Bildverarbeitungsmodells verbessern— unabhängig von dessen Trainingsaufgabe oder der Anzahl der Bilder, die es während des Trainings gesehen hat. Die neuen Modelle können die menschlichen Konzepte besser verstehen und erreichen eine wesentlich bessere Generalisierungsleistung und Robustheit gegenüber Verteilungsfehlern als die nicht-menschenähnlichen Basismodelle. Unsere Ergebnisse unterstreichen, wie wichtig es ist, die hierarchische Natur des menschlichen Konzeptwissens in die Darstellungen neuronaler Netze einzubeziehen, um denWeg für robustere, interpretierbare und auf den Menschen abgestimmte künstliche Intelligenzsysteme zu ebnen."],"dc:identifier.uri":["https://depositonce.tu-berlin.de/handle/11303/24674","https://doi.org/10.14279/depositonce-23490"],"dc:language.iso":["en"],"dc:rights.uri":["https://creativecommons.org/licenses/by/4.0/"],"dc:title":["Representational alignment of humans and machines for computer vision"],"dc:type":["Doctoral Thesis"]},"updated_at":"2026-07-27T21:28:40Z"}