{"id":{"repo_id":"tu-berlin","oai_identifier":"oai:depositonce.tu-berlin.de:11303/26655"},"canonical_url":"https://search.dev.ndltd.org/etd/tu-berlin/oai:depositonce.tu-berlin.de:11303/26655","repository":{"repo_id":"tu-berlin","name":"Technische Universität Berlin","base_url":"https://api-depositonce.tu-berlin.de/server/oai/request"},"display":{"title":"From local explanations to comprehensive mechanistic understanding of deep vision models","abstract":"Deep learning models have evolved into a cornerstone of modern industry and science, enabling applications from medical diagnosis and perception to conversational systems. Over the past two decades, both models and datasets have grown substantially in scale: models now reach trillions of parameters, trained on datasets with billions of samples. Despite their success, our understanding and control of these models remain limited, hindering safe and robust deployment in critical domains. Regulations such as the EU AI Act further emphasize the need for transparent and accountable AI systems. The field of eXplainable Artificial Intelligence (XAI) has introduced techniques like attribution maps and feature visualizations to illuminate singular aspects of model behavior. Yet, achieving a comprehensive understanding that enables validation and control of the complex mechanisms inside AI models requires the combination of multiple XAI perspectives. This is already challenging, and as most approaches rely on manual inspection of individual explanations, they fail to scale with the size and complexity of today’s models and datasets. This dissertation develops an explainability framework that is (i) mechanistic, by providing component-level insights (ii) comprehensive, by integrating multiple interpretability perspectives, (iii) scalable, by aggregating and summarizing explanations across data and model components while flagging outliers and deviations, and (iv) actionable, by directly informing practical strategies for refining and improving model behavior. Key contributions of this thesis include: (1) A foundation for comprehensive mechanistic explanations that integrate component-level attributions, input localizations, and feature visualizations, validated via a user study. (2) Measuring and improving interpretability by introducing multiple measures to estimate human interpretability of components, further validated through a user study, and methods to mitigate issues such as polysemanticity. (3) Prototypical Concept-based Explanations that summarize model behavior across entire datasets using a small set of concept-level prototypes. (4) Semantic component mbeddings that enable text-based semantic search, labeling, clustering, and comparison of model components. (5) Automated auditing methods such as outlier detection and concept alignment analysis to flag spurious or unexpected behaviors. (6) Interpretability-informed correction techniques that refine and correct models based on mechanistic insights. Through experiments on state-of-the-art vision models, this work demonstrates that mechanistic explanations enable the identification, understanding, and correction of model failures, providing a path toward more transparent, robust and controllable AI systems.","abstract_html":"Deep learning models have evolved into a cornerstone of modern industry and science, enabling applications from medical diagnosis and perception to conversational systems. Over the past two decades, both models and datasets have grown substantially in scale: models now reach trillions of parameters, trained on datasets with billions of samples. Despite their success, our understanding and control of these models remain limited, hindering safe and robust deployment in critical domains. Regulations such as the EU AI Act further emphasize the need for transparent and accountable AI systems. The field of eXplainable Artificial Intelligence (XAI) has introduced techniques like attribution maps and feature visualizations to illuminate singular aspects of model behavior. Yet, achieving a comprehensive understanding that enables validation and control of the complex mechanisms inside AI models requires the combination of multiple XAI perspectives. This is already challenging, and as most approaches rely on manual inspection of individual explanations, they fail to scale with the size and complexity of today’s models and datasets. This dissertation develops an explainability framework that is (i) mechanistic, by providing component-level insights (ii) comprehensive, by integrating multiple interpretability perspectives, (iii) scalable, by aggregating and summarizing explanations across data and model components while flagging outliers and deviations, and (iv) actionable, by directly informing practical strategies for refining and improving model behavior. Key contributions of this thesis include: (1) A foundation for comprehensive mechanistic explanations that integrate component-level attributions, input localizations, and feature visualizations, validated via a user study. (2) Measuring and improving interpretability by introducing multiple measures to estimate human interpretability of components, further validated through a user study, and methods to mitigate issues such as polysemanticity. (3) Prototypical Concept-based Explanations that summarize model behavior across entire datasets using a small set of concept-level prototypes. (4) Semantic component mbeddings that enable text-based semantic search, labeling, clustering, and comparison of model components. (5) Automated auditing methods such as outlier detection and concept alignment analysis to flag spurious or unexpected behaviors. (6) Interpretability-informed correction techniques that refine and correct models based on mechanistic insights. Through experiments on state-of-the-art vision models, this work demonstrates that mechanistic explanations enable the identification, understanding, and correction of model failures, providing a path toward more transparent, robust and controllable AI systems.","abstract_has_math":false,"creators":["Dreyer, Maximilian"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Samek, Wojciech","Lapuschkin, Sebastian"],"committee_chairs":[],"committee_members":[],"year":2026,"date_issued":"2026","date_published":"2026","updated_at":"2026-07-27T21:28:44Z","subjects":[],"languages":["en"],"rights":[],"rights_urls":["https://creativecommons.org/licenses/by-sa/4.0/"],"identifier_entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://doi.org/10.14279/depositonce-25485"],"render_values":[{"text":"https://doi.org/10.14279/depositonce-25485","href":"https://doi.org/10.14279/depositonce-25485","code":true}]}]},"links":{"outbound_url":"https://depositonce.tu-berlin.de/handle/11303/26655","outbound_label":"Repository record","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Samek, Wojciech","Lapuschkin, Sebastian"]},{"key":"dc:creator","label":"Author","values":["Dreyer, Maximilian"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2026-03-16T08:44:53Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2026-03-16T08:44:53Z"]},{"key":"dc:date.issued","label":"Date","values":["2026"]},{"key":"dc:type","label":"Dc Type","values":["Doctoral Thesis"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]},{"key":"dc:rights.uri","label":"Rights URI","values":["https://creativecommons.org/licenses/by-sa/4.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://depositonce.tu-berlin.de/handle/11303/26655","https://doi.org/10.14279/depositonce-25485"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Deep learning models have evolved into a cornerstone of modern industry and science, enabling applications from medical diagnosis and perception to conversational systems. Over the past two decades, both models and datasets have grown substantially in scale: models now reach trillions of parameters, trained on datasets with billions of samples. Despite their success, our understanding and control of these models remain limited, hindering safe and robust deployment in critical domains. Regulations such as the EU AI Act further emphasize the need for transparent and accountable AI systems. The field of eXplainable Artificial Intelligence (XAI) has introduced techniques like attribution maps and feature visualizations to illuminate singular aspects of model behavior. Yet, achieving a comprehensive understanding that enables validation and control of the complex mechanisms inside AI models requires the combination of multiple XAI perspectives. This is already challenging, and as most approaches rely on manual inspection of individual explanations, they fail to scale with the size and complexity of today’s models and datasets. This dissertation develops an explainability framework that is (i) mechanistic, by providing component-level insights (ii) comprehensive, by integrating multiple interpretability perspectives, (iii) scalable, by aggregating and summarizing explanations across data and model components while flagging outliers and deviations, and (iv) actionable, by directly informing practical strategies for refining and improving model behavior. Key contributions of this thesis include: (1) A foundation for comprehensive mechanistic explanations that integrate component-level attributions, input localizations, and feature visualizations, validated via a user study. (2) Measuring and improving interpretability by introducing multiple measures to estimate human interpretability of components, further validated through a user study, and methods to mitigate issues such as polysemanticity. (3) Prototypical Concept-based Explanations that summarize model behavior across entire datasets using a small set of concept-level prototypes. (4) Semantic component mbeddings that enable text-based semantic search, labeling, clustering, and comparison of model components. (5) Automated auditing methods such as outlier detection and concept alignment analysis to flag spurious or unexpected behaviors. (6) Interpretability-informed correction techniques that refine and correct models based on mechanistic insights. Through experiments on state-of-the-art vision models, this work demonstrates that mechanistic explanations enable the identification, understanding, and correction of model failures, providing a path toward more transparent, robust and controllable AI systems.","Tiefe neuronale Netze sind ein Eckpfeiler moderner Industrie und Wissenschaft, mit Anwendungen von der medizinischen Diagnose bis zu Konversationssystemen. Modelle und Datensätze sind in den letzten zwei Jahrzehnten stark gewachsen: Modelle erreichen Billionen von Parametern und werden auf Milliarden von Datenpunkten trainiert. Trotz ihres Erfolgs bleiben Verständnis und Kontrolle über diese Modelle begrenzt, was ihren sicheren Einsatz in kritischen Bereichen erschwert. Vorschriften wie das EU-KI-Gesetz betonen dabei die Notwendigkeit transparenter KI-Systeme. Techniken der erklärbaren künstlichen Intelligenz (XAI) wie Heatmaps und Konzeptvisualisierung beleuchten einzelne Aspekte des Modellverhaltens. Um jedoch ein umfassendes Verständnis zu erreichen, das die Validierung und Kontrolle komplexer Modellmechanismen ermöglicht, ist die Kombination mehrerer XAI-Perspektiven erforderlich. Diese erhöht bereits die kognitive Komplexität aufgrund der hohen Informationsdichte, und eine manuelle Überprüfung einzelner Erklärungen skaliert somit nicht mit der Größe aktueller Modelle und Datensätze. Diese Dissertation entwickelt ein Erklärbarkeits-Framework, das (i) mechanistische Einblicke auf Ebene einzelner Modellkomponenten liefert, (ii) umfassend mehrere Interpretierbarkeitsperspektiven integriert, (iii) skalierbar Erklärungen aggregiert und Ausreißer markiert, und (iv) handlungsorientiert Strategien zur Modellverbesserung bietet. Zentrale Beiträge dieser Arbeit sind: (1) Eine Grundlage für umfassende mechanistische Erklärungen, die komponentenspezifische Relevanzwerte, Lokalisierungen und Konzeptvisualisierungen integriert und in einer Benutzerstudie validiert wurde. (2) Messung und Verbesserung der Interpretierbarkeit durch Methoden zur Bewertung der menschlichen Interpretierbarkeit von Komponenten, ebenfalls in einer Benutzerstudie validiert, sowie zur Reduktion von Polysemantik. (3) Protoypische konzeptbasierte Erklärungen, die das Modellverhalten über Datensätze hinweg anhand weniger Konzept-basierten Prototypen zusammenfassen. (4) Semantische Komponenten-Einbettung, die textbasierte Suche, Beschreibungen, Gruppierung und Vergleiche von Modellkomponenten ermöglicht. (5) Automatisierte Prüfmethoden wie Ausreißererkennung und Konzeptabgleichanalyse zur Identifikation problematischer oder unerwarteter Verhaltensweisen. (6) Interpretierbarkeitsbasierte Korrekturtechniken, die Modelle anhand mechanistischer Einsichten gezielt korrigieren. Experimente mit modernen KI Modellen zeigen, dass mechanistische Erklärungen die Identifikation, das Verständnis und die Korrektur von Modellfehlern ermöglichen und damit den Weg zu transparenteren, robusteren und besser kontrollierbaren KI-Systemen ebnen."]},{"key":"dc:title","label":"Title","values":["From local explanations to comprehensive mechanistic understanding of deep vision models"]}]}],"canonical_facts":{"dc:contributor.advisor":["Samek, Wojciech","Lapuschkin, Sebastian"],"dc:creator":["Dreyer, Maximilian"],"dc:date.accessioned":["2026-03-16T08:44:53Z"],"dc:date.available":["2026-03-16T08:44:53Z"],"dc:date.issued":["2026"],"dc:description.abstract":["Deep learning models have evolved into a cornerstone of modern industry and science, enabling applications from medical diagnosis and perception to conversational systems. Over the past two decades, both models and datasets have grown substantially in scale: models now reach trillions of parameters, trained on datasets with billions of samples. Despite their success, our understanding and control of these models remain limited, hindering safe and robust deployment in critical domains. Regulations such as the EU AI Act further emphasize the need for transparent and accountable AI systems. The field of eXplainable Artificial Intelligence (XAI) has introduced techniques like attribution maps and feature visualizations to illuminate singular aspects of model behavior. Yet, achieving a comprehensive understanding that enables validation and control of the complex mechanisms inside AI models requires the combination of multiple XAI perspectives. This is already challenging, and as most approaches rely on manual inspection of individual explanations, they fail to scale with the size and complexity of today’s models and datasets. This dissertation develops an explainability framework that is (i) mechanistic, by providing component-level insights (ii) comprehensive, by integrating multiple interpretability perspectives, (iii) scalable, by aggregating and summarizing explanations across data and model components while flagging outliers and deviations, and (iv) actionable, by directly informing practical strategies for refining and improving model behavior. Key contributions of this thesis include: (1) A foundation for comprehensive mechanistic explanations that integrate component-level attributions, input localizations, and feature visualizations, validated via a user study. (2) Measuring and improving interpretability by introducing multiple measures to estimate human interpretability of components, further validated through a user study, and methods to mitigate issues such as polysemanticity. (3) Prototypical Concept-based Explanations that summarize model behavior across entire datasets using a small set of concept-level prototypes. (4) Semantic component mbeddings that enable text-based semantic search, labeling, clustering, and comparison of model components. (5) Automated auditing methods such as outlier detection and concept alignment analysis to flag spurious or unexpected behaviors. (6) Interpretability-informed correction techniques that refine and correct models based on mechanistic insights. Through experiments on state-of-the-art vision models, this work demonstrates that mechanistic explanations enable the identification, understanding, and correction of model failures, providing a path toward more transparent, robust and controllable AI systems.","Tiefe neuronale Netze sind ein Eckpfeiler moderner Industrie und Wissenschaft, mit Anwendungen von der medizinischen Diagnose bis zu Konversationssystemen. Modelle und Datensätze sind in den letzten zwei Jahrzehnten stark gewachsen: Modelle erreichen Billionen von Parametern und werden auf Milliarden von Datenpunkten trainiert. Trotz ihres Erfolgs bleiben Verständnis und Kontrolle über diese Modelle begrenzt, was ihren sicheren Einsatz in kritischen Bereichen erschwert. Vorschriften wie das EU-KI-Gesetz betonen dabei die Notwendigkeit transparenter KI-Systeme. Techniken der erklärbaren künstlichen Intelligenz (XAI) wie Heatmaps und Konzeptvisualisierung beleuchten einzelne Aspekte des Modellverhaltens. Um jedoch ein umfassendes Verständnis zu erreichen, das die Validierung und Kontrolle komplexer Modellmechanismen ermöglicht, ist die Kombination mehrerer XAI-Perspektiven erforderlich. Diese erhöht bereits die kognitive Komplexität aufgrund der hohen Informationsdichte, und eine manuelle Überprüfung einzelner Erklärungen skaliert somit nicht mit der Größe aktueller Modelle und Datensätze. Diese Dissertation entwickelt ein Erklärbarkeits-Framework, das (i) mechanistische Einblicke auf Ebene einzelner Modellkomponenten liefert, (ii) umfassend mehrere Interpretierbarkeitsperspektiven integriert, (iii) skalierbar Erklärungen aggregiert und Ausreißer markiert, und (iv) handlungsorientiert Strategien zur Modellverbesserung bietet. Zentrale Beiträge dieser Arbeit sind: (1) Eine Grundlage für umfassende mechanistische Erklärungen, die komponentenspezifische Relevanzwerte, Lokalisierungen und Konzeptvisualisierungen integriert und in einer Benutzerstudie validiert wurde. (2) Messung und Verbesserung der Interpretierbarkeit durch Methoden zur Bewertung der menschlichen Interpretierbarkeit von Komponenten, ebenfalls in einer Benutzerstudie validiert, sowie zur Reduktion von Polysemantik. (3) Protoypische konzeptbasierte Erklärungen, die das Modellverhalten über Datensätze hinweg anhand weniger Konzept-basierten Prototypen zusammenfassen. (4) Semantische Komponenten-Einbettung, die textbasierte Suche, Beschreibungen, Gruppierung und Vergleiche von Modellkomponenten ermöglicht. (5) Automatisierte Prüfmethoden wie Ausreißererkennung und Konzeptabgleichanalyse zur Identifikation problematischer oder unerwarteter Verhaltensweisen. (6) Interpretierbarkeitsbasierte Korrekturtechniken, die Modelle anhand mechanistischer Einsichten gezielt korrigieren. Experimente mit modernen KI Modellen zeigen, dass mechanistische Erklärungen die Identifikation, das Verständnis und die Korrektur von Modellfehlern ermöglichen und damit den Weg zu transparenteren, robusteren und besser kontrollierbaren KI-Systemen ebnen."],"dc:identifier.uri":["https://depositonce.tu-berlin.de/handle/11303/26655","https://doi.org/10.14279/depositonce-25485"],"dc:language.iso":["en"],"dc:rights.uri":["https://creativecommons.org/licenses/by-sa/4.0/"],"dc:title":["From local explanations to comprehensive mechanistic understanding of deep vision models"],"dc:type":["Doctoral Thesis"]},"updated_at":"2026-07-27T21:28:44Z"}