{"id":{"repo_id":"cambridge","oai_identifier":"oai:www.repository.cam.ac.uk:1810/399323"},"canonical_url":"https://search.dev.ndltd.org/etd/cambridge/oai:www.repository.cam.ac.uk:1810/399323","repository":{"repo_id":"cambridge","name":"Cambridge University","base_url":"https://api.repository.cam.ac.uk/server/oai/request"},"display":{"title":"Help Wanted: Robust Concept Interventions for Interpretable Deep Neural Networks","abstract":"Artificial Intelligence (AI) systems are at their most powerful not when they replace humans, but when they collaborate with them. Yet, in critical domains such as healthcare and law, where expert knowledge is abundant, powerful AI systems driven by Deep Neural Networks (DNNs) operate under the rigid assumption that they cannot receive feedback once their training concludes. This myopic paradigm prevents models from soliciting or benefiting from external guidance when it matters most. In contrast, humans routinely consult one another, ask for help, and incorporate new evidence on the fly. This thesis asks: how can we design DNNs that capitalise on human feedback available during deployment? Work in explainable artificial intelligence (XAI) has taken a first step towards designing DNNs that support human-in-the-loop feedback via so-called concept interventions, which are operations where an expert communicates the presence or absence of a high-level concept to the model through a direct manipulation of its latent space. These interventions provide a practical approach to deploying models that remain transparent and interactive in high-stakes environments. However, the effectiveness of current intervenable models rests upon four unrealistic assumptions: (1) concept annotations are available for training; (2) those concepts are a complete description of the downstream task of interest; (3) all concept interventions are equally valuable; and (4) test samples stay within the training distribution. This thesis shows that such assumptions do not hold in real-world scenarios and argues that, when they are violated, interventions become ineffective. To address this, we introduce a series of methods that make concept interventions robust to conditions faced during real-world deployment. First, by discovering simple functions over small feature subsets that can explain a tabular task of interest, we show how to perform interventions in tabular domains that lack training concept labels. Second, we demonstrate that interventions may backfire when important concepts are missing during training, and introduce Concept Embedding Models (CEMs) as a solution to this problem. CEMs learn high-dimensional, interpretable concept representations and use them to preserve intervenability even when trained with incomplete concept sets. Third, we relax the assumption that all concepts are equally valuable and propose an intervention-aware training paradigm that teaches CEMs to prioritise requesting specific concepts from experts, reducing the amount of help needed in budget-constrained setups. Finally, we extend this framework to handle out-of-distribution test samples, proposing a decomposition of concept embeddings into sample-specific and concept-specific components that preserves intervention robustness under distribution shifts. Overall, the methodologies proposed in this thesis provide a principled approach for designing DNNs that are accurate, interpretable, and capable of significantly increasing their accuracy when experts can provide test-time feedback.","abstract_html":"Artificial Intelligence (AI) systems are at their most powerful not when they replace humans, but when they collaborate with them. Yet, in critical domains such as healthcare and law, where expert knowledge is abundant, powerful AI systems driven by Deep Neural Networks (DNNs) operate under the rigid assumption that they cannot receive feedback once their training concludes. This myopic paradigm prevents models from soliciting or benefiting from external guidance when it matters most. In contrast, humans routinely consult one another, ask for help, and incorporate new evidence on the fly. This thesis asks: how can we design DNNs that capitalise on human feedback available during deployment? Work in explainable artificial intelligence (XAI) has taken a first step towards designing DNNs that support human-in-the-loop feedback via so-called concept interventions, which are operations where an expert communicates the presence or absence of a high-level concept to the model through a direct manipulation of its latent space. These interventions provide a practical approach to deploying models that remain transparent and interactive in high-stakes environments. However, the effectiveness of current intervenable models rests upon four unrealistic assumptions: (1) concept annotations are available for training; (2) those concepts are a complete description of the downstream task of interest; (3) all concept interventions are equally valuable; and (4) test samples stay within the training distribution. This thesis shows that such assumptions do not hold in real-world scenarios and argues that, when they are violated, interventions become ineffective. To address this, we introduce a series of methods that make concept interventions robust to conditions faced during real-world deployment. First, by discovering simple functions over small feature subsets that can explain a tabular task of interest, we show how to perform interventions in tabular domains that lack training concept labels. Second, we demonstrate that interventions may backfire when important concepts are missing during training, and introduce Concept Embedding Models (CEMs) as a solution to this problem. CEMs learn high-dimensional, interpretable concept representations and use them to preserve intervenability even when trained with incomplete concept sets. Third, we relax the assumption that all concepts are equally valuable and propose an intervention-aware training paradigm that teaches CEMs to prioritise requesting specific concepts from experts, reducing the amount of help needed in budget-constrained setups. Finally, we extend this framework to handle out-of-distribution test samples, proposing a decomposition of concept embeddings into sample-specific and concept-specific components that preserves intervention robustness under distribution shifts. Overall, the methodologies proposed in this thesis provide a principled approach for designing DNNs that are accurate, interpretable, and capable of significantly increasing their accuracy when experts can provide test-time feedback.","abstract_has_math":false,"creators":["Espinosa Zarlenga, Mateo"],"institution":"University of Cambridge","degree_name":"Doctor of Philosophy (PhD)","degree_level":"Doctoral","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Jamnik, Mateja"],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-07-25","date_published":"2025-07-25","updated_at":"2026-07-22T22:24:16Z","subjects":["Computer Science","Machine Learning","Representation Learning","Explainable AI","Artificial Intelligence","AI","Concept Interventions"],"languages":["eng"],"rights":[],"rights_urls":["https://www.repository.cam.ac.uk/bitstreams/dea3f386-d06c-4a92-9ce5-f8b78f540a1f/download","https://creativecommons.org/licenses/by/4.0/"],"identifier_entries":[{"key":"dc:creator.authoridentifier","label":"Author Identifier","values":["0009000673335727"],"render_values":[{"text":"0009-0006-7333-5727","href":"https://orcid.org/0009-0006-7333-5727","code":true}]}]},"links":{"outbound_url":"https://doi.org/10.17863/CAM.127898","outbound_label":"DOI","outbound_source":"dc:identifier.doi"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Jamnik, Mateja"]},{"key":"dc:contributor.sponsor","label":"Sponsor","values":["Gates Cambridge Scholarship"]},{"key":"dc:creator","label":"Author","values":["Espinosa Zarlenga, Mateo"]},{"key":"dc:creator.authoridentifier","label":"Author Identifier","values":["0009000673335727"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2025-07-25"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["University of Cambridge"]},{"key":"dc:relation.isreferencedby.uri","label":"Dc Relation Isreferencedby URI","values":["https://www.repository.cam.ac.uk/handle/1810/399323"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["Doctoral"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["Doctor of Philosophy (PhD)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer Science","Machine Learning","Representation Learning","Explainable AI","Artificial Intelligence","AI","Concept Interventions"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["https://www.repository.cam.ac.uk/bitstreams/dea3f386-d06c-4a92-9ce5-f8b78f540a1f/download","https://creativecommons.org/licenses/by/4.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.17863/CAM.127898"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://www.repository.cam.ac.uk/bitstreams/be1df77d-5eb4-48ca-8489-ed5a304f5602/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Artificial Intelligence (AI) systems are at their most powerful not when they replace humans, but when they collaborate with them. Yet, in critical domains such as healthcare and law, where expert knowledge is abundant, powerful AI systems driven by Deep Neural Networks (DNNs) operate under the rigid assumption that they cannot receive feedback once their training concludes. This myopic paradigm prevents models from soliciting or benefiting from external guidance when it matters most. In contrast, humans routinely consult one another, ask for help, and incorporate new evidence on the fly. This thesis asks: how can we design DNNs that capitalise on human feedback available during deployment? Work in explainable artificial intelligence (XAI) has taken a first step towards designing DNNs that support human-in-the-loop feedback via so-called concept interventions, which are operations where an expert communicates the presence or absence of a high-level concept to the model through a direct manipulation of its latent space. These interventions provide a practical approach to deploying models that remain transparent and interactive in high-stakes environments. However, the effectiveness of current intervenable models rests upon four unrealistic assumptions: (1) concept annotations are available for training; (2) those concepts are a complete description of the downstream task of interest; (3) all concept interventions are equally valuable; and (4) test samples stay within the training distribution. This thesis shows that such assumptions do not hold in real-world scenarios and argues that, when they are violated, interventions become ineffective. To address this, we introduce a series of methods that make concept interventions robust to conditions faced during real-world deployment. First, by discovering simple functions over small feature subsets that can explain a tabular task of interest, we show how to perform interventions in tabular domains that lack training concept labels. Second, we demonstrate that interventions may backfire when important concepts are missing during training, and introduce Concept Embedding Models (CEMs) as a solution to this problem. CEMs learn high-dimensional, interpretable concept representations and use them to preserve intervenability even when trained with incomplete concept sets. Third, we relax the assumption that all concepts are equally valuable and propose an intervention-aware training paradigm that teaches CEMs to prioritise requesting specific concepts from experts, reducing the amount of help needed in budget-constrained setups. Finally, we extend this framework to handle out-of-distribution test samples, proposing a decomposition of concept embeddings into sample-specific and concept-specific components that preserves intervention robustness under distribution shifts. Overall, the methodologies proposed in this thesis provide a principled approach for designing DNNs that are accurate, interpretable, and capable of significantly increasing their accuracy when experts can provide test-time feedback."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["c937b5bae1ddd2046cc6f434a465675d","87eda9de84448d1f82354d60eee3eb5f"]},{"key":"dc:title","label":"Title","values":["Help Wanted: Robust Concept Interventions for Interpretable Deep Neural Networks"]}]}],"canonical_facts":{"dc:contributor.advisor":["Jamnik, Mateja"],"dc:contributor.sponsor":["Gates Cambridge Scholarship"],"dc:creator":["Espinosa Zarlenga, Mateo"],"dc:creator.authoridentifier":["0009000673335727"],"dc:date.issued":["2025-07-25"],"dc:description.abstract":["Artificial Intelligence (AI) systems are at their most powerful not when they replace humans, but when they collaborate with them. Yet, in critical domains such as healthcare and law, where expert knowledge is abundant, powerful AI systems driven by Deep Neural Networks (DNNs) operate under the rigid assumption that they cannot receive feedback once their training concludes. This myopic paradigm prevents models from soliciting or benefiting from external guidance when it matters most. In contrast, humans routinely consult one another, ask for help, and incorporate new evidence on the fly. This thesis asks: how can we design DNNs that capitalise on human feedback available during deployment? Work in explainable artificial intelligence (XAI) has taken a first step towards designing DNNs that support human-in-the-loop feedback via so-called concept interventions, which are operations where an expert communicates the presence or absence of a high-level concept to the model through a direct manipulation of its latent space. These interventions provide a practical approach to deploying models that remain transparent and interactive in high-stakes environments. However, the effectiveness of current intervenable models rests upon four unrealistic assumptions: (1) concept annotations are available for training; (2) those concepts are a complete description of the downstream task of interest; (3) all concept interventions are equally valuable; and (4) test samples stay within the training distribution. This thesis shows that such assumptions do not hold in real-world scenarios and argues that, when they are violated, interventions become ineffective. To address this, we introduce a series of methods that make concept interventions robust to conditions faced during real-world deployment. First, by discovering simple functions over small feature subsets that can explain a tabular task of interest, we show how to perform interventions in tabular domains that lack training concept labels. Second, we demonstrate that interventions may backfire when important concepts are missing during training, and introduce Concept Embedding Models (CEMs) as a solution to this problem. CEMs learn high-dimensional, interpretable concept representations and use them to preserve intervenability even when trained with incomplete concept sets. Third, we relax the assumption that all concepts are equally valuable and propose an intervention-aware training paradigm that teaches CEMs to prioritise requesting specific concepts from experts, reducing the amount of help needed in budget-constrained setups. Finally, we extend this framework to handle out-of-distribution test samples, proposing a decomposition of concept embeddings into sample-specific and concept-specific components that preserves intervention robustness under distribution shifts. Overall, the methodologies proposed in this thesis provide a principled approach for designing DNNs that are accurate, interpretable, and capable of significantly increasing their accuracy when experts can provide test-time feedback."],"dc:format.checksum.md5":["c937b5bae1ddd2046cc6f434a465675d","87eda9de84448d1f82354d60eee3eb5f"],"dc:identifier.doi":["https://doi.org/10.17863/CAM.127898"],"dc:identifier.uri":["https://www.repository.cam.ac.uk/bitstreams/be1df77d-5eb4-48ca-8489-ed5a304f5602/download"],"dc:language":["eng"],"dc:publisher.institution":["University of Cambridge"],"dc:relation.isreferencedby.uri":["https://www.repository.cam.ac.uk/handle/1810/399323"],"dc:rights":["https://www.repository.cam.ac.uk/bitstreams/dea3f386-d06c-4a92-9ce5-f8b78f540a1f/download","https://creativecommons.org/licenses/by/4.0/"],"dc:subject":["Computer Science","Machine Learning","Representation Learning","Explainable AI","Artificial Intelligence","AI","Concept Interventions"],"dc:title":["Help Wanted: Robust Concept Interventions for Interpretable Deep Neural Networks"],"dc:type":["Thesis"],"dc:type.qualificationlevel":["Doctoral"],"dc:type.qualificationname":["Doctor of Philosophy (PhD)"]},"updated_at":"2026-07-22T22:24:16Z"}