Back to results

University of Cambridge

Help Wanted: Robust Concept Interventions for Interpretable Deep Neural Networks

Abstract

dc:description.abstract

Artificial Intelligence (AI) systems are at their most powerful not when they replace humans, but when they collaborate with them. Yet, in critical domains such as healthcare and law, where expert knowledge is abundant, powerful AI systems driven by Deep Neural Networks (DNNs) operate under the rigid assumption that they cannot receive feedback once their training concludes. This myopic paradigm prevents models from soliciting or benefiting from external guidance when it matters most. In contrast, humans routinely consult one another, ask for help, and incorporate new evidence on the fly. This thesis asks: how can we design DNNs that capitalise on human feedback available during deployment? Work in explainable artificial intelligence (XAI) has taken a first step towards designing DNNs that support human-in-the-loop feedback via so-called concept interventions, which are operations where an expert communicates the presence or absence of a high-level concept to the model through a direct manipulation of its latent space. These interventions provide a practical approach to deploying models that remain transparent and interactive in high-stakes environments. However, the effectiveness of current intervenable models rests upon four unrealistic assumptions: (1) concept annotations are available for training; (2) those concepts are a complete description of the downstream task of interest; (3) all concept interventions are equally valuable; and (4) test samples stay within the training distribution. This thesis shows that such assumptions do not hold in real-world scenarios and argues that, when they are violated, interventions become ineffective. To address this, we introduce a series of methods that make concept interventions robust to conditions faced during real-world deployment. First, by discovering simple functions over small feature subsets that can explain a tabular task of interest, we show how to perform interventions in tabular domains that lack training concept labels. Second, we demonstrate that interventions may backfire when important concepts are missing during training, and introduce Concept Embedding Models (CEMs) as a solution to this problem. CEMs learn high-dimensional, interpretable concept representations and use them to preserve intervenability even when trained with incomplete concept sets. Third, we relax the assumption that all concepts are equally valuable and propose an intervention-aware training paradigm that teaches CEMs to prioritise requesting specific concepts from experts, reducing the amount of help needed in budget-constrained setups. Finally, we extend this framework to handle out-of-distribution test samples, proposing a decomposition of concept embeddings into sample-specific and concept-specific components that preserves intervention robustness under distribution shifts. Overall, the methodologies proposed in this thesis provide a principled approach for designing DNNs that are accurate, interpretable, and capable of significantly increasing their accuracy when experts can provide test-time feedback.

Degree

thesis:*
Name dc:type.qualificationname
Doctor of Philosophy (PhD)
Level dc:type.qualificationlevel
Doctoral
Grantor dc:publisher.institution
University of Cambridge
Year dc:date.issued
2025

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Espinosa Zarlenga, Mateo
Advisor dc:contributor.advisor
  • Jamnik, Mateja

Subjects

dc:subject × 7

Rights

dc:rights
Language dc:language
eng

Identifiers

dc:identifier.*
Author Identifier
0009-0006-7333-5727
OAI identifier oai:identifier
oai:www.repository.cam.ac.uk:1810/399323

Chain of custody

source
Harvested from
Cambridge University
Base URL
api.repository.cam.ac.uk/server/oai/request
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Espinosa Zarlenga, Mateo. Help Wanted: Robust Concept Interventions for Interpretable Deep Neural Networks. Doctoral thesis, University of Cambridge, 2025. https://doi.org/10.17863/CAM.127898