University of Minnesota
Causal and system-theoretic approaches to interpretable machine learning
Abstract
dc:description.abstractWhile machine learning algorithms are achieving increasingly impressive levels of decision-making and predictive performance, the way such algorithms operate is also becoming more and more inscrutable. As more decisions are delegated to complex systems that are not susceptible to any form of human supervision or scrutiny, it is only natural toquestion their fairness, validity, and trustworthiness. These concerns are amplified in scientific domains where data are scarce, biased toward specific conditions, or difficult to obtain due to logistical or ethical constraints. In such settings, machine learning models frequently fail to generalize and may even produce results that contradict well-established scientific principles. The lack of interpretability not only undermines trust but also hinders our ability to diagnose, debug, and improve these models. This dissertation addresses these challenges by exploring both post-hoc explanation methods and augmentation based strategies for increasing interpretability. We begin by examining existing explainable AI (XAI) methods, which aim to attribute a model’s output to its input features. Focusing on the widely used Local Interpretable Model-Agnostic Explanations (LIME), we adapt this technique for the time series setting and demonstrate how it can be extended through tools from causal inference to reveal cause-effect relationships among input and output variables. We next consider Shapley values, another prominent XAI method that has demonstrated the ability to generate coherent explanations across different types of data, including time series data. We find that the computation of Shapley values has striking similarities to causal discovery algorithms and only differ in the fact that Shapley values averages marginal contributions over all possible feature orderings, while causal discovery techniques typically select orderings that yield the sparsest graphs. We reconcile these approaches by interpreting them through the lens of maximum entropy: Shapley values operate under the assumption that no prior knowledge exists to justify preferring one variable ordering over another. In contrast, causal discovery methods rely on assumptions about the data’s probability distribution, which guide them toward favoring variable orderings that yield simpler graphs. This reinterpretation allows us to develop new, theoretically grounded variations of Shapley values that can incorporate prior knowledge into the explanation process. In the second part of the dissertation, we shift focus to model augmentation as an alternative route to interpretability. In contrast to post-hoc methods, model augmentation is primarily driven by the goal of enhancing generalizability in scenarios with limited data but available domain knowledge, rather than aiming for interpretability as its central objective. These techniques integrate structured prior knowledge, such as physical laws, into the learning process by combining a primary model with an auxiliary one. We introduce a novel augmentation framework, called the Tutor-Pupil scheme, designed to enhance both interpretability and predictive accuracy. In this architecture, the Pupil is a fixed, interpretable model dedicated to the core task, while the Tutor is a flexible auxiliary model trained to apply minimal input-level corrections that improve the Pupil’s performance. By analyzing the Tutor’s interventions, we gain insight into where the Pupil fails to generalize, identify latent structure in the data, and expose higher-order patterns not captured by the original model.
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Biparva, Darya
Subjects
dc:subject × 3Rights
- Language dc:language.iso
- en
Identifiers
dc:identifier.*- Handle dc:identifier.uri
- https://hdl.handle.net/11299/278137
- OAI identifier oai:identifier
- oai:conservancy.umn.edu:11299/278137