Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 20 of 547 for “"Interpretability"”.
-
Machine Learning Interpretability in Malware Detection
… One of these challenges is the need for interpretability in machine learning. Machine learning interpretability is the process of giving explanations of a machine learning model's predictions to humans. Rather than trying to understand everything that is learnt by the model, it is an …
-
Interpretability of Neural Networks Latent Representations
… the models appear as black- boxes. The field of interpretability has developed with the aim to make these models more transparent by explaining their predictions. However, the field has historically focused on explaining the output of these models, leaving their internal latent representations …
-
Automated Mechanistic Interpretability for Neural Networks
Mechanistic interpretability research aims to deconstruct the underlying algorithms that neural networks use to perform computations, such that we can modify their components, causing them to change behavior in predictable and positive ways. This thesis details three novel methods for automating …
-
Contrasting contrastive and supervised models interpretability
… model using several deep neural network interpretability methods: network dissection, sparsity experiments, and saliency maps. Network dissections of self-supervised contrastive and supervised models show that the neurons of the contrastive model tend to learn about different parts of an …
-
Towards AI Safety via Interpretability and Oversight
… thesis, we advance AI safety through mechanistic interpretability and oversight methodologies across three key areas: mathematical reasoning in large language models (LLMs), the validity of sparse autoencoders, and scalable oversight. First, we reverse-engineer addition within mid-sized LLMs and …
-
Nonparametric High-dimensional Models: Sparsity, Efficiency, Interpretability
… enhance training efficiency, inference, and/or interpretability. In the first part, we consider additive models with interactions under component selection constraints and additional structural constraints e.g., hierarchical interactions. We consider different optimization based formulations and …
-
Enhancing Robustness and Interpretability in Computer Vision AI
L'abstract è presente nell'allegato / the abstract is in the attachment
-
Mechanistic Interpretability for Progress Towards Quantitative AI Safety
… program synthesis methods using mechanistic interpretability. We explore the robustness of LLMs through layer-level interventions such as zero-ablations and layer swapping, revealing that these models maintain high accuracy despite perturbations. As a result, we hypothesize the stages of …
-
ONLINE LEARNING, UNIFORM CONVERGENCE, AND A THEORY OF INTERPRETABILITY
… with existing lower bounds. Finally, regarding interpretability, we establish a taxonomy for approximating complex binary concepts with interpretable models such as shallow decision trees. Leveraging uniform convergence for Vapnik-Chervonenkis classes and von Neumann's minimax theorem, we …
-
Improving neuron-level interpretability with white-box language models
Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2027-05-01
-
Improving utilization, granularity, and interpretability in visual representation learning
… image datasets. The third work improves the interpretability of the deep model features on sparse image data. We integrate a decomposition method known as shift-invariant probabilistic latent component analysis (PLCA) into deep convolutional neural nets (CNNs). Hence we call our method Deep …
-
Deep concept reasoning: beyond the accuracy-interpretability trade-off
… in striking a balance between task accuracy and interpretability. Existing solutions either sacrifice model transparency for task accuracy or vice versa, making it difficult to optimize both objectives simultaneously. This thesis includes four research works contributing in addressing these …
-
Towards an Artificial Neuroscience: Analytics for Language Model Interpretability
The growing deployment of neural language models demands greater understanding of their internal mechanisms. The goal of this thesis is to make progress on understanding the latent computations within large language models (LLMs) to lay the groundwork for monitoring, controlling, and aligning …
-
Statistical learning for decision making : interpretability, uncertainty, and inference
… interpretable models from association rules. Interpretability is important for decision makers to understand why a prediction is made. First we show how linear mixtures of rules can be used to make sequential predictions. Then we develop Bayesian Rule Lists, a method for learning small, …
-
Techniques for Interpretability and Transparency of Black-Box Models
… factors have spurred research that increases the interpretability and transparency of these models. This thesis makes several contributions in this direction. We start with a concise but practical overview of the current techniques for defining and evaluating explanations for model predictions. …
-
Interpretability and Debugging for Distributed Privacy Preserving Machine Learning
… learning benefits from rich debugging and interpretability techniques enabled by transparent access to training data. However, FL removes this transparency, rendering traditional techniques ineffective and making debugging and interpretability a challenging open problem. This thesis …
-
Design and Interpretability of Contour Lines for Visualizing Multivariate Data
… stylizing contour lines, and investigate their interpretability using three crowdsourced studies. We evaluated the designs through a set of common geospatial data analysis tasks on a four-dimensional dataset. Our first two studies examined how the contour line width and the number of contour …
-
Interpretability by Design: New Interpretable Machine Learning Models and Methods
… important roles in many real-life scenarios, interpretability has become a key issue for whether we can trust the predictions made by these models, especially when we are making some high-stakes decisions. Lack of transparency has long been a concern for predictive models in criminal justice …
Page 1 of 28