Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 6 of 6 for “"AI Interpretability"”.

  1. Towards AI Safety via Interpretability and Oversight

    In this thesis, we advance AI safety through mechanistic interpretability and oversight methodologies across three key areas: mathematical reasoning in large language models (LLMs), the validity of sparse autoencoders, and scalable oversight. First, we reverse-engineer addition within mid-sized …

    mit Repository record for Towards AI Safety via Interpretability and Oversight (opens in a new tab)

  2. Probabilistic Programming for Postoperative Bleeding

    … existing research has focused on making medical AI systems interpretable to clinicians, we explore what it might take for clinicians to be authors of their own decision support tools. We focus specifically on the challenge of modelling uncertainty and study this in the context of postoperative …

    cambridge Repository record for Probabilistic Programming for Postoperative Bleeding (opens in a new tab)

  3. Predicting the Effectiveness of Medical Interventions

    … medical effectiveness. In Chapter 5, I argue against the prominent view that AI models are not explainable by showing how four familiar accounts of scientific explanation can be applied to neural networks. The confusion about explaining AI models is due to the conflation of ‘explainability’, …

    cambridge Repository record for Predicting the Effectiveness of Medical Interventions (opens in a new tab)

  4. Practical Diagnostic Tools for Deep Neural Networks

    The most common way to evaluate AI systems is by analyzing their performance on a test set. However, test sets can fail to identify some problems (such as out-of-distribution failures) and can actively reinforce others (such as dataset biases). Identifying problems like these requires techniques …

    mit Repository record for Practical Diagnostic Tools for Deep Neural Networks (opens in a new tab)

  5. Distributed AI-defense for Cyber Threats on Edge Computing Systems

    … cyber security scenarios and the lack of available representative datasets for threat detection. 2. State-of-the-art works presented in the literature focus on developing models via a centralized-learning approach. While most works show great effectiveness in detecting anomalies in network …

    tdl Repository record for Distributed AI-defense for Cyber Threats on Edge Computing Systems (opens in a new tab)

  6. Developing Interpretable Data Mining Frameworks for Addressing Biomedical Challenges in Genomics and Imaging

    … infections can progress to fatal multi-organ failure, while untreated DR may result in vision impairment or complete blindness. Given the impacts of these conditions on individual quality of life and healthcare systems worldwide, it is imperative to develop a comprehensive set of computational …

    iupui Repository record for Developing Interpretable Data Mining Frameworks for Addressing Biomedical Challenges in Genomics and Imaging (opens in a new tab)