Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 16 of 16 for “"AI Safety"”.

  1. Towards AI Safety via Interpretability and Oversight

    In this thesis, we advance AI safety through mechanistic interpretability and oversight methodologies across three key areas: mathematical reasoning in large language models (LLMs), the validity of sparse autoencoders, and scalable oversight. First, we reverse-engineer addition within mid-sized …

    mit Repository record for Towards AI Safety via Interpretability and Oversight (opens in a new tab)

  2. Mechanistic Interpretability for Progress Towards Quantitative AI Safety

    In this thesis, we conduct a detailed investigation into the dynamics of neural networks, focusing on two key areas: inference stages in large language models (LLMs) and novel program synthesis methods using mechanistic interpretability. We explore the robustness of LLMs through layer-level …

    mit Repository record for Mechanistic Interpretability for Progress Towards Quantitative AI Safety (opens in a new tab)

  3. Learning to Improve Clinical Decisions and AI Safety by Leveraging Structure

    The availability of large collections of digitized healthcare data along with the increasing power of computation has allowed machine learning (ML) for healthcare to become one of the key applied research domains in ML. ML for health has great potential in providing clinical decision-making support …

    mit Repository record for Learning to Improve Clinical Decisions and AI Safety by Leveraging Structure (opens in a new tab)

  4. Advancing Information Extraction with Large Language Models: The Role of Structured Understanding in Knowledge Management and AI Safety

    … an invaluable source of knowledge, yet it remains challenging for machines to interpret and transform into actionable insights. Information Extraction (IE) offers a promising solution by converting raw text into machine-readable representations. However, traditional IE systems are limited by …

    cagliari Repository record for Advancing Information Extraction with Large Language Models: The Role of Structured Understanding in Knowledge Management and AI Safety (opens in a new tab)

  5. Private, Verifiable, and Auditable AI Systems

    … verifiability, and auditability in modern AI, particularly in foundation models. It argues that technical solutions that integrate these elements are critical for responsible AI innovation. Drawing from international policy contributions and technical research to identify key risks in the …

    mit Repository record for Private, Verifiable, and Auditable AI Systems (opens in a new tab)

  6. Advancing Cross-Domain Fake News Detection: Enhanced Models to Improve Generalization and Tackle the Class Imbalance Problem

    The rapid proliferation of fake news across domains, such as politics, health, and social media, poses a significant threat to the integrity of information dissemination, leading to misinformation that can affect public perception and decision-making. Detecting fake news is critical to preserving …

    ottawa-retro Repository record for Advancing Cross-Domain Fake News Detection: Enhanced Models to Improve Generalization and Tackle the Class Imbalance Problem (opens in a new tab)

  7. Evolving Threats and Defenses in Machine Learning: Focus on Model Inversion and Beyond

    … into critical real-world applications, raising concerns about security, privacy, and trustworthiness. Among various emerging threats, model inversion (MI) attacks stand out due to their potential to compromise the confidentiality of training data. This dissertation investigates evolving …

    vt Repository record for Evolving Threats and Defenses in Machine Learning: Focus on Model Inversion and Beyond (opens in a new tab)

  8. Detection and mitigation of misbehaviour in LLMs

    Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms

    uiuc Repository record for Detection and mitigation of misbehaviour in LLMs (opens in a new tab)

  9. Nothing to See Here: Generative AI, Neoliberal Crisis, and Intensified Counterinsurgency

    Examining generative artificial intelligence (genAI) as a discursive and technical agent deployed in response to the ongoing crisis of neoliberal legitimacy, this dissertation advances the claim that so-called “safe” models extend the counterinsurgent (COIN) mode of governance, both as an effect of …

    york Repository record for Nothing to See Here: Generative AI, Neoliberal Crisis, and Intensified Counterinsurgency (opens in a new tab)

  10. Building Reliable AI under Distribution Shifts

    … where distribution shifts—differences between training and deployment data—can significantly impact their reliability. These shifts affect models in multiple ways, leading to degraded generalization, fairness collapse, loss of robustness, and new safety vulnerabilities. This dissertation …

    maryland Repository record for Building Reliable AI under Distribution Shifts (opens in a new tab)

  11. Table manners at scale: introducing Lora Lens for efficient analysis & detoxification of language models

    … toxicity without requiring complete model retraining. Despite its effectiveness, Attention Lens faces severe limitations in scalability and computational efficiency. To overcome these limitations, we propose and implement the Lora Lens, an innovative adaptation of Attention Lens employing …

    colo-mines Repository record for Table manners at scale: introducing Lora Lens for efficient analysis & detoxification of language models (opens in a new tab)

  12. Understanding and Mitigating Data-Centric Vulnerabilities in Modern AI Systems

    Modern artificial intelligence (AI) systems, trained on vast internet-scale datasets, demonstrate remarkable performance and emergent capabilities. However, this reliance on large datasets that are expensive or difficult to quality-control exposes AI systems to critical vulnerabilities, including …

    vt Repository record for Understanding and Mitigating Data-Centric Vulnerabilities in Modern AI Systems (opens in a new tab)

  13. Adversarial Déjà Vu: Jailbreak Dictionary Learning for Stronger Generalization to Unseen Attacks

    … across a wide range of tasks, yet they remain vulnerable to jailbreak attacks that bypass safety guardrails and elicit harmful behavior. Defending against such attacks is particularly challenging when adversaries introduce novel strategies that differ from those observed during training or …

    vt Repository record for Adversarial Déjà Vu: Jailbreak Dictionary Learning for Stronger Generalization to Unseen Attacks (opens in a new tab)

  14. Crafting safe human-centric agents with risk intelligence

    Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms

    uiuc Repository record for Crafting safe human-centric agents with risk intelligence (opens in a new tab)

  15. Toward managing catastrophic AI risks

    Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-09-16 without embargo terms

    uiuc Repository record for Toward managing catastrophic AI risks (opens in a new tab)

  16. Artificial Intelligence in Human Spaceflight Safety-Critical Systems: A Requirements Framework for AI-Enabled Computer-Based Control Systems

    … standard that addresses the safe integration of AI into computer-based control systems (CBCS) in human spaceflight. The computer-based control expectations of those safety-critical systems on the International Space Station (ISS) are captured in SSP 50038, Computer-Based Control System Safety

    embry-riddle Repository record for Artificial Intelligence in Human Spaceflight Safety-Critical Systems: A Requirements Framework for AI-Enabled Computer-Based Control Systems (opens in a new tab)