Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 16 of 16 for “"AI Safety"”.
-
Towards AI Safety via Interpretability and Oversight
In this thesis, we advance AI safety through mechanistic interpretability and oversight methodologies across three key areas: mathematical reasoning in large language models (LLMs), the validity of sparse autoencoders, and scalable oversight. First, we reverse-engineer addition within mid-sized …
-
Mechanistic Interpretability for Progress Towards Quantitative AI Safety
In this thesis, we conduct a detailed investigation into the dynamics of neural networks, focusing on two key areas: inference stages in large language models (LLMs) and novel program synthesis methods using mechanistic interpretability. We explore the robustness of LLMs through layer-level …
-
Learning to Improve Clinical Decisions and AI Safety by Leveraging Structure
The availability of large collections of digitized healthcare data along with the increasing power of computation has allowed machine learning (ML) for healthcare to become one of the key applied research domains in ML. ML for health has great potential in providing clinical decision-making support …
-
Advancing Information Extraction with Large Language Models: The Role of Structured Understanding in Knowledge Management and AI Safety
… an invaluable source of knowledge, yet it remains challenging for machines to interpret and transform into actionable insights. Information Extraction (IE) offers a promising solution by converting raw text into machine-readable representations. However, traditional IE systems are limited by …
-
Private, Verifiable, and Auditable AI Systems
… verifiability, and auditability in modern AI, particularly in foundation models. It argues that technical solutions that integrate these elements are critical for responsible AI innovation. Drawing from international policy contributions and technical research to identify key risks in the …
-
Advancing Cross-Domain Fake News Detection: Enhanced Models to Improve Generalization and Tackle the Class Imbalance Problem
The rapid proliferation of fake news across domains, such as politics, health, and social media, poses a significant threat to the integrity of information dissemination, leading to misinformation that can affect public perception and decision-making. Detecting fake news is critical to preserving …
-
Evolving Threats and Defenses in Machine Learning: Focus on Model Inversion and Beyond
… into critical real-world applications, raising concerns about security, privacy, and trustworthiness. Among various emerging threats, model inversion (MI) attacks stand out due to their potential to compromise the confidentiality of training data. This dissertation investigates evolving …
-
Detection and mitigation of misbehaviour in LLMs
Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms
-
Nothing to See Here: Generative AI, Neoliberal Crisis, and Intensified Counterinsurgency
Examining generative artificial intelligence (genAI) as a discursive and technical agent deployed in response to the ongoing crisis of neoliberal legitimacy, this dissertation advances the claim that so-called “safe” models extend the counterinsurgent (COIN) mode of governance, both as an effect of …
-
Building Reliable AI under Distribution Shifts
… where distribution shifts—differences between training and deployment data—can significantly impact their reliability. These shifts affect models in multiple ways, leading to degraded generalization, fairness collapse, loss of robustness, and new safety vulnerabilities. This dissertation …
-
Table manners at scale: introducing Lora Lens for efficient analysis & detoxification of language models
… toxicity without requiring complete model retraining. Despite its effectiveness, Attention Lens faces severe limitations in scalability and computational efficiency. To overcome these limitations, we propose and implement the Lora Lens, an innovative adaptation of Attention Lens employing …
-
Understanding and Mitigating Data-Centric Vulnerabilities in Modern AI Systems
Modern artificial intelligence (AI) systems, trained on vast internet-scale datasets, demonstrate remarkable performance and emergent capabilities. However, this reliance on large datasets that are expensive or difficult to quality-control exposes AI systems to critical vulnerabilities, including …
-
Adversarial Déjà Vu: Jailbreak Dictionary Learning for Stronger Generalization to Unseen Attacks
… across a wide range of tasks, yet they remain vulnerable to jailbreak attacks that bypass safety guardrails and elicit harmful behavior. Defending against such attacks is particularly challenging when adversaries introduce novel strategies that differ from those observed during training or …
-
Crafting safe human-centric agents with risk intelligence
Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms
-
Toward managing catastrophic AI risks
Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-09-16 without embargo terms
-
Artificial Intelligence in Human Spaceflight Safety-Critical Systems: A Requirements Framework for AI-Enabled Computer-Based Control Systems
… standard that addresses the safe integration of AI into computer-based control systems (CBCS) in human spaceflight. The computer-based control expectations of those safety-critical systems on the International Space Station (ISS) are captured in SSP 50038, Computer-Based Control System Safety …