Abstract
dc:descriptionArtificial intelligence (AI) has rapidly improved over the past decade, leading to widespread adoption of AI systems and demonstrating the potential for AI to greatly benefit society. However, as with any powerful new technology, AI introduces risks that must be managed to fully realize these benefits. Recent breakthroughs in the generality of AI systems have drawn increased attention to AI risks, including those of a potentially catastrophic nature. To help manage these anticipated risks, we take a defense in depth approach, combining different areas of AI safety research to address different aspects of AI risk. We present research on making AI systems more robust to adversarial influence, monitoring AIs for hidden behavior and trojans, enabling AIs to understand and adhere to human values, and finally addressing systemic problems to enable increased transparency.
Degree
thesis:*- Name thesis:degree_name
- Ph.D.
- Level thesis:degree_level
- Dissertation
- Discipline thesis:degree_discipline
- Computer Science
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2024
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Mazeika, Mantas
- Contributors dc:contributor
-
- Forsyth, David
- Li, Bo
- Lazebnik, Svetlana
- Krueger, David
Subjects
dc:subject × 8Rights
dc:rights- Statement dc:rights
-
- Copyright 2024 Mantas Mazeika
- Language dc:language
- en, eng
Identifiers
dc:identifier.*- Handle dc:identifier
- https://hdl.handle.net/2142/124375