{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/124375"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/124375","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Toward managing catastrophic AI risks","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-09-16 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2024-09-16 without embargo terms","abstract_has_math":false,"creators":["Mazeika, Mantas"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Forsyth, David","Li, Bo","Lazebnik, Svetlana","Krueger, David"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-05","date_published":"2024-05","updated_at":"2026-07-22T22:25:00Z","subjects":["Ai Safety","Ai Risk","Robustness","Red Teaming","Neural Trojans","Trojan Detection","Alignment","Model Stealing"],"languages":["en","eng"],"rights":["Copyright 2024 Mantas Mazeika"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/124375","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Forsyth, David","Li, Bo","Lazebnik, Svetlana","Krueger, David"]},{"key":"dc:creator","label":"Author","values":["Mazeika, Mantas"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2024-05","2024-04-24"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Ai Safety","Ai Risk","Robustness","Red Teaming","Neural Trojans","Trojan Detection","Alignment","Model Stealing"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2024 Mantas Mazeika"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/124375"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-09-16 without embargo terms","The student, Mantas Mazeika, accepted the attached license on 2024-04-23 at 09:49.","The student, Mantas Mazeika, submitted this Dissertation for approval on 2024-04-23 at 09:57.","This Dissertation was approved for publication on 2024-04-24 at 15:16.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20566 on 2024-09-16 at 00:35:51","Artificial intelligence (AI) has rapidly improved over the past decade, leading to widespread adoption of AI systems and demonstrating the potential for AI to greatly benefit society. However, as with any powerful new technology, AI introduces risks that must be managed to fully realize these benefits. Recent breakthroughs in the generality of AI systems have drawn increased attention to AI risks, including those of a potentially catastrophic nature. To help manage these anticipated risks, we take a defense in depth approach, combining different areas of AI safety research to address different aspects of AI risk. We present research on making AI systems more robust to adversarial influence, monitoring AIs for hidden behavior and trojans, enabling AIs to understand and adhere to human values, and finally addressing systemic problems to enable increased transparency."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Toward managing catastrophic AI risks"]}]}],"canonical_facts":{"dc:contributor":["Forsyth, David","Li, Bo","Lazebnik, Svetlana","Krueger, David"],"dc:creator":["Mazeika, Mantas"],"dc:date":["2024-05","2024-04-24"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-09-16 without embargo terms","The student, Mantas Mazeika, accepted the attached license on 2024-04-23 at 09:49.","The student, Mantas Mazeika, submitted this Dissertation for approval on 2024-04-23 at 09:57.","This Dissertation was approved for publication on 2024-04-24 at 15:16.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20566 on 2024-09-16 at 00:35:51","Artificial intelligence (AI) has rapidly improved over the past decade, leading to widespread adoption of AI systems and demonstrating the potential for AI to greatly benefit society. However, as with any powerful new technology, AI introduces risks that must be managed to fully realize these benefits. Recent breakthroughs in the generality of AI systems have drawn increased attention to AI risks, including those of a potentially catastrophic nature. To help manage these anticipated risks, we take a defense in depth approach, combining different areas of AI safety research to address different aspects of AI risk. We present research on making AI systems more robust to adversarial influence, monitoring AIs for hidden behavior and trojans, enabling AIs to understand and adhere to human values, and finally addressing systemic problems to enable increased transparency."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/124375"],"dc:language":["en","eng"],"dc:rights":["Copyright 2024 Mantas Mazeika"],"dc:subject":["Ai Safety","Ai Risk","Robustness","Red Teaming","Neural Trojans","Trojan Detection","Alignment","Model Stealing"],"dc:title":["Toward managing catastrophic AI risks"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:00Z"}