Back to results

University of Toronto

Understanding Adversarial Robustness in Deep Learning

Abstract

dc:description.abstract

This thesis studies the adversarial robustness of deep learning models. Our investigation covers various aspects of this phenomenon, including the development of two new defense algorithms, two new attack algorithms, a novel definition of hierarchical adversarial robustness, and an analysis of how the optimization process affects model robustness. We begin by introducing the Second-Order Adversarial Regularizer (SOAR) as a defense strategy to improve model robustness. Unlike traditional data augmentation approaches that rely on computationally expensive algorithms to generate adversarially perturbed training samples, we derive a regularizer that mimics the effect of data augmentation, eliminating the need for it during training. We empirically demonstrate the improved adversarial robustness of SOAR-regularized models against white-box and transferred perturbations. In practice, many machine learning systems still forgo robustification techniques due to the additional computational overhead. This motivates our investigation into the robustness of models obtained through standard training regimes. Focusing on the optimization process, we examine the robustness of models trained with different algorithms. As we will see, models trained using stochastic gradient descent exhibit far greater robustness compared to those trained with adaptive gradient methods, such as Adam and RMSProp. Through a frequency-domain analysis, we discover that specific properties of datasets, seemingly irrelevant to generalization, can actually result in vulnerabilities when models are trained using certain optimizers. These insights underscore the importance of considering both the optimization strategy and dataset characteristics in improving model robustness. Extending our exploration of dataset properties, we recognize that as datasets grow in size and complexity, the number of classes and their hierarchical relationships become increasingly significant. However, current ways of evaluating adversarial robustness, which treat all misclassifications equally, risk overestimating robustness or underestimating attack effectiveness. To address this, we introduce the concept of hierarchical adversarial robustness. For datasets with hierarchical structures, we define hierarchical adversarial examples as those that lead to misclassifications at the meta-class level. Building on this, we develop an attack algorithm designed to generate such examples and propose an architectural solution to improve models' robustness against them. Finally, we take on the role of an adversary. Current robustification strategies predominantly involve replacing the training samples with their adversarially perturbed counterparts, highlighting the importance of understanding and improving adversarial example generation. Therefore, in our last work, we focus on improving the transferability of perturbations. We develop a fine-tuning method called model alignment to transform any model into one from which any attack algorithms can generate more transferable perturbations.

Degree

thesis:*
Department dc:contributor.department
Computer Science
Year dc:date.issued
2025

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Ma, Bojie
Advisors dc:contributor.advisor
  • Farahmand, Amir-massoud
  • Zemel, Rich

Rights

dc:rights
Statement dc:rights
  • Attribution 4.0 International

Identifiers

dc:identifier.*
Handle dc:identifier.uri
https://hdl.handle.net/1807/149895
OAI identifier oai:identifier
oai:utoronto.scholaris.ca:1807/149895

Chain of custody

source
Harvested from
University of Toronto
Base URL
utoronto.scholaris.ca/server/oai/request
Last updated
2026-07-27
Source record
OAI-PMH GetRecord
related terms
citation

Ma, Bojie. Understanding Adversarial Robustness in Deep Learning. 2025. https://hdl.handle.net/1807/149895