University of Illinois at Urbana-Champaign
The implicit bias of gradient descent: from linear classifiers to deep networks
Abstract
dc:descriptionGradient descent and other first-order methods have been extensively used in deep learning. They can find solutions with not only low training error, but also low test error. This motivates the study of the implicit bias of training algorithms such as gradient descent, meaning we want to understand what special properties are satisfied by gradient descent solutions that lead to good generalization. In this thesis, we focus on gradient descent and its derivatives, show test error bounds and characterize the implicit biases. In detail, for linear classifiers, we consider both the separable setting and nonseparable setting. In the separable case, we first give an 1/t test error bound for stochastic gradient descent. Then for gradient descent, we present a primal-dual framework to analyze its implicit bias, which leads to fast margin maximization algorithms. While the previous results mostly require exponentially-tailed losses, we also show that for general decreasing losses, the implicit bias can still be characterized in terms of the regularization path. In the nonseparable case, we design a nearly-optimal algorithm by combining logistic regression and perceptron. We also characterize the implicit bias of gradient descent via a unique decomposition of the training set. For neural networks, we first provide test error bounds for shallow ReLU networks, using the recent idea of neural tangent kernel, but only requiring a mild overparameterization. Then we show implicit bias results for deep linear networks and deep homogeneous networks, in the form of alignment properties.
Degree
thesis:*- Name thesis:degree_name
- Ph.D.
- Level thesis:degree_level
- Dissertation
- Discipline thesis:degree_discipline
- Computer Science
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2022
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Ji, Ziwei
- Contributors dc:contributor
-
- Telgarsky, Matus
- Chekuri, Chandra
- Raginsky, Maxim
- Hsu, Daniel
Subjects
dc:subject × 5Rights
dc:rights- Statement dc:rights
-
- Copyright 2022 Ziwei Ji
- Language dc:language
- en, eng
Identifiers
dc:identifier.*- Handle dc:identifier
- https://hdl.handle.net/2142/115475