Back to results

University of Illinois at Urbana-Champaign

The implicit bias of gradient descent: from linear classifiers to deep networks

Abstract

dc:description

Gradient descent and other first-order methods have been extensively used in deep learning. They can find solutions with not only low training error, but also low test error. This motivates the study of the implicit bias of training algorithms such as gradient descent, meaning we want to understand what special properties are satisfied by gradient descent solutions that lead to good generalization. In this thesis, we focus on gradient descent and its derivatives, show test error bounds and characterize the implicit biases. In detail, for linear classifiers, we consider both the separable setting and nonseparable setting. In the separable case, we first give an 1/t test error bound for stochastic gradient descent. Then for gradient descent, we present a primal-dual framework to analyze its implicit bias, which leads to fast margin maximization algorithms. While the previous results mostly require exponentially-tailed losses, we also show that for general decreasing losses, the implicit bias can still be characterized in terms of the regularization path. In the nonseparable case, we design a nearly-optimal algorithm by combining logistic regression and perceptron. We also characterize the implicit bias of gradient descent via a unique decomposition of the training set. For neural networks, we first provide test error bounds for shallow ReLU networks, using the recent idea of neural tangent kernel, but only requiring a mild overparameterization. Then we show implicit bias results for deep linear networks and deep homogeneous networks, in the form of alignment properties.

Degree

thesis:*
Name thesis:degree_name
Ph.D.
Level thesis:degree_level
Dissertation
Discipline thesis:degree_discipline
Computer Science
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2022

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Ji, Ziwei
Contributors dc:contributor
  • Telgarsky, Matus
  • Chekuri, Chandra
  • Raginsky, Maxim
  • Hsu, Daniel

Subjects

dc:subject × 5

Rights

dc:rights
Statement dc:rights
  • Copyright 2022 Ziwei Ji
Language dc:language
en, eng

Identifiers

dc:identifier.*
Handle dc:identifier
https://hdl.handle.net/2142/115475

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Ji, Ziwei. The implicit bias of gradient descent: from linear classifiers to deep networks. Dissertation thesis, University of Illinois at Urbana-Champaign, 2022. https://hdl.handle.net/2142/115475