University of Illinois at Urbana-Champaign
Completion of hinge loss has an implicit bias
Abstract
dc:descriptionA new loss function is proposed which learns the hinge loss function an infinite number of times pushing f(xi)yi \to \infty. It is proven that for a linear model on linearly separable data this modified hinge loss function converges in the direction of the \ell2 max-margin separator at a rate of $\bigO\left( \sqrt{d/t} \right)$ where $d$ is the dimension of the data. Then, an explicit formula for the underlying dynamical system of the gradient descent iterates for two-layer linear networks on the inner product loss function is derived. Using the derived dynamical system, a precise explicit algorithm is developed which when implemented reproduces the gradient descent iterates of two-layer ReLU nets on the inner product exactly. This result is studied further to extrapolate conclusions for neural network optimization.
Degree
thesis:*- Name thesis:degree_name
- M.S.
- Level thesis:degree_level
- Thesis
- Discipline thesis:degree_discipline
- Computer Science
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2020
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Lizama, Justin N
- Contributors dc:contributor
-
- Telgarsky, Matus J
Subjects
dc:subject × 4Rights
dc:rights- Statement dc:rights
-
- Copyright 2020 Justin Lizama
- Language dc:language
- en
Identifiers
dc:identifier.*- Handle dc:identifier
- http://hdl.handle.net/2142/108189
- OAI identifier oai:identifier
- oai:www.ideals.illinois.edu:2142/108189