Back to results

Massachusetts Institute of Technology

Reductions of ReLU neural networks to linear neural networks and their applications

Abstract

dc:description.abstract

Deep neural networks are the main subject of interest in the study of theoretical deep learning, which aims to rigorously explain the incredible performance of these function classes in practice. Although a lot are understood about deep linear network (neural network with all linear activations), nonlinearities in the activation make it very challenging to extend existing techniques to more realistic types of neural networks. In this thesis, we describe reductions of ReLU neural networks to linear neural networks, under various general condition of network architectures, loss functions and datasets. When such conditions are met, one can adapt techniques used in the theory of linear neural networks to study ReLU neural networks. To this end, we provide two applications in which we put the reduction to use: the first characterizes an implicit regularization behavior of ReLU neural networks trained with gradient descent and the second characterizes their convergence under gradient-based algorithms.

Degree

thesis:*
Name thesis:degree_name
Master
Department dc:contributor.department
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
Grantor dc:publisher
Massachusetts Institute of Technology
Year dc:date.issued
2022

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Le, Thien
Advisor dc:contributor.advisor
  • Jegelka, Stefanie

Rights

dc:rights
Statement dc:rights
  • In Copyright - Educational Use Permitted
  • Copyright MIT

Identifiers

dc:identifier.*
Handle dc:identifier.uri
https://hdl.handle.net/1721.1/143196
OAI identifier oai:identifier
oai:dspace.mit.edu:1721.1/143196

Chain of custody

source
Harvested from
MIT
Base URL
dspace.mit.edu/oai/request
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
related terms
citation

Le, Thien. Reductions of ReLU neural networks to linear neural networks and their applications. Massachusetts Institute of Technology, 2022. https://hdl.handle.net/1721.1/143196