Back to results

Massachusetts Institute of Technology

The Lottery Ticket Hypothesis: On Sparse, Trainable Neural Networks

Abstract

dc:description.abstract

In this thesis, I show that, from an early point in training, typical neural networks for computer vision contain subnetworks capable of training in isolation to the same accuracy as the original unpruned network. These subnetworks—which I find retroactively by pruning after training and rewinding weights to their values from earlier in training—are the same size as those produced by state-of-the-art pruning techniques from after training. They rely on a combination of structure and initialization: if either is modified (by reinitializing the network or shuffling which weights are pruned in each layer), accuracy drops. In small-scale settings, I show that these subnetworks exist from initialization; in large-scale settings, I show that they exist early in training (< 5% of the way through). In general, I find these subnetworks when the outcome of optimizing them becomes robust to the sample of SGD noise used to train them; that is, when they train to the same convex region of the loss landscape regardless of data order. This occurs at initialization in small-scale settings and early in training in large-scale settings. The implication of these findings is that it may be possible to prune neural networks early in training, which would create an opportunity to substantially reduce the cost of training from that point forward. In service of this goal, I establish a framework for what success would look like in solving this problem and survey existing techniques for pruning neural networks at initialization and early in training. I find that magnitude pruning at initialization matches state-of-the-art performance for this task. In addition, the only information that existing techniques extract are the per-layer proportions in which to prune the network; in the case of magnitude pruning, this means that the only signals necessary to achieve state-of-the-art results are the per-layer widths used by variance-scaled initialization techniques.

Degree

thesis:*
Name thesis:degree_name
Doctoral
Department dc:contributor.department
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
Grantor dc:publisher
Massachusetts Institute of Technology
Year dc:date.issued
2023

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Frankle, Jonathan
Advisor dc:contributor.advisor
  • Carbin, Michael

Rights

dc:rights
Statement dc:rights
  • In Copyright - Educational Use Permitted
  • Copyright MIT

Identifiers

dc:identifier.*
Handle dc:identifier.uri
https://hdl.handle.net/1721.1/150166
OAI identifier oai:identifier
oai:dspace.mit.edu:1721.1/150166

Chain of custody

source
Harvested from
MIT
Base URL
dspace.mit.edu/oai/request
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
related terms
citation

Frankle, Jonathan. The Lottery Ticket Hypothesis: On Sparse, Trainable Neural Networks. Massachusetts Institute of Technology, 2023. https://hdl.handle.net/1721.1/150166