Back to search

University of Washington

Linearization of Neural Networks as a Dual Framework for Stability and Sparsity

Abstract

dc:description.abstract

Neural networks are powerful but difficult to reason about. Understanding why a model is stable, why it is trained effectively, or how it is being compressed without performance degradation requires tools that go beyond inspecting weights or monitoring loss curves. Linearization, which characterizes how a nonlinear system responds to small perturbations, offers a principled foundation for these questions. While a well known tool in fields such as control and dynamical systems, its potential as a unifying design and optimization language for neural networks has remained largely unexplored. This dissertation develops linearization as a dual framework for networks of neurons, addressing two complementary challenges: 1) Stability and hyperparameter sensitivity of Recurrent Neural Networks (RNNs), and 2) Performance recovery of compressed generative diffusion models. My thesis introduces the following contributions to the above goals. For RNNs, I use linearization as a design target. By shaping the eigenvalue spectrum of the recurrent Jacobian at initialization, I proposed to obtain networks whose dynamics remain stable across a broad range of regimes, making FORCE-based training robust to network configuration rather than confined to a narrow edge-of-chaos band. Turning from design to analysis, I then show that the Lyapunov Exponent spectrum, a signature of a network's long-run dynamical behavior derived from its linearization, encodes rich information about model performance. I introduce a representation learning approach that maps these spectra into a compact latent space, enabling reliable cross-architecture performance prediction from early in training and substantially reducing the cost of hyperparameter search. Building on this, I further leverage the Lyapunov Spectrum distance as a more reliable proxy than loss-based metrics for guiding automated search over pruned RNN variants, consistently identifying better-performing configurations at lower computational cost. For diffusion models, I demonstrate that Jacobian matching, aligning the input-output derivatives between dense teacher and compressed student, provides an effective finetuning signal that standard training objectives do not exploit. Since na\"{i}ve Jacobian matching is incompatible with diffusion training's inherent noise injection, I develop a second-order extension inspired by finite-time Lyapunov exponents that resolves this incompatibility and recovers generation quality after pruning. Together, these contributions establish linearization as a versatile theoretical language for both analysis and deployment of neural networks.

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Zheng, Caleb
Advisor dc:contributor.advisor
  • Shlizerman, Eli

Subjects

dc:subject × 9

Rights

dc:rights
Statement dc:rights
  • none
Language dc:language.iso
en_US

Identifiers

dc:identifier.*
Handle dc:identifier.uri
https://hdl.handle.net/1773/57322

Chain of custody

source
Harvested from
University of Washington
Base URL
digital.lib.washington.edu/server/oai/request
Last updated
2026-08-21
Source record
OAI-PMH GetRecord
citation

Zheng, Caleb. Linearization of Neural Networks as a Dual Framework for Stability and Sparsity. 2026. https://hdl.handle.net/1773/57322