{"id":{"repo_id":"washington","oai_identifier":"oai:digital.lib.washington.edu:1773/57322"},"canonical_url":"https://search.dev.ndltd.org/etd/washington/oai:digital.lib.washington.edu:1773/57322","repository":{"repo_id":"washington","name":"University of Washington","base_url":"https://digital.lib.washington.edu/server/oai/request"},"display":{"title":"Linearization of Neural Networks as a Dual Framework for Stability and Sparsity","abstract":"Neural networks are powerful but difficult to reason about. Understanding why a model is stable, why it is trained effectively, or how it is being compressed without performance degradation requires tools that go beyond inspecting weights or monitoring loss curves. Linearization, which characterizes how a nonlinear system responds to small perturbations, offers a principled foundation for these questions. While a well known tool in fields such as control and dynamical systems, its potential as a unifying design and optimization language for neural networks has remained largely unexplored. This dissertation develops linearization as a dual framework for networks of neurons, addressing two complementary challenges: 1) Stability and hyperparameter sensitivity of Recurrent Neural Networks (RNNs), and 2) Performance recovery of compressed generative diffusion models. My thesis introduces the following contributions to the above goals. For RNNs, I use linearization as a design target. By shaping the eigenvalue spectrum of the recurrent Jacobian at initialization, I proposed to obtain networks whose dynamics remain stable across a broad range of regimes, making FORCE-based training robust to network configuration rather than confined to a narrow edge-of-chaos band. Turning from design to analysis, I then show that the Lyapunov Exponent spectrum, a signature of a network's long-run dynamical behavior derived from its linearization, encodes rich information about model performance. I introduce a representation learning approach that maps these spectra into a compact latent space, enabling reliable cross-architecture performance prediction from early in training and substantially reducing the cost of hyperparameter search. Building on this, I further leverage the Lyapunov Spectrum distance as a more reliable proxy than loss-based metrics for guiding automated search over pruned RNN variants, consistently identifying better-performing configurations at lower computational cost. For diffusion models, I demonstrate that Jacobian matching, aligning the input-output derivatives between dense teacher and compressed student, provides an effective finetuning signal that standard training objectives do not exploit. Since na\\\"{i}ve Jacobian matching is incompatible with diffusion training's inherent noise injection, I develop a second-order extension inspired by finite-time Lyapunov exponents that resolves this incompatibility and recovers generation quality after pruning. Together, these contributions establish linearization as a versatile theoretical language for both analysis and deployment of neural networks.","abstract_html":"Neural networks are powerful but difficult to reason about. Understanding why a model is stable, why it is trained effectively, or how it is being compressed without performance degradation requires tools that go beyond inspecting weights or monitoring loss curves. Linearization, which characterizes how a nonlinear system responds to small perturbations, offers a principled foundation for these questions. While a well known tool in fields such as control and dynamical systems, its potential as a unifying design and optimization language for neural networks has remained largely unexplored. This dissertation develops linearization as a dual framework for networks of neurons, addressing two complementary challenges: 1) Stability and hyperparameter sensitivity of Recurrent Neural Networks (RNNs), and 2) Performance recovery of compressed generative diffusion models. My thesis introduces the following contributions to the above goals. For RNNs, I use linearization as a design target. By shaping the eigenvalue spectrum of the recurrent Jacobian at initialization, I proposed to obtain networks whose dynamics remain stable across a broad range of regimes, making FORCE-based training robust to network configuration rather than confined to a narrow edge-of-chaos band. Turning from design to analysis, I then show that the Lyapunov Exponent spectrum, a signature of a network&#x27;s long-run dynamical behavior derived from its linearization, encodes rich information about model performance. I introduce a representation learning approach that maps these spectra into a compact latent space, enabling reliable cross-architecture performance prediction from early in training and substantially reducing the cost of hyperparameter search. Building on this, I further leverage the Lyapunov Spectrum distance as a more reliable proxy than loss-based metrics for guiding automated search over pruned RNN variants, consistently identifying better-performing configurations at lower computational cost. For diffusion models, I demonstrate that Jacobian matching, aligning the input-output derivatives between dense teacher and compressed student, provides an effective finetuning signal that standard training objectives do not exploit. Since na\\&quot;{i}ve Jacobian matching is incompatible with diffusion training&#x27;s inherent noise injection, I develop a second-order extension inspired by finite-time Lyapunov exponents that resolves this incompatibility and recovers generation quality after pruning. Together, these contributions establish linearization as a versatile theoretical language for both analysis and deployment of neural networks.","abstract_has_math":false,"creators":["Zheng, Caleb"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Shlizerman, Eli"],"committee_chairs":[],"committee_members":[],"year":2026,"date_issued":"2026-08-11","date_published":"2026-08-11","updated_at":"2026-08-21T16:50:44Z","subjects":["Deep Learning","Diffusion Model","Generative Model","Linearization","Lyapunox Exponents","Recurrect Neural Network","Artificial intelligence","Computer science","Engineering"],"languages":["en_US"],"rights":["none"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/1773/57322","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"source_record":{"url":"https://digital.lib.washington.edu/server/oai/request?verb=GetRecord&metadataPrefix=dim&identifier=oai%3Adigital.lib.washington.edu%3A1773%2F57322","prefix":"dim"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Shlizerman, Eli"]},{"key":"dc:creator","label":"Author","values":["Zheng, Caleb"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2026-08-11T19:28:43Z"]},{"key":"dc:date.issued","label":"Date","values":["2026-08-11"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Deep Learning","Diffusion Model","Generative Model","Linearization","Lyapunox Exponents","Recurrect Neural Network","Artificial intelligence","Computer science","Engineering"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en_US"]},{"key":"dc:rights","label":"Dc Rights","values":["none"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["Zheng_washington_0250E_29815.pdf"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/1773/57322"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Thesis (Ph.D.)--University of Washington, 2026"]},{"key":"dc:description.abstract","label":"Abstract","values":["Neural networks are powerful but difficult to reason about. Understanding why a model is stable, why it is trained effectively, or how it is being compressed without performance degradation requires tools that go beyond inspecting weights or monitoring loss curves. Linearization, which characterizes how a nonlinear system responds to small perturbations, offers a principled foundation for these questions. While a well known tool in fields such as control and dynamical systems, its potential as a unifying design and optimization language for neural networks has remained largely unexplored. This dissertation develops linearization as a dual framework for networks of neurons, addressing two complementary challenges: 1) Stability and hyperparameter sensitivity of Recurrent Neural Networks (RNNs), and 2) Performance recovery of compressed generative diffusion models. My thesis introduces the following contributions to the above goals. For RNNs, I use linearization as a design target. By shaping the eigenvalue spectrum of the recurrent Jacobian at initialization, I proposed to obtain networks whose dynamics remain stable across a broad range of regimes, making FORCE-based training robust to network configuration rather than confined to a narrow edge-of-chaos band. Turning from design to analysis, I then show that the Lyapunov Exponent spectrum, a signature of a network's long-run dynamical behavior derived from its linearization, encodes rich information about model performance. I introduce a representation learning approach that maps these spectra into a compact latent space, enabling reliable cross-architecture performance prediction from early in training and substantially reducing the cost of hyperparameter search. Building on this, I further leverage the Lyapunov Spectrum distance as a more reliable proxy than loss-based metrics for guiding automated search over pruned RNN variants, consistently identifying better-performing configurations at lower computational cost. For diffusion models, I demonstrate that Jacobian matching, aligning the input-output derivatives between dense teacher and compressed student, provides an effective finetuning signal that standard training objectives do not exploit. Since na\\\"{i}ve Jacobian matching is incompatible with diffusion training's inherent noise injection, I develop a second-order extension inspired by finite-time Lyapunov exponents that resolves this incompatibility and recovers generation quality after pruning. Together, these contributions establish linearization as a versatile theoretical language for both analysis and deployment of neural networks."]},{"key":"dc:format.mimetype","label":"Dc Format Mimetype","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Linearization of Neural Networks as a Dual Framework for Stability and Sparsity"]}]}],"canonical_facts":{"dc:contributor.advisor":["Shlizerman, Eli"],"dc:creator":["Zheng, Caleb"],"dc:date.accessioned":["2026-08-11T19:28:43Z"],"dc:date.issued":["2026-08-11"],"dc:description":["Thesis (Ph.D.)--University of Washington, 2026"],"dc:description.abstract":["Neural networks are powerful but difficult to reason about. Understanding why a model is stable, why it is trained effectively, or how it is being compressed without performance degradation requires tools that go beyond inspecting weights or monitoring loss curves. Linearization, which characterizes how a nonlinear system responds to small perturbations, offers a principled foundation for these questions. While a well known tool in fields such as control and dynamical systems, its potential as a unifying design and optimization language for neural networks has remained largely unexplored. This dissertation develops linearization as a dual framework for networks of neurons, addressing two complementary challenges: 1) Stability and hyperparameter sensitivity of Recurrent Neural Networks (RNNs), and 2) Performance recovery of compressed generative diffusion models. My thesis introduces the following contributions to the above goals. For RNNs, I use linearization as a design target. By shaping the eigenvalue spectrum of the recurrent Jacobian at initialization, I proposed to obtain networks whose dynamics remain stable across a broad range of regimes, making FORCE-based training robust to network configuration rather than confined to a narrow edge-of-chaos band. Turning from design to analysis, I then show that the Lyapunov Exponent spectrum, a signature of a network's long-run dynamical behavior derived from its linearization, encodes rich information about model performance. I introduce a representation learning approach that maps these spectra into a compact latent space, enabling reliable cross-architecture performance prediction from early in training and substantially reducing the cost of hyperparameter search. Building on this, I further leverage the Lyapunov Spectrum distance as a more reliable proxy than loss-based metrics for guiding automated search over pruned RNN variants, consistently identifying better-performing configurations at lower computational cost. For diffusion models, I demonstrate that Jacobian matching, aligning the input-output derivatives between dense teacher and compressed student, provides an effective finetuning signal that standard training objectives do not exploit. Since na\\\"{i}ve Jacobian matching is incompatible with diffusion training's inherent noise injection, I develop a second-order extension inspired by finite-time Lyapunov exponents that resolves this incompatibility and recovers generation quality after pruning. Together, these contributions establish linearization as a versatile theoretical language for both analysis and deployment of neural networks."],"dc:format.mimetype":["application/pdf"],"dc:identifier.other":["Zheng_washington_0250E_29815.pdf"],"dc:identifier.uri":["https://hdl.handle.net/1773/57322"],"dc:language.iso":["en_US"],"dc:rights":["none"],"dc:subject":["Deep Learning","Diffusion Model","Generative Model","Linearization","Lyapunox Exponents","Recurrect Neural Network","Artificial intelligence","Computer science","Engineering"],"dc:title":["Linearization of Neural Networks as a Dual Framework for Stability and Sparsity"],"dc:type":["Thesis"]},"updated_at":"2026-08-21T16:50:44Z"}