{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/125587"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/125587","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Dynamic sparsity: enabling efficient, interpretable and generalizable sequence models","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-02-04 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2025-02-04 without embargo terms","abstract_has_math":false,"creators":["Ren, Liliang"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Zhai, ChengXiang","Peng, Hao","Zhao, Han","Chang, Kevin Chen-Chuan","Liu, Yang"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-07-12","date_published":"2024-07-12","updated_at":"2026-07-22T22:25:02Z","subjects":["Dynamic Sparsity","Neural Network","Deep Learning Architecture","Natural Language Processing","Machine Learning"],"languages":["en","eng"],"rights":["Copyright 2024 Liliang Ren"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/125587","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Zhai, ChengXiang","Peng, Hao","Zhao, Han","Chang, Kevin Chen-Chuan","Liu, Yang"]},{"key":"dc:creator","label":"Author","values":["Ren, Liliang"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2024-07-12","2024-08"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Dynamic Sparsity","Neural Network","Deep Learning Architecture","Natural Language Processing","Machine Learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2024 Liliang Ren"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/125587"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-02-04 without embargo terms","The student, Liliang Ren, accepted the attached license on 2024-07-09 at 01:58.","The student, Liliang Ren, submitted this Dissertation for approval on 2024-07-11 at 16:30.","This Dissertation was approved for publication on 2024-07-12 at 09:33.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20995 on 2025-02-04 at 21:04:44","As humans, we consume natural signals and generate structured information to understand our world and propagate knowledge. Equipped with dynamically and sparsely activated neural networks, we conduct the learning process efficiently, make interpretable decisions, and generalize the learned knowledge quickly to unseen scenarios. Recent advances in large language models (LLMs) have shown impressive results of artificial neural networks in understanding the real world through sequential data. Inspired by the dynamic sparsity presented in biological neural networks, the aim of this thesis is to improve the fundamental architecture of the artificial sequence models, so that they can complete the tasks with improved efficiency, interpretability, and generalizability. We start our research by investigating how to develop efficient neural architectures that can extract structured information from natural text sequences under human supervision. Particularly, we focus on extracting a knowledge graph that provides a sparse representation of the relationships between the important entities in a sentence. We show that a neural network that generates an alternating sequence of nodes and edges with hybrid span decoding can achieve state-of-the-art performance with linear time and space complexities. Supervised learning of knowledge graph extraction requires expensive human efforts for tagging relations and entity types in text sequences. This problem becomes more severe when it comes to the science discovery domain because of the cost of finding capable human experts for data annotation. We propose a pre-training objective together with a neural sequence model to sparsely and dynamically extract sentence-level keyword representations with diverse latent types. We show that our model is able to learn interpretable tagging of the sentence elements at the latent level without using any human annotations, and its learned knowledge can be quickly adapted to new domains with few labels. We further explore how to introduce dynamic sparsity to general sequence models for efficient long-sequence modeling. We first propose a general mechanism that enables neural networks to activate submodules sparsely and dynamically for sequence elements with end-to-end differentiability. We then design a neural architecture employing this mechanism to sparsely activate an attention module based on the representations learned from a linear state-space model. The input is sparsely selected into a First-In-First-Out memory that is dynamically updated throughout the generation process. We show that our model can bring significantly better training and inference quality-efficiency trade-offs than state-of-the-art models on various sequence modeling tasks, while at the same time revealing the amount of attention needed for each data sample through the learned sparse activation patterns. Lastly, to establish a large-scale baseline for future research on dynamic sparsity, we propose and scale a simple hybrid model that layer-wise combines sliding window attention and Mamba, a selective state-space model, up to 3.8 billion parameters with 3.2 trillion training tokens. The selective gate in Mamba can be viewed as an input selection mechanism implemented with soft gating. The proposed model substantially outperforms the state-of-the-art transformer architectures on major downstream benchmarks while having a bunch of good properties for efficient long-context processing: it can be length extrapolated infinitely in zero-shot and have perfect memory recall on tasks with instruction-tuning, while maintaining the linear time complexity and the constant memory complexity for sequence generation."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Dynamic sparsity: enabling efficient, interpretable and generalizable sequence models"]}]}],"canonical_facts":{"dc:contributor":["Zhai, ChengXiang","Peng, Hao","Zhao, Han","Chang, Kevin Chen-Chuan","Liu, Yang"],"dc:creator":["Ren, Liliang"],"dc:date":["2024-07-12","2024-08"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-02-04 without embargo terms","The student, Liliang Ren, accepted the attached license on 2024-07-09 at 01:58.","The student, Liliang Ren, submitted this Dissertation for approval on 2024-07-11 at 16:30.","This Dissertation was approved for publication on 2024-07-12 at 09:33.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20995 on 2025-02-04 at 21:04:44","As humans, we consume natural signals and generate structured information to understand our world and propagate knowledge. Equipped with dynamically and sparsely activated neural networks, we conduct the learning process efficiently, make interpretable decisions, and generalize the learned knowledge quickly to unseen scenarios. Recent advances in large language models (LLMs) have shown impressive results of artificial neural networks in understanding the real world through sequential data. Inspired by the dynamic sparsity presented in biological neural networks, the aim of this thesis is to improve the fundamental architecture of the artificial sequence models, so that they can complete the tasks with improved efficiency, interpretability, and generalizability. We start our research by investigating how to develop efficient neural architectures that can extract structured information from natural text sequences under human supervision. Particularly, we focus on extracting a knowledge graph that provides a sparse representation of the relationships between the important entities in a sentence. We show that a neural network that generates an alternating sequence of nodes and edges with hybrid span decoding can achieve state-of-the-art performance with linear time and space complexities. Supervised learning of knowledge graph extraction requires expensive human efforts for tagging relations and entity types in text sequences. This problem becomes more severe when it comes to the science discovery domain because of the cost of finding capable human experts for data annotation. We propose a pre-training objective together with a neural sequence model to sparsely and dynamically extract sentence-level keyword representations with diverse latent types. We show that our model is able to learn interpretable tagging of the sentence elements at the latent level without using any human annotations, and its learned knowledge can be quickly adapted to new domains with few labels. We further explore how to introduce dynamic sparsity to general sequence models for efficient long-sequence modeling. We first propose a general mechanism that enables neural networks to activate submodules sparsely and dynamically for sequence elements with end-to-end differentiability. We then design a neural architecture employing this mechanism to sparsely activate an attention module based on the representations learned from a linear state-space model. The input is sparsely selected into a First-In-First-Out memory that is dynamically updated throughout the generation process. We show that our model can bring significantly better training and inference quality-efficiency trade-offs than state-of-the-art models on various sequence modeling tasks, while at the same time revealing the amount of attention needed for each data sample through the learned sparse activation patterns. Lastly, to establish a large-scale baseline for future research on dynamic sparsity, we propose and scale a simple hybrid model that layer-wise combines sliding window attention and Mamba, a selective state-space model, up to 3.8 billion parameters with 3.2 trillion training tokens. The selective gate in Mamba can be viewed as an input selection mechanism implemented with soft gating. The proposed model substantially outperforms the state-of-the-art transformer architectures on major downstream benchmarks while having a bunch of good properties for efficient long-context processing: it can be length extrapolated infinitely in zero-shot and have perfect memory recall on tasks with instruction-tuning, while maintaining the linear time complexity and the constant memory complexity for sequence generation."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/125587"],"dc:language":["en","eng"],"dc:rights":["Copyright 2024 Liliang Ren"],"dc:subject":["Dynamic Sparsity","Neural Network","Deep Learning Architecture","Natural Language Processing","Machine Learning"],"dc:title":["Dynamic sparsity: enabling efficient, interpretable and generalizable sequence models"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:02Z"}