Back to results

University of Illinois at Urbana-Champaign

Dynamic sparsity: enabling efficient, interpretable and generalizable sequence models

Abstract

dc:description

As humans, we consume natural signals and generate structured information to understand our world and propagate knowledge. Equipped with dynamically and sparsely activated neural networks, we conduct the learning process efficiently, make interpretable decisions, and generalize the learned knowledge quickly to unseen scenarios. Recent advances in large language models (LLMs) have shown impressive results of artificial neural networks in understanding the real world through sequential data. Inspired by the dynamic sparsity presented in biological neural networks, the aim of this thesis is to improve the fundamental architecture of the artificial sequence models, so that they can complete the tasks with improved efficiency, interpretability, and generalizability. We start our research by investigating how to develop efficient neural architectures that can extract structured information from natural text sequences under human supervision. Particularly, we focus on extracting a knowledge graph that provides a sparse representation of the relationships between the important entities in a sentence. We show that a neural network that generates an alternating sequence of nodes and edges with hybrid span decoding can achieve state-of-the-art performance with linear time and space complexities. Supervised learning of knowledge graph extraction requires expensive human efforts for tagging relations and entity types in text sequences. This problem becomes more severe when it comes to the science discovery domain because of the cost of finding capable human experts for data annotation. We propose a pre-training objective together with a neural sequence model to sparsely and dynamically extract sentence-level keyword representations with diverse latent types. We show that our model is able to learn interpretable tagging of the sentence elements at the latent level without using any human annotations, and its learned knowledge can be quickly adapted to new domains with few labels. We further explore how to introduce dynamic sparsity to general sequence models for efficient long-sequence modeling. We first propose a general mechanism that enables neural networks to activate submodules sparsely and dynamically for sequence elements with end-to-end differentiability. We then design a neural architecture employing this mechanism to sparsely activate an attention module based on the representations learned from a linear state-space model. The input is sparsely selected into a First-In-First-Out memory that is dynamically updated throughout the generation process. We show that our model can bring significantly better training and inference quality-efficiency trade-offs than state-of-the-art models on various sequence modeling tasks, while at the same time revealing the amount of attention needed for each data sample through the learned sparse activation patterns. Lastly, to establish a large-scale baseline for future research on dynamic sparsity, we propose and scale a simple hybrid model that layer-wise combines sliding window attention and Mamba, a selective state-space model, up to 3.8 billion parameters with 3.2 trillion training tokens. The selective gate in Mamba can be viewed as an input selection mechanism implemented with soft gating. The proposed model substantially outperforms the state-of-the-art transformer architectures on major downstream benchmarks while having a bunch of good properties for efficient long-context processing: it can be length extrapolated infinitely in zero-shot and have perfect memory recall on tasks with instruction-tuning, while maintaining the linear time complexity and the constant memory complexity for sequence generation.

Degree

thesis:*
Name thesis:degree_name
Ph.D.
Level thesis:degree_level
Dissertation
Discipline thesis:degree_discipline
Computer Science
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2024

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Ren, Liliang
Contributors dc:contributor
  • Zhai, ChengXiang
  • Peng, Hao
  • Zhao, Han
  • Chang, Kevin Chen-Chuan
  • Liu, Yang

Subjects

dc:subject × 5

Rights

dc:rights
Statement dc:rights
  • Copyright 2024 Liliang Ren
Language dc:language
en, eng

Identifiers

dc:identifier.*
Handle dc:identifier
https://hdl.handle.net/2142/125587

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Ren, Liliang. Dynamic sparsity: enabling efficient, interpretable and generalizable sequence models. Dissertation thesis, University of Illinois at Urbana-Champaign, 2024. https://hdl.handle.net/2142/125587