Back to search

University of Illinois at Urbana-Champaign

LPC: Lossless parameter compression for deploying large language model inference on edge devices

Abstract

dc:description

The deployment of large language models (LLMs) with billions of parameters, particularly on edge systems, is challenging because of the limited memory capacity of edge accelerators, necessitating offloading the storage of parameters to the host memory. This leads to the frequent transfer of parameters over low-bandwidth PCIe, hindering the system’s ability to meet latency and throughput requirements for inference. Many lossy model compression techniques have been proposed to facilitate inference in resource-constrained systems, but they compromise model performance and require time-consuming, model-specific tuning. To tackle this challenge, we present a Lossless Parameter Compression (LPC) technique that exploits the unique numerical characteristics of LLM parameters. Using simple bit-level manipulation of parameters with the popular and hardware complexity-efficient LZ4 algorithm, LPC achieves high compression ratios and high decompression throughput. To demonstrate the efficiency of LPC, we redesign the hardware and software stack of an open-source CNN accelerator to accelerate PyTorch-based LLMs in a host-device full-system environment and integrate it with LPC. By transferring compressed parameters over PCIe and decompressing them on-the-fly within the accelerator, LPC alleviates the PCIe bottlenecks. In evaluations with the OPT-13B and Llama2-13B models, LPC reduces parameter sizes by 26.0–26.2% for BF16 and 19.4–41.7% for INT8 formats, translating into equivalent PCIe bandwidth savings and inference latency speedups of 1.31–1.32× and 1.19–1.43×, respectively.

Degree

thesis:*
Name thesis:degree_name
M.S.
Level thesis:degree_level
Thesis
Discipline thesis:degree_discipline
Electrical & Computer Engr
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2024

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Wang, Nachuan
Contributors dc:contributor
  • Kim, Nam Sung

Subjects

dc:subject × 3

Rights

dc:rights
Statement dc:rights
  • Copyright 2024 Nachuan Wang
Language dc:language
en, eng

Identifiers

dc:identifier.*
Handle dc:identifier
https://hdl.handle.net/2142/127242

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Wang, Nachuan. LPC: Lossless parameter compression for deploying large language model inference on edge devices. Thesis thesis, University of Illinois at Urbana-Champaign, 2024. https://hdl.handle.net/2142/127242