Back to results

Technische Universität Berlin

Workload-aware compressed linear algebra for data-centric machine learning pipelines

Abstract

dc:description.abstract

Compression is an effective technique for fitting data in available memory, reducing I/O across the storage-memory-cache hierarchy, decreasing energy consumption, and increasing instruction parallelism. Modern machine learning (ML) systems exploit the approximate nature of ML and mostly use lossy compression via low-precision floating- or fixed-point quantized representations. The lossy techniques have an unknown impact on convergence and model accuracy compared to full precision training and create trust concerns. Furthermore, exploratory refinement of such lossy decisions is difficult to define in declarative ML pipelines. To use declarative language abstractions, lossless matrix compression applies lightweight compression schemes to numeric matrices and enables compressed linear algebra operations such as matrix-vector multiplications directly on compressed representations. Traditional ML pipelines containing feature transformations and model training are increasingly extended into so-called data-centric ML pipelines with additional preprocessing steps for data cleaning, augmentation, and feature engineering to create higher-quality results. Given the trend towards increasingly complex composite pipelines, it is hard to infer the impact of compound ML pipeline primitives. Therefore, multiple variations of stacked techniques must be evaluated to find the best combinations. The evaluation of many pipeline variations is expensive but contains data redundancy that is exploitable via compression techniques. However, current compression techniques struggle to detect these redundancies. Individual pipeline stages, such as data cleaning, augmentation, and feature transformations, collect data characteristics, such as distinct items, column sparsity, and column correlations. These properties are core components in selecting compression schemes. Current compression algorithms redundantly rediscover these statistics in compression planning while compressing pipeline intermediates. Some systems already exploit redundancy using sparsity exploitation on intermediates, which is a form of lossless compression. This thesis aims to evolve sparsity exploitation to general redundancy-exploiting lossless compression that exploits common values instead of only zero values. Existing work on lossless compression and compressed linear algebra enable such exploitation to a degree but face challenges for general applicability. To solve the challenges, we introduce a workload-aware compression framework comprising a broad spectrum of new compression schemes to exploit different redundancy patterns and compressed kernels that can process long sequences of instructions with compressed intermediates and limited decompressions. The framework seamlessly fits into declarative ML pipelines by returning equivalent results to uncompressed linear algebra. We propose new feature transformation and engineering techniques that leverage information about the structural transformations collected in preprocessing pipelines. Furthermore, we develop a lightweight morphing technique adapting compressed intermediates to their subsequent linear algebra workloads. Instead of using a memory-centric approach that optimizes compression ratios, our workload-aware compression summarizes the workload of an ML pipeline and optimizes the compression scheme to minimize execution time. All presented contributions are integrated into Apache SystemDS, an open-source ML system for the end-to-end data science lifecycle. We evaluate our implementation on micro benchmarks of components, end-to-end ML pipelines, and distributed federated linear algebra. Our evaluation shows asymptotic improvements in operations performed on workload-aware compressed data. The asymptotic changes and data size reduction translate to real-time gains on individual operations on real datasets up to 10,000x compared to uncompressed and 20,000x compared to previous compressed linear algebra. Furthermore, end-to-end ML algorithms improve by 6.6x, and a data-centric pipeline reduces power consumption by 3.6x.

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Baunsgaard, Sebastian
Advisor dc:contributor.advisor
  • Boehm, Matthias

Rights

Language dc:language.iso
en

Identifiers

dc:identifier.*
OAI identifier oai:identifier
oai:depositonce.tu-berlin.de:11303/25907

Chain of custody

source
Harvested from
Technische Universität Berlin
Base URL
api-depositonce.tu-berlin.de/server/oai/request
Last updated
2026-07-27
Source record
OAI-PMH GetRecord
related terms
citation

Baunsgaard, Sebastian. Workload-aware compressed linear algebra for data-centric machine learning pipelines. 2025. https://depositonce.tu-berlin.de/handle/11303/25907