University of Illinois Urbana-Champaign
HarmonySched: Dynamically scheduling multiple concurrent machine learning models across a heterogeneous system on chip
Abstract
dc:descriptionIncreasing interest in Heterogeneous SoCs has led to the need to find more optimal ways to effectively use all the resources provided by such chips. Motivated by various AI applications, modern SoC systems integrate components such as GPUs and NPUs. Very few studies address scheduling multiple concurrent workloads and dynamically arriving workloads on these SoCs. This study introduces a dual-layer scheduling algorithm that directs workloads to either an iGPU or NPU on the SoC using a novel machine learning algorithm that aims to maximize latency as well as throughput. The accelerator chosen (embedded GPU or NPU) performs its own scheduling. The GPU uses temporal time slicing in a queue to ensure fair resource sharing amongst workloads and the NPU executes workloads sequentially in a queue. Workloads executing on both accelerators are ordered by first priority and then deadline. Unfinished GPU workloads are re-queued to allow for better resource sharing. Experiments show that this approach leads to a 2.7x reduction in tail latencies, 1.5x improvement in throughput, and 2.38x reduction in deadline violations compared to existing schedulers.
Degree
thesis:*- Name thesis:degree_name
- M.S.
- Level thesis:degree_level
- Thesis
- Discipline thesis:degree_discipline
- Electrical & Computer Engr
- Grantor
- University of Illinois Urbana-Champaign
- Year dc:date
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Pingali, Sanjana
- Contributors dc:contributor
-
- Chen, Deming
Subjects
dc:subject × 5Rights
dc:rights- Statement dc:rights
-
- Copyright 2025 Sanjana Pingali
- Language dc:language
- en
Identifiers
dc:identifier.*- Handle dc:identifier
- https://hdl.handle.net/2142/132818
- OAI identifier oai:identifier
- oai:www.ideals.illinois.edu:2142/132818