{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/132818"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/132818","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"HarmonySched: Dynamically scheduling multiple concurrent machine learning models across a heterogeneous system on chip","abstract":"Increasing interest in Heterogeneous SoCs has led to the need to find more optimal ways to effectively use all the resources provided by such chips. Motivated by various AI applications, modern SoC systems integrate components such as GPUs and NPUs. Very few studies address scheduling multiple concurrent workloads and dynamically arriving workloads on these SoCs. This study introduces a dual-layer scheduling algorithm that directs workloads to either an iGPU or NPU on the SoC using a novel machine learning algorithm that aims to maximize latency as well as throughput. The accelerator chosen (embedded GPU or NPU) performs its own scheduling. The GPU uses temporal time slicing in a queue to ensure fair resource sharing amongst workloads and the NPU executes workloads sequentially in a queue. Workloads executing on both accelerators are ordered by first priority and then deadline. Unfinished GPU workloads are re-queued to allow for better resource sharing. Experiments show that this approach leads to a 2.7x reduction in tail latencies, 1.5x improvement in throughput, and 2.38x reduction in deadline violations compared to existing schedulers.","abstract_html":"Increasing interest in Heterogeneous SoCs has led to the need to find more optimal ways to effectively use all the resources provided by such chips. Motivated by various AI applications, modern SoC systems integrate components such as GPUs and NPUs. Very few studies address scheduling multiple concurrent workloads and dynamically arriving workloads on these SoCs. This study introduces a dual-layer scheduling algorithm that directs workloads to either an iGPU or NPU on the SoC using a novel machine learning algorithm that aims to maximize latency as well as throughput. The accelerator chosen (embedded GPU or NPU) performs its own scheduling. The GPU uses temporal time slicing in a queue to ensure fair resource sharing amongst workloads and the NPU executes workloads sequentially in a queue. Workloads executing on both accelerators are ordered by first priority and then deadline. Unfinished GPU workloads are re-queued to allow for better resource sharing. Experiments show that this approach leads to a 2.7x reduction in tail latencies, 1.5x improvement in throughput, and 2.38x reduction in deadline violations compared to existing schedulers.","abstract_has_math":false,"creators":["Pingali, Sanjana"],"institution":"University of Illinois Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Chen, Deming"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-12","date_published":"2025-12","updated_at":"2026-07-22T22:25:07Z","subjects":["Accelerator","Heterogeneous","Scheduling","Contention","Machine Learning"],"languages":["en"],"rights":["Copyright 2025 Sanjana Pingali"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/132818","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Chen, Deming"]},{"key":"dc:creator","label":"Author","values":["Pingali, Sanjana"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-12","2025-12-12"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Accelerator","Heterogeneous","Scheduling","Contention","Machine Learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Sanjana Pingali"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/132818"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Increasing interest in Heterogeneous SoCs has led to the need to find more optimal ways to effectively use all the resources provided by such chips. Motivated by various AI applications, modern SoC systems integrate components such as GPUs and NPUs. Very few studies address scheduling multiple concurrent workloads and dynamically arriving workloads on these SoCs. This study introduces a dual-layer scheduling algorithm that directs workloads to either an iGPU or NPU on the SoC using a novel machine learning algorithm that aims to maximize latency as well as throughput. The accelerator chosen (embedded GPU or NPU) performs its own scheduling. The GPU uses temporal time slicing in a queue to ensure fair resource sharing amongst workloads and the NPU executes workloads sequentially in a queue. Workloads executing on both accelerators are ordered by first priority and then deadline. Unfinished GPU workloads are re-queued to allow for better resource sharing. Experiments show that this approach leads to a 2.7x reduction in tail latencies, 1.5x improvement in throughput, and 2.38x reduction in deadline violations compared to existing schedulers.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2027-12-01","The student, Sanjana Pingali, accepted the attached license on 2025-12-12 at 11:44.","The student, Sanjana Pingali, submitted this Thesis for approval on 2025-12-12 at 12:06.","This Thesis was approved for publication on 2025-12-12 at 15:12.","DSpace SAF Submission Ingestion Package generated from Vireo submission #23149 on 2026-02-19 at 20:10:14"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["HarmonySched: Dynamically scheduling multiple concurrent machine learning models across a heterogeneous system on chip"]}]}],"canonical_facts":{"dc:contributor":["Chen, Deming"],"dc:creator":["Pingali, Sanjana"],"dc:date":["2025-12","2025-12-12"],"dc:description":["Increasing interest in Heterogeneous SoCs has led to the need to find more optimal ways to effectively use all the resources provided by such chips. Motivated by various AI applications, modern SoC systems integrate components such as GPUs and NPUs. Very few studies address scheduling multiple concurrent workloads and dynamically arriving workloads on these SoCs. This study introduces a dual-layer scheduling algorithm that directs workloads to either an iGPU or NPU on the SoC using a novel machine learning algorithm that aims to maximize latency as well as throughput. The accelerator chosen (embedded GPU or NPU) performs its own scheduling. The GPU uses temporal time slicing in a queue to ensure fair resource sharing amongst workloads and the NPU executes workloads sequentially in a queue. Workloads executing on both accelerators are ordered by first priority and then deadline. Unfinished GPU workloads are re-queued to allow for better resource sharing. Experiments show that this approach leads to a 2.7x reduction in tail latencies, 1.5x improvement in throughput, and 2.38x reduction in deadline violations compared to existing schedulers.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2027-12-01","The student, Sanjana Pingali, accepted the attached license on 2025-12-12 at 11:44.","The student, Sanjana Pingali, submitted this Thesis for approval on 2025-12-12 at 12:06.","This Thesis was approved for publication on 2025-12-12 at 15:12.","DSpace SAF Submission Ingestion Package generated from Vireo submission #23149 on 2026-02-19 at 20:10:14"],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/132818"],"dc:language":["en"],"dc:rights":["Copyright 2025 Sanjana Pingali"],"dc:subject":["Accelerator","Heterogeneous","Scheduling","Contention","Machine Learning"],"dc:title":["HarmonySched: Dynamically scheduling multiple concurrent machine learning models across a heterogeneous system on chip"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:07Z"}