Back to results

Virginia Tech

Toward Predictable and Efficient Deep Neural Network Inference on Graphics Processing Units

Abstract

dc:description.abstract

GPUs dominate DNN inference but remain difficult to control predictably under multi-tenant load. This thesis presents a practical, end-to-end approach for predictable, efficient single-GPU inference built around a closed loop of predict → allocate → power-tune. First, we introduce SGPRS, a spatio–temporal scheduler that combines spatial partitioning with stream-based temporal concurrency and explicit staging for coarse preemption and partition reuse without costly reconfiguration. Second, we develop GRAIL, a lightweight online predictor of latency and throughput under varying TPC (SM-group) and clock settings, and GRAIL-A, a zero-overhead allocator that turns those predictions into fast TPC and clock decisions. Third, we design SAGE, a power-aware runtime that coordinates DVFS and TPC control via a central partitioner and per-tenant local schedulers. Across diverse CNN and Transformer models and workload mixes, the system improves total throughput, tightens P95 and P99 latency, and reduces energy compared to temporal-only, spatial-only, and framework baselines, while preserving deadline behavior. The result is a resource-aware, empirically validated path to predictable multi-tenant inference on a single NVIDIA GPU.

Degree

thesis:*
Name thesis:degree_name
Doctor of Philosophy
Level thesis:degree_level
doctoral
Discipline thesis:degree_discipline
Computer Engineering
Department dc:contributor.department
Electrical and Computer Engineering
Grantor dc:publisher
Virginia Tech
Year dc:date.issued
2025

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Fakhim Babaei, Amir
Chair dc:contributor.committeechair
  • Chantem, Thidapat
Committee members dc:contributor.committeemember
  • Wang, Yue J.
  • Stavrou, Angelos
  • Tilevich, Eli
  • Dimarino, Christina Marie

Subjects

dc:subject × 4

Rights

dc:rights
Statement dc:rights
  • In Copyright
Language dc:language.iso
en

Identifiers

dc:identifier.*
Dc Identifier Other
vt_gsexam:45061
OAI identifier oai:identifier
oai:vtechworks.lib.vt.edu:10919/139710

Chain of custody

source
Harvested from
Virginia Tech
Base URL
vtechworks.lib.vt.edu/oai/request
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Fakhim Babaei, Amir. Toward Predictable and Efficient Deep Neural Network Inference on Graphics Processing Units. doctoral thesis, Virginia Tech, 2025. https://hdl.handle.net/10919/139710