Back to results

Virginia Tech

Optimizing Systems for Deep Learning Applications

Abstract

dc:description.abstract

Modern systems for Machine Learning (ML) workloads support heterogeneous workloads and resources. However, existing resource managers in these systems do not differentiate between heterogeneous GPU resources. Moreover, users are usually unaware of the sufficient and appropriate type and amount of GPU resources to request for their ML jobs. In this thesis, we analyze the performance of ML training and inference jobs and identify ML model and GPU characteristics that impact this performance. We then propose ML-based prediction models to accurately determine appropriate and sufficient resource requirements to ensure improved job latency and GPU utilization in the cluster.

Degree

thesis:*
Name thesis:degree_name
Doctor of Philosophy
Level thesis:degree_level
doctoral
Discipline thesis:degree_discipline
Computer Engineering
Department dc:contributor.department
Electrical and Computer Engineering
Grantor dc:publisher
Virginia Tech
Year dc:date.issued
2023

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Albahar, Hadeel Ahmad
Chair dc:contributor.committeechair
  • Butt, Ali R.
Committee members dc:contributor.committeemember
  • Anwar, Ali
  • Chantem, Thidapat
  • Min, Chang Woo
  • Tilevich, Eli

Subjects

dc:subject × 5

Rights

dc:rights
Statement dc:rights
  • In Copyright
Language dc:language.iso
en

Identifiers

dc:identifier.*
Dc Identifier Other
vt_gsexam:36650
OAI identifier oai:identifier
oai:vtechworks.lib.vt.edu:10919/114021

Chain of custody

source
Harvested from
Virginia Tech
Base URL
vtechworks.lib.vt.edu/oai/request
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Albahar, Hadeel Ahmad. Optimizing Systems for Deep Learning Applications. doctoral thesis, Virginia Tech, 2023. http://hdl.handle.net/10919/114021