University of Illinois Urbana-Champaign
Towards efficient and reliable infrastructure for machine learning
Abstract
dc:descriptionThe rapid advancement of machine learning, particularly large language and generative models, has enabled transformative applications but at immense computational cost. Training and serving these models requires tens of thousands of specialized accelerators that consume megawatts of power and costs hundreds of millions of dollars. The efficiency of the underlying compute infrastructure is, therefore, critical for making machine learning more accessible and sustainable. This thesis addresses infrastructure efficiency challenges through three interconnected dimensions: core infrastructure optimization, workload adaptation, and reliability enhancement. We present five contributions that introduce novel methodologies and system designs for improving efficiency, resource utilization, and reliability. First, in core infrastructure optimization, we address resource fragmentation by enabling resource disaggregation, and resolving network contention that arises in disaggregated systems. INDIGO addresses memory disaggregation challenges through network-aware page migration. The system uses contextual multi-armed bandits trained on historical application data to make migration decisions that account for network transfer costs and memory access locality benefits. Netscope introduces a delay sensitivity-driven congestion mitigation framework that quantifies how applications are affected by network congestion using probabilistic regression models. The framework dynamically adjusts congestion control parameters based on estimated delay sensitivity, and selectively throttles applications with low sensitivity while protecting delay-sensitive ones. Second, in workload adaptation, we address the unique characteristics of modern ML workloads. QLM improves efficiency of distributed inference for large language models by multiplexing interactive and batch requests. QLM leverages statistical properties of continuous batching to estimate request waiting times in queues and groups requests with similar performance characteristics to enable efficient decision-making and orchestrates request pulling, eviction, load balancing, and model swapping operations. Complementary to QLM, Chiron introduces hierarchical autoscaling that employs multi-level backpressure mechanisms that distinguishes between interactive and batch requests. The framework dynamically adjusts batch sizes at the local level using reactive backpressure and makes global scaling decisions based on request waiting time estimation to enable meeting the service-level objectives while maximizing efficiency. Third, in reliability enhancement, we conduct a comprehensive characterization of GPU failures in modern AI accelerators through analysis of failure data from large-scale production clusters. Our methodology examines failure patterns across hardware components, failure types, propagation mechanisms, and system-wide impacts, thus providing insights into the interplay between hardware failures, resilience mechanisms, and application-level fault tolerance. Together, these contributions provide a holistic approach towards making AI infrastructure more efficient, adaptive, and resilient.
Degree
thesis:*- Name thesis:degree_name
- Ph.D.
- Level thesis:degree_level
- Dissertation
- Discipline thesis:degree_discipline
- Electrical & Computer Engr
- Grantor
- University of Illinois Urbana-Champaign
- Year dc:date
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Patke, Archit
- Contributors dc:contributor
-
- Iyer, Ravishankar K
- Kim, Nam Sung
- Huang, Jian
- Srivatsa, Mudhakar
Subjects
dc:subject × 8Rights
dc:rights- Statement dc:rights
-
- Copyright 2025 Archit Patke
- Language dc:language
- en
Identifiers
dc:identifier.*- Handle dc:identifier
- https://hdl.handle.net/2142/132496
- OAI identifier oai:identifier
- oai:www.ideals.illinois.edu:2142/132496