{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/132496"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/132496","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Towards efficient and reliable infrastructure for machine learning","abstract":"The rapid advancement of machine learning, particularly large language and generative models, has enabled transformative applications but at immense computational cost. Training and serving these models requires tens of thousands of specialized accelerators that consume megawatts of power and costs hundreds of millions of dollars. The efficiency of the underlying compute infrastructure is, therefore, critical for making machine learning more accessible and sustainable. This thesis addresses infrastructure efficiency challenges through three interconnected dimensions: core infrastructure optimization, workload adaptation, and reliability enhancement. We present five contributions that introduce novel methodologies and system designs for improving efficiency, resource utilization, and reliability. First, in core infrastructure optimization, we address resource fragmentation by enabling resource disaggregation, and resolving network contention that arises in disaggregated systems. INDIGO addresses memory disaggregation challenges through network-aware page migration. The system uses contextual multi-armed bandits trained on historical application data to make migration decisions that account for network transfer costs and memory access locality benefits. Netscope introduces a delay sensitivity-driven congestion mitigation framework that quantifies how applications are affected by network congestion using probabilistic regression models. The framework dynamically adjusts congestion control parameters based on estimated delay sensitivity, and selectively throttles applications with low sensitivity while protecting delay-sensitive ones. Second, in workload adaptation, we address the unique characteristics of modern ML workloads. QLM improves efficiency of distributed inference for large language models by multiplexing interactive and batch requests. QLM leverages statistical properties of continuous batching to estimate request waiting times in queues and groups requests with similar performance characteristics to enable efficient decision-making and orchestrates request pulling, eviction, load balancing, and model swapping operations. Complementary to QLM, Chiron introduces hierarchical autoscaling that employs multi-level backpressure mechanisms that distinguishes between interactive and batch requests. The framework dynamically adjusts batch sizes at the local level using reactive backpressure and makes global scaling decisions based on request waiting time estimation to enable meeting the service-level objectives while maximizing efficiency. Third, in reliability enhancement, we conduct a comprehensive characterization of GPU failures in modern AI accelerators through analysis of failure data from large-scale production clusters. Our methodology examines failure patterns across hardware components, failure types, propagation mechanisms, and system-wide impacts, thus providing insights into the interplay between hardware failures, resilience mechanisms, and application-level fault tolerance. Together, these contributions provide a holistic approach towards making AI infrastructure more efficient, adaptive, and resilient.","abstract_html":"The rapid advancement of machine learning, particularly large language and generative models, has enabled transformative applications but at immense computational cost. Training and serving these models requires tens of thousands of specialized accelerators that consume megawatts of power and costs hundreds of millions of dollars. The efficiency of the underlying compute infrastructure is, therefore, critical for making machine learning more accessible and sustainable. This thesis addresses infrastructure efficiency challenges through three interconnected dimensions: core infrastructure optimization, workload adaptation, and reliability enhancement. We present five contributions that introduce novel methodologies and system designs for improving efficiency, resource utilization, and reliability. First, in core infrastructure optimization, we address resource fragmentation by enabling resource disaggregation, and resolving network contention that arises in disaggregated systems. INDIGO addresses memory disaggregation challenges through network-aware page migration. The system uses contextual multi-armed bandits trained on historical application data to make migration decisions that account for network transfer costs and memory access locality benefits. Netscope introduces a delay sensitivity-driven congestion mitigation framework that quantifies how applications are affected by network congestion using probabilistic regression models. The framework dynamically adjusts congestion control parameters based on estimated delay sensitivity, and selectively throttles applications with low sensitivity while protecting delay-sensitive ones. Second, in workload adaptation, we address the unique characteristics of modern ML workloads. QLM improves efficiency of distributed inference for large language models by multiplexing interactive and batch requests. QLM leverages statistical properties of continuous batching to estimate request waiting times in queues and groups requests with similar performance characteristics to enable efficient decision-making and orchestrates request pulling, eviction, load balancing, and model swapping operations. Complementary to QLM, Chiron introduces hierarchical autoscaling that employs multi-level backpressure mechanisms that distinguishes between interactive and batch requests. The framework dynamically adjusts batch sizes at the local level using reactive backpressure and makes global scaling decisions based on request waiting time estimation to enable meeting the service-level objectives while maximizing efficiency. Third, in reliability enhancement, we conduct a comprehensive characterization of GPU failures in modern AI accelerators through analysis of failure data from large-scale production clusters. Our methodology examines failure patterns across hardware components, failure types, propagation mechanisms, and system-wide impacts, thus providing insights into the interplay between hardware failures, resilience mechanisms, and application-level fault tolerance. Together, these contributions provide a holistic approach towards making AI infrastructure more efficient, adaptive, and resilient.","abstract_has_math":false,"creators":["Patke, Archit"],"institution":"University of Illinois Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Iyer, Ravishankar K","Kim, Nam Sung","Huang, Jian","Srivatsa, Mudhakar"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-12","date_published":"2025-12","updated_at":"2026-07-22T22:25:07Z","subjects":["Machine learning infrastructure","Resource disaggregation","Network congestion control","Large language model serving","GPU reliability","Autoscaling","Page migration","High-performance computing"],"languages":["en"],"rights":["Copyright 2025 Archit Patke"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/132496","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Iyer, Ravishankar K","Kim, Nam Sung","Huang, Jian","Srivatsa, Mudhakar"]},{"key":"dc:creator","label":"Author","values":["Patke, Archit"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-12","2025-12-05"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Machine learning infrastructure","Resource disaggregation","Network congestion control","Large language model serving","GPU reliability","Autoscaling","Page migration","High-performance computing"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Archit Patke"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/132496"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The rapid advancement of machine learning, particularly large language and generative models, has enabled transformative applications but at immense computational cost. Training and serving these models requires tens of thousands of specialized accelerators that consume megawatts of power and costs hundreds of millions of dollars. The efficiency of the underlying compute infrastructure is, therefore, critical for making machine learning more accessible and sustainable. This thesis addresses infrastructure efficiency challenges through three interconnected dimensions: core infrastructure optimization, workload adaptation, and reliability enhancement. We present five contributions that introduce novel methodologies and system designs for improving efficiency, resource utilization, and reliability. First, in core infrastructure optimization, we address resource fragmentation by enabling resource disaggregation, and resolving network contention that arises in disaggregated systems. INDIGO addresses memory disaggregation challenges through network-aware page migration. The system uses contextual multi-armed bandits trained on historical application data to make migration decisions that account for network transfer costs and memory access locality benefits. Netscope introduces a delay sensitivity-driven congestion mitigation framework that quantifies how applications are affected by network congestion using probabilistic regression models. The framework dynamically adjusts congestion control parameters based on estimated delay sensitivity, and selectively throttles applications with low sensitivity while protecting delay-sensitive ones. Second, in workload adaptation, we address the unique characteristics of modern ML workloads. QLM improves efficiency of distributed inference for large language models by multiplexing interactive and batch requests. QLM leverages statistical properties of continuous batching to estimate request waiting times in queues and groups requests with similar performance characteristics to enable efficient decision-making and orchestrates request pulling, eviction, load balancing, and model swapping operations. Complementary to QLM, Chiron introduces hierarchical autoscaling that employs multi-level backpressure mechanisms that distinguishes between interactive and batch requests. The framework dynamically adjusts batch sizes at the local level using reactive backpressure and makes global scaling decisions based on request waiting time estimation to enable meeting the service-level objectives while maximizing efficiency. Third, in reliability enhancement, we conduct a comprehensive characterization of GPU failures in modern AI accelerators through analysis of failure data from large-scale production clusters. Our methodology examines failure patterns across hardware components, failure types, propagation mechanisms, and system-wide impacts, thus providing insights into the interplay between hardware failures, resilience mechanisms, and application-level fault tolerance. Together, these contributions provide a holistic approach towards making AI infrastructure more efficient, adaptive, and resilient.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2026-02-19 without embargo terms","The student, Archit Patke, accepted the attached license on 2025-12-05 at 09:05.","The student, Archit Patke, submitted this Dissertation for approval on 2025-12-05 at 09:12.","This Dissertation was approved for publication on 2025-12-05 at 10:44.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22875 on 2026-02-19 at 18:24:49"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Towards efficient and reliable infrastructure for machine learning"]}]}],"canonical_facts":{"dc:contributor":["Iyer, Ravishankar K","Kim, Nam Sung","Huang, Jian","Srivatsa, Mudhakar"],"dc:creator":["Patke, Archit"],"dc:date":["2025-12","2025-12-05"],"dc:description":["The rapid advancement of machine learning, particularly large language and generative models, has enabled transformative applications but at immense computational cost. Training and serving these models requires tens of thousands of specialized accelerators that consume megawatts of power and costs hundreds of millions of dollars. The efficiency of the underlying compute infrastructure is, therefore, critical for making machine learning more accessible and sustainable. This thesis addresses infrastructure efficiency challenges through three interconnected dimensions: core infrastructure optimization, workload adaptation, and reliability enhancement. We present five contributions that introduce novel methodologies and system designs for improving efficiency, resource utilization, and reliability. First, in core infrastructure optimization, we address resource fragmentation by enabling resource disaggregation, and resolving network contention that arises in disaggregated systems. INDIGO addresses memory disaggregation challenges through network-aware page migration. The system uses contextual multi-armed bandits trained on historical application data to make migration decisions that account for network transfer costs and memory access locality benefits. Netscope introduces a delay sensitivity-driven congestion mitigation framework that quantifies how applications are affected by network congestion using probabilistic regression models. The framework dynamically adjusts congestion control parameters based on estimated delay sensitivity, and selectively throttles applications with low sensitivity while protecting delay-sensitive ones. Second, in workload adaptation, we address the unique characteristics of modern ML workloads. QLM improves efficiency of distributed inference for large language models by multiplexing interactive and batch requests. QLM leverages statistical properties of continuous batching to estimate request waiting times in queues and groups requests with similar performance characteristics to enable efficient decision-making and orchestrates request pulling, eviction, load balancing, and model swapping operations. Complementary to QLM, Chiron introduces hierarchical autoscaling that employs multi-level backpressure mechanisms that distinguishes between interactive and batch requests. The framework dynamically adjusts batch sizes at the local level using reactive backpressure and makes global scaling decisions based on request waiting time estimation to enable meeting the service-level objectives while maximizing efficiency. Third, in reliability enhancement, we conduct a comprehensive characterization of GPU failures in modern AI accelerators through analysis of failure data from large-scale production clusters. Our methodology examines failure patterns across hardware components, failure types, propagation mechanisms, and system-wide impacts, thus providing insights into the interplay between hardware failures, resilience mechanisms, and application-level fault tolerance. Together, these contributions provide a holistic approach towards making AI infrastructure more efficient, adaptive, and resilient.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2026-02-19 without embargo terms","The student, Archit Patke, accepted the attached license on 2025-12-05 at 09:05.","The student, Archit Patke, submitted this Dissertation for approval on 2025-12-05 at 09:12.","This Dissertation was approved for publication on 2025-12-05 at 10:44.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22875 on 2026-02-19 at 18:24:49"],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/132496"],"dc:language":["en"],"dc:rights":["Copyright 2025 Archit Patke"],"dc:subject":["Machine learning infrastructure","Resource disaggregation","Network congestion control","Large language model serving","GPU reliability","Autoscaling","Page migration","High-performance computing"],"dc:title":["Towards efficient and reliable infrastructure for machine learning"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:07Z"}