{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/132593"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/132593","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"SLO-aware optimization and stateful orchestration for LLM systems","abstract":"The rapid evolution of Large Language Models (LLMs) has shifted the focus of AI infrastructure from simple text generation to complex, multi-turn agentic workflows. As these applications become increasingly sensitive to latency and dependencies, existing serving systems—which primarily optimize for aggregate throughput—fail to meet application-specific Service Level Objectives (SLOs). Furthermore, as workloads evolve into multi-agent systems (MAS), the lack of robust state management and error recovery in current runtimes creates a bottleneck for reliable orchestration. This thesis addresses these challenges by proposing a comprehensive optimization of the LLM runtime stack. First, we present \\name, an SLO-aware serving system designed to maximize service \"goodput\" (the rate of requests served within strict performance goals) under imprecise request information. \\name employs a novel iterative scheduling algorithm and Criticality-Aware Length Matching (CALM) to dynamically refine resource allocation as generation progresses. Evaluation across diverse realistic workloads, including chat, deep research, and agentic pipelines, demonstrates that \\name improves service goodput by 1.4×–6.3× and achieves 28.5%–83.2% resource savings compared to state-of-the-art designs. Building upon this optimized serving layer, the thesis concludes by exploring the future of Stateful Agent Orchestration. We propose the design of an ML Agent Compiler, a runtime environment akin to a JVM for agents. This proposed framework addresses the limitations of current stateless orchestration by introducing graph-based checkpointing, forking engines, and deduplication of partial executions. Together, these works chart a path toward a unified, efficient, and fault-tolerant infrastructure for the next generation of AI applications.","abstract_html":"The rapid evolution of Large Language Models (LLMs) has shifted the focus of AI infrastructure from simple text generation to complex, multi-turn agentic workflows. As these applications become increasingly sensitive to latency and dependencies, existing serving systems—which primarily optimize for aggregate throughput—fail to meet application-specific Service Level Objectives (SLOs). Furthermore, as workloads evolve into multi-agent systems (MAS), the lack of robust state management and error recovery in current runtimes creates a bottleneck for reliable orchestration. This thesis addresses these challenges by proposing a comprehensive optimization of the LLM runtime stack. First, we present \\name, an SLO-aware serving system designed to maximize service &quot;goodput&quot; (the rate of requests served within strict performance goals) under imprecise request information. \\name employs a novel iterative scheduling algorithm and Criticality-Aware Length Matching (CALM) to dynamically refine resource allocation as generation progresses. Evaluation across diverse realistic workloads, including chat, deep research, and agentic pipelines, demonstrates that \\name improves service goodput by 1.4×–6.3× and achieves 28.5%–83.2% resource savings compared to state-of-the-art designs. Building upon this optimized serving layer, the thesis concludes by exploring the future of Stateful Agent Orchestration. We propose the design of an ML Agent Compiler, a runtime environment akin to a JVM for agents. This proposed framework addresses the limitations of current stateless orchestration by introducing graph-based checkpointing, forking engines, and deduplication of partial executions. Together, these works chart a path toward a unified, efficient, and fault-tolerant infrastructure for the next generation of AI applications.","abstract_has_math":false,"creators":["Wu, Zhiyu"],"institution":"University of Illinois Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Lai, Fan"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-12","date_published":"2025-12","updated_at":"2026-07-22T22:25:07Z","subjects":["Large Language Models","Multi-Agent Systems","Agentic Workflow","SLO-Aware","scheduling","JVM"],"languages":["en"],"rights":["Copyright 2025 Zhiyu Wu"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/132593","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Lai, Fan"]},{"key":"dc:creator","label":"Author","values":["Wu, Zhiyu"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-12","2025-12-09"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Large Language Models","Multi-Agent Systems","Agentic Workflow","SLO-Aware","scheduling","JVM"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Zhiyu Wu"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/132593"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The rapid evolution of Large Language Models (LLMs) has shifted the focus of AI infrastructure from simple text generation to complex, multi-turn agentic workflows. As these applications become increasingly sensitive to latency and dependencies, existing serving systems—which primarily optimize for aggregate throughput—fail to meet application-specific Service Level Objectives (SLOs). Furthermore, as workloads evolve into multi-agent systems (MAS), the lack of robust state management and error recovery in current runtimes creates a bottleneck for reliable orchestration. This thesis addresses these challenges by proposing a comprehensive optimization of the LLM runtime stack. First, we present \\name, an SLO-aware serving system designed to maximize service \"goodput\" (the rate of requests served within strict performance goals) under imprecise request information. \\name employs a novel iterative scheduling algorithm and Criticality-Aware Length Matching (CALM) to dynamically refine resource allocation as generation progresses. Evaluation across diverse realistic workloads, including chat, deep research, and agentic pipelines, demonstrates that \\name improves service goodput by 1.4×–6.3× and achieves 28.5%–83.2% resource savings compared to state-of-the-art designs. Building upon this optimized serving layer, the thesis concludes by exploring the future of Stateful Agent Orchestration. We propose the design of an ML Agent Compiler, a runtime environment akin to a JVM for agents. This proposed framework addresses the limitations of current stateless orchestration by introducing graph-based checkpointing, forking engines, and deduplication of partial executions. Together, these works chart a path toward a unified, efficient, and fault-tolerant infrastructure for the next generation of AI applications.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2026-02-19 without embargo terms","The student, Zhiyu Wu, accepted the attached license on 2025-12-08 at 20:48.","The student, Zhiyu Wu, submitted this Thesis for approval on 2025-12-08 at 21:06.","This Thesis was approved for publication on 2025-12-09 at 08:38.","DSpace SAF Submission Ingestion Package generated from Vireo submission #23105 on 2026-02-19 at 18:30:08"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["SLO-aware optimization and stateful orchestration for LLM systems"]}]}],"canonical_facts":{"dc:contributor":["Lai, Fan"],"dc:creator":["Wu, Zhiyu"],"dc:date":["2025-12","2025-12-09"],"dc:description":["The rapid evolution of Large Language Models (LLMs) has shifted the focus of AI infrastructure from simple text generation to complex, multi-turn agentic workflows. As these applications become increasingly sensitive to latency and dependencies, existing serving systems—which primarily optimize for aggregate throughput—fail to meet application-specific Service Level Objectives (SLOs). Furthermore, as workloads evolve into multi-agent systems (MAS), the lack of robust state management and error recovery in current runtimes creates a bottleneck for reliable orchestration. This thesis addresses these challenges by proposing a comprehensive optimization of the LLM runtime stack. First, we present \\name, an SLO-aware serving system designed to maximize service \"goodput\" (the rate of requests served within strict performance goals) under imprecise request information. \\name employs a novel iterative scheduling algorithm and Criticality-Aware Length Matching (CALM) to dynamically refine resource allocation as generation progresses. Evaluation across diverse realistic workloads, including chat, deep research, and agentic pipelines, demonstrates that \\name improves service goodput by 1.4×–6.3× and achieves 28.5%–83.2% resource savings compared to state-of-the-art designs. Building upon this optimized serving layer, the thesis concludes by exploring the future of Stateful Agent Orchestration. We propose the design of an ML Agent Compiler, a runtime environment akin to a JVM for agents. This proposed framework addresses the limitations of current stateless orchestration by introducing graph-based checkpointing, forking engines, and deduplication of partial executions. Together, these works chart a path toward a unified, efficient, and fault-tolerant infrastructure for the next generation of AI applications.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2026-02-19 without embargo terms","The student, Zhiyu Wu, accepted the attached license on 2025-12-08 at 20:48.","The student, Zhiyu Wu, submitted this Thesis for approval on 2025-12-08 at 21:06.","This Thesis was approved for publication on 2025-12-09 at 08:38.","DSpace SAF Submission Ingestion Package generated from Vireo submission #23105 on 2026-02-19 at 18:30:08"],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/132593"],"dc:language":["en"],"dc:rights":["Copyright 2025 Zhiyu Wu"],"dc:subject":["Large Language Models","Multi-Agent Systems","Agentic Workflow","SLO-Aware","scheduling","JVM"],"dc:title":["SLO-aware optimization and stateful orchestration for LLM systems"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:07Z"}