{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/132596"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/132596","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"SKYAPI: structural-aware orchestration for LLM-based multi-agent systems","abstract":"The proliferation of Large Language Models (LLMs) has catalyzed the development of Multi-Agent Systems (MAS) capable of autonomous reasoning and complex workflow execution. These systems heavily rely on third-party inference APIs, where providers exhibit significant heterogeneity in latency and pricing. However, existing API routing strategies typically optimize requests in isolation, failing to account for the structural dependencies and synchronization barriers inherent in multi-agent collaboration. In this paper, we introduce SkyAPI, a structural-aware routing framework that shifts the optimization paradigm from per-query selection to stage-level orchestration. By decomposing dynamic agent workflows into ordered execution stages, SkyAPI employs a Mixed-Integer Linear Programming (MILP) formulation to explicitly optimize the trade-off between stage makespan and monetary cost. Furthermore, we propose a prefix-aware scheduling mechanism with Time-To-First-Token (TTFT) deferral to maximize KV-cache reuse among collaborative agents. Extensive evaluations on complex benchmarks, including DeepResearch and Gama-Bench, demonstrate that SkyAPI significantly outperforms state-of-the-art baselines, reducing operational costs by up to 3× while satisfying strict latency constraints across diverse model architectures.","abstract_html":"The proliferation of Large Language Models (LLMs) has catalyzed the development of Multi-Agent Systems (MAS) capable of autonomous reasoning and complex workflow execution. These systems heavily rely on third-party inference APIs, where providers exhibit significant heterogeneity in latency and pricing. However, existing API routing strategies typically optimize requests in isolation, failing to account for the structural dependencies and synchronization barriers inherent in multi-agent collaboration. In this paper, we introduce SkyAPI, a structural-aware routing framework that shifts the optimization paradigm from per-query selection to stage-level orchestration. By decomposing dynamic agent workflows into ordered execution stages, SkyAPI employs a Mixed-Integer Linear Programming (MILP) formulation to explicitly optimize the trade-off between stage makespan and monetary cost. Furthermore, we propose a prefix-aware scheduling mechanism with Time-To-First-Token (TTFT) deferral to maximize KV-cache reuse among collaborative agents. Extensive evaluations on complex benchmarks, including DeepResearch and Gama-Bench, demonstrate that SkyAPI significantly outperforms state-of-the-art baselines, reducing operational costs by up to 3× while satisfying strict latency constraints across diverse model architectures.","abstract_has_math":false,"creators":["Wang, Phillip"],"institution":"University of Illinois Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Lai, Fan"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-12","date_published":"2025-12","updated_at":"2026-07-22T22:25:07Z","subjects":["Large Language Model","Multi-agent System","Optimization"],"languages":["en"],"rights":["Copyright 2025 Phillip Wang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/132596","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Lai, Fan"]},{"key":"dc:creator","label":"Author","values":["Wang, Phillip"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-12","2025-12-09"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Large Language Model","Multi-agent System","Optimization"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Phillip Wang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/132596"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The proliferation of Large Language Models (LLMs) has catalyzed the development of Multi-Agent Systems (MAS) capable of autonomous reasoning and complex workflow execution. These systems heavily rely on third-party inference APIs, where providers exhibit significant heterogeneity in latency and pricing. However, existing API routing strategies typically optimize requests in isolation, failing to account for the structural dependencies and synchronization barriers inherent in multi-agent collaboration. In this paper, we introduce SkyAPI, a structural-aware routing framework that shifts the optimization paradigm from per-query selection to stage-level orchestration. By decomposing dynamic agent workflows into ordered execution stages, SkyAPI employs a Mixed-Integer Linear Programming (MILP) formulation to explicitly optimize the trade-off between stage makespan and monetary cost. Furthermore, we propose a prefix-aware scheduling mechanism with Time-To-First-Token (TTFT) deferral to maximize KV-cache reuse among collaborative agents. Extensive evaluations on complex benchmarks, including DeepResearch and Gama-Bench, demonstrate that SkyAPI significantly outperforms state-of-the-art baselines, reducing operational costs by up to 3× while satisfying strict latency constraints across diverse model architectures.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2026-02-19 without embargo terms","The student, Phillip Wang, accepted the attached license on 2025-12-09 at 12:06.","The student, Phillip Wang, submitted this Thesis for approval on 2025-12-09 at 12:17.","This Thesis was approved for publication on 2025-12-09 at 13:25.","DSpace SAF Submission Ingestion Package generated from Vireo submission #23112 on 2026-02-19 at 18:30:10"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["SKYAPI: structural-aware orchestration for LLM-based multi-agent systems"]}]}],"canonical_facts":{"dc:contributor":["Lai, Fan"],"dc:creator":["Wang, Phillip"],"dc:date":["2025-12","2025-12-09"],"dc:description":["The proliferation of Large Language Models (LLMs) has catalyzed the development of Multi-Agent Systems (MAS) capable of autonomous reasoning and complex workflow execution. These systems heavily rely on third-party inference APIs, where providers exhibit significant heterogeneity in latency and pricing. However, existing API routing strategies typically optimize requests in isolation, failing to account for the structural dependencies and synchronization barriers inherent in multi-agent collaboration. In this paper, we introduce SkyAPI, a structural-aware routing framework that shifts the optimization paradigm from per-query selection to stage-level orchestration. By decomposing dynamic agent workflows into ordered execution stages, SkyAPI employs a Mixed-Integer Linear Programming (MILP) formulation to explicitly optimize the trade-off between stage makespan and monetary cost. Furthermore, we propose a prefix-aware scheduling mechanism with Time-To-First-Token (TTFT) deferral to maximize KV-cache reuse among collaborative agents. Extensive evaluations on complex benchmarks, including DeepResearch and Gama-Bench, demonstrate that SkyAPI significantly outperforms state-of-the-art baselines, reducing operational costs by up to 3× while satisfying strict latency constraints across diverse model architectures.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2026-02-19 without embargo terms","The student, Phillip Wang, accepted the attached license on 2025-12-09 at 12:06.","The student, Phillip Wang, submitted this Thesis for approval on 2025-12-09 at 12:17.","This Thesis was approved for publication on 2025-12-09 at 13:25.","DSpace SAF Submission Ingestion Package generated from Vireo submission #23112 on 2026-02-19 at 18:30:10"],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/132596"],"dc:language":["en"],"dc:rights":["Copyright 2025 Phillip Wang"],"dc:subject":["Large Language Model","Multi-agent System","Optimization"],"dc:title":["SKYAPI: structural-aware orchestration for LLM-based multi-agent systems"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:07Z"}