{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/129257"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/129257","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Efficient visual generation for community player and production service","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2025-10-19 without embargo terms","abstract_has_math":false,"creators":["Zhang, Chengsong"],"institution":"University of Illinois Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Lai, Fan"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-04-28","date_published":"2025-04-28","updated_at":"2026-07-22T22:25:04Z","subjects":["Image generation","ML system"],"languages":["en","eng"],"rights":["Copyright 2025 Chengsong Zhang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/129257","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Lai, Fan"]},{"key":"dc:creator","label":"Author","values":["Zhang, Chengsong"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-04-28","2025-05"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Image generation","ML system"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Chengsong Zhang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/129257"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","The student, Chengsong Zhang, accepted the attached license on 2025-04-25 at 17:19.","The student, Chengsong Zhang, submitted this Thesis for approval on 2025-04-25 at 17:23.","This Thesis was approved for publication on 2025-04-28 at 12:38.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21994 on 2025-10-19 at 18:11:03","Morpheus improves the performance and throughput of serving diffusion transformer models (DiTs) for high-resolution image generation by strategically distributing computation across GPUs and adaptively skipping less salient regions. Unlike Large Language Models (LLMs), which are often I/O bound, image generation models like DiTs are predominantly compute-bound, especially at high resolutions. Furthermore, different regions within an image generation task exhibit varying levels of computational difficulty and semantic importance. To address these challenges, we introduce several technical contributions in Morpheus. First, we leverage a cross-attention-based analysis to dynamically identify and partition the image into salient and less salient regions. Second, we develop an algorithm that partitions the total computation into sub-tasks, considering both regional saliency and available compute resources. Third, we propose an adaptive skip algorithm that dynamically decides, on-the-fly during the diffusion process, when to skip computation for less salient regions without significant quality loss. Fourth, a centralized task scheduler dynamically assigns sub-tasks to workers, intelligently filling potential idle time arising from skipping with sub-tasks from other queued requests to boost overall throughput. Finally, we employ an asynchronous communication pattern for KV cache management and migration, crucial for enabling efficient parallel execution and hiding communication latency behind computation. Our evaluation of Morpheus serving the Flux model on 4x A100 40GB GPUs demonstrates a 2.1x speed-up compared to single-GPU execution, while maintaining minimal image quality degradation."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Efficient visual generation for community player and production service"]}]}],"canonical_facts":{"dc:contributor":["Lai, Fan"],"dc:creator":["Zhang, Chengsong"],"dc:date":["2025-04-28","2025-05"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","The student, Chengsong Zhang, accepted the attached license on 2025-04-25 at 17:19.","The student, Chengsong Zhang, submitted this Thesis for approval on 2025-04-25 at 17:23.","This Thesis was approved for publication on 2025-04-28 at 12:38.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21994 on 2025-10-19 at 18:11:03","Morpheus improves the performance and throughput of serving diffusion transformer models (DiTs) for high-resolution image generation by strategically distributing computation across GPUs and adaptively skipping less salient regions. Unlike Large Language Models (LLMs), which are often I/O bound, image generation models like DiTs are predominantly compute-bound, especially at high resolutions. Furthermore, different regions within an image generation task exhibit varying levels of computational difficulty and semantic importance. To address these challenges, we introduce several technical contributions in Morpheus. First, we leverage a cross-attention-based analysis to dynamically identify and partition the image into salient and less salient regions. Second, we develop an algorithm that partitions the total computation into sub-tasks, considering both regional saliency and available compute resources. Third, we propose an adaptive skip algorithm that dynamically decides, on-the-fly during the diffusion process, when to skip computation for less salient regions without significant quality loss. Fourth, a centralized task scheduler dynamically assigns sub-tasks to workers, intelligently filling potential idle time arising from skipping with sub-tasks from other queued requests to boost overall throughput. Finally, we employ an asynchronous communication pattern for KV cache management and migration, crucial for enabling efficient parallel execution and hiding communication latency behind computation. Our evaluation of Morpheus serving the Flux model on 4x A100 40GB GPUs demonstrates a 2.1x speed-up compared to single-GPU execution, while maintaining minimal image quality degradation."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/129257"],"dc:language":["en","eng"],"dc:rights":["Copyright 2025 Chengsong Zhang"],"dc:subject":["Image generation","ML system"],"dc:title":["Efficient visual generation for community player and production service"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:04Z"}