{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/115786"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/115786","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Analysis of communication and computation overlap in accelerated programs","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-11 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2022-11-11 without embargo terms","abstract_has_math":false,"creators":["Yellapragada, Sushma"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Snir, Marc"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-05","date_published":"2022-05","updated_at":"2026-07-22T22:24:55Z","subjects":["Graphics Processing Units (GPU(s))","Compute Unified Device Architecture (CUDA)"],"languages":["en","eng"],"rights":["Copyright 2022 Sushma Yellapragada"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/115786","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Snir, Marc"]},{"key":"dc:creator","label":"Author","values":["Yellapragada, Sushma"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022-05","2022-04-27"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Graphics Processing Units (GPU(s))","Compute Unified Device Architecture (CUDA)"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2022 Sushma Yellapragada"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/115786"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-11 without embargo terms","The student, Sushma Yellapragada, accepted the attached license on 2022-04-27 at 15:00.","The student, Sushma Yellapragada, submitted this Thesis for approval on 2022-04-27 at 15:23.","This Thesis was approved for publication on 2022-04-27 at 15:43.","DSpace SAF Submission Ingestion Package generated from Vireo submission #17290 on 2022-11-11 at 17:53:53","Accelerated computing has revolutionized a broad range of industries by applying Graphics Processing Units (GPU(s)) to optimize workloads for efficient performance. Any application that is designed to run on such specialized devices has a typical flow that begins and ends with data movement between Central Processing Unit (CPU) and GPU. Previous research has shown that this data movement is indeed a major bottleneck with even most optimized applications hitting only 30% of the peak performance [1]. While most recent efforts have been to manually tune communication to improve performance, in this work, we present three different general optimization techniques to introduce overlap between communication and computation. The optimized model thus developed on a stencil application is evaluated against the baseline model using various performance metrics including time and throughput. Sets of experiments performed on two different GPUs reported that a certain combination of these optimizations result in a maximum speedup of 6.5 and a 160% increase in bandwidth utilization when compared to the baseline model in Compute Unified Device Architecture (CUDA)."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Analysis of communication and computation overlap in accelerated programs"]}]}],"canonical_facts":{"dc:contributor":["Snir, Marc"],"dc:creator":["Yellapragada, Sushma"],"dc:date":["2022-05","2022-04-27"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-11 without embargo terms","The student, Sushma Yellapragada, accepted the attached license on 2022-04-27 at 15:00.","The student, Sushma Yellapragada, submitted this Thesis for approval on 2022-04-27 at 15:23.","This Thesis was approved for publication on 2022-04-27 at 15:43.","DSpace SAF Submission Ingestion Package generated from Vireo submission #17290 on 2022-11-11 at 17:53:53","Accelerated computing has revolutionized a broad range of industries by applying Graphics Processing Units (GPU(s)) to optimize workloads for efficient performance. Any application that is designed to run on such specialized devices has a typical flow that begins and ends with data movement between Central Processing Unit (CPU) and GPU. Previous research has shown that this data movement is indeed a major bottleneck with even most optimized applications hitting only 30% of the peak performance [1]. While most recent efforts have been to manually tune communication to improve performance, in this work, we present three different general optimization techniques to introduce overlap between communication and computation. The optimized model thus developed on a stencil application is evaluated against the baseline model using various performance metrics including time and throughput. Sets of experiments performed on two different GPUs reported that a certain combination of these optimizations result in a maximum speedup of 6.5 and a 160% increase in bandwidth utilization when compared to the baseline model in Compute Unified Device Architecture (CUDA)."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/115786"],"dc:language":["en","eng"],"dc:rights":["Copyright 2022 Sushma Yellapragada"],"dc:subject":["Graphics Processing Units (GPU(s))","Compute Unified Device Architecture (CUDA)"],"dc:title":["Analysis of communication and computation overlap in accelerated programs"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:55Z"}