University of Illinois at Urbana-Champaign
Analysis of communication and computation overlap in accelerated programs
Abstract
dc:descriptionAccelerated computing has revolutionized a broad range of industries by applying Graphics Processing Units (GPU(s)) to optimize workloads for efficient performance. Any application that is designed to run on such specialized devices has a typical flow that begins and ends with data movement between Central Processing Unit (CPU) and GPU. Previous research has shown that this data movement is indeed a major bottleneck with even most optimized applications hitting only 30% of the peak performance [1]. While most recent efforts have been to manually tune communication to improve performance, in this work, we present three different general optimization techniques to introduce overlap between communication and computation. The optimized model thus developed on a stencil application is evaluated against the baseline model using various performance metrics including time and throughput. Sets of experiments performed on two different GPUs reported that a certain combination of these optimizations result in a maximum speedup of 6.5 and a 160% increase in bandwidth utilization when compared to the baseline model in Compute Unified Device Architecture (CUDA).
Degree
thesis:*- Name thesis:degree_name
- M.S.
- Level thesis:degree_level
- Thesis
- Discipline thesis:degree_discipline
- Computer Science
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2022
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Yellapragada, Sushma
- Contributors dc:contributor
-
- Snir, Marc
Subjects
dc:subject × 2Rights
dc:rights- Statement dc:rights
-
- Copyright 2022 Sushma Yellapragada
- Language dc:language
- en, eng
Identifiers
dc:identifier.*- Handle dc:identifier
- https://hdl.handle.net/2142/115786