Back to results

Rice University

Compiling reductions in data parallel programs for distributed memory multiprocessors

Abstract

dc:description.abstract

Reduction recognition and optimization are crucial techniques in parallelizing compilers. They are used to detect the recurrences in a program and transform the originally sequential code into parallel code. Because of the expensive interprocessor communication cost, reduction recognition and efficient code generation become even more important for distributed-memory multiprocessors. As part of the dHPF parallel compiler project, we developed reduction recognition and parallel code generation strategies for distributed-memory multiprocessors. Using combination of dependence analysis and pattern matching our reduction recognition technique can detect a broad range of reduction operations in complex loop or control flow structures. It recognizes reductions into scalar variable as well as reductions into one or more elements of all array. To generate efficient code, we generalize dHPF's computation partitioning algorithm to assign computations for reduction statements and apply various techniques to reduce the number of collective communications. Furthermore, we "factor" some reduction statements into a group of reduction statements to better exploit data locality. We have evaluated the effectiveness of our optimization oil an IBM Scalable PowerParallel System SP2. The results indicate that without reduction optimization, reduction operations are sequentially executed and become an execution bottleneck while with reduction support, we get good speedups and the time for reductions is almost negligible as compared to that for the non-optimized version. Compared to IBM's xlhpf compiler, dHPF can recognize a larger set of reduction operations and dHPF achieves 111% to 270% increase in efficiency for combined extreme value/location reductions. Our experiments show that factorization is an effective approach which reduces the reduction operation time by 50% or more where it is applicable.

Degree

thesis:*
Name thesis:degree_name
Master of Science
Level thesis:degree_level
Masters
Discipline thesis:degree_discipline
Engineering
Grantor
Rice University
Year dc:date.issued
1998

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Lu, Bo
Advisor dc:contributor.advisor
  • Mellor-Crummey, John

Subjects

dc:subject × 1

Rights

dc:rights
Statement dc:rights
  • Copyright is held by the author, unless otherwise indicated. Permission to reuse, publish, or reproduce the work beyond the bounds of fair use or other exemptions to copyright law must be obtained from the copyright holder.
Language dc:language.iso
eng

Identifiers

dc:identifier.*
Handle dc:identifier.uri
https://hdl.handle.net/1911/17397
OAI identifier oai:identifier
oai:repository.rice.edu:1911/17397

Chain of custody

source
Harvested from
Rice University
Base URL
repository.rice.edu/server/oai/request
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Lu, Bo. Compiling reductions in data parallel programs for distributed memory multiprocessors. Masters thesis, Rice University, 1998. https://hdl.handle.net/1911/17397