University of Illinois at Urbana-Champaign
Parallel merge for many-core architectures
Abstract
dc:descriptionThis thesis proposes a novel GPU implementation for merging two sorted arrays. We consider the problem of merging two arrays A and B into a single array C. Each element in the arrays has a key. An ordering relation denoted by is defined on the keys. Array A and array B have m and n elements, respectively, where m and n do not have to be equal. Both array A and array B are sorted based on the ordering relation. The task is to produce the output array C of size m + n. Array C consists of all the input elements from array A and array B, and is sorted by the ordering relation. We applied several GPU-specific optimizations to a parallel merge algorithm. The optimizations include coordinating the memory access pattern, making full use of the shared memory and reducing the thread divergence. Our implementation achieves up to 10x and 40x speedup on Titan-Z and GTX 980 GPU respectively compared to thrust merge implementation.
Degree
thesis:*- Name thesis:degree_name
- M.S.
- Level thesis:degree_level
- Thesis
- Discipline thesis:degree_discipline
- Electrical & Computer Engr
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2016
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Lv, Jie
- Contributors dc:contributor
-
- Hwu, Wen-Mei W.
Subjects
dc:subject × 2Rights
dc:rights- Statement dc:rights
-
- Copyright 2016 Jie Lv
- Language dc:language
- en
Identifiers
dc:identifier.*- Handle dc:identifier
- http://hdl.handle.net/2142/90824
- OAI identifier oai:identifier
- oai:www.ideals.illinois.edu:2142/90824