University of Illinois at Urbana-Champaign
Processor parallelism considerations and memory latency reduction in shared memory multiprocessors
Abstract
dc:descriptionA wide variety of computer architectures have been proposed to exploit parallelism at different granularities. These architectures have significant differences in instruction scheduling constraints, memory latencies, and synchronization overhead, making it difficult to determine which architecture can achieve the best performance on a given program. Trace-driven simulations and analytic models are used to compare the instruction-level parallelism of a superscalar processor and a pipelined processor with the loop-level parallelism of a shared memory multiprocessor. It is shown that the maximum speedup for a loop with a cyclic dependence graph is limited by its critical dependence ratio, independent of the number of iterations in the loop. The fine-grained processors are better suited for executing these loops with cyclic dependence graphs, while the multiprocessor has better performance on the very parallel loops with acyclic dependence graphs. When executing programs with a variety of loops and sequential code, the best performance is obtained using a multiprocessor architecture in which each individual processor has a fine-grained parallelism of two to four.
Degree
thesis:*- Name thesis:degree_name
- Ph.D.
- Level thesis:degree_level
- Dissertation
- Discipline thesis:degree_discipline
- Electrical and Computer Engineering
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2011
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Lilja, David John
- Contributors dc:contributor
-
- Yew, Pen-Chung
Subjects
dc:subject × 2Rights
dc:rights- Statement dc:rights
-
- Copyright 1991 Lilja, David John
- Language dc:language
- eng
Identifiers
dc:identifier.*- Identifier
-
AAI9210894
(UMI)AAI9210894 - OAI identifier oai:identifier
- oai:www.ideals.illinois.edu:2142/22386