Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 20 of 121 for “"distributed memory"”.
-
Resource Management for Distributed Memory Multicomputers
… The remap algorithm provides an efficient, distributed implementation with each processor analyzing the tasks assigned to it. Finally, the template strategy provided a static allocation to dynamic tree-based flow graphs.
-
Automatic Data Partitioning on Distributed Memory Multicomputers
Distributed-memory parallel computers are increasingly being used to provide high levels of performance for scientific applications. Unfortunately, such machines are not very easy to program. A number of research efforts seek to alleviate this problem by developing compilers that take over the task …
-
Reconfiguration and recovery in distributed memory multicomputers
Restriction data tranferred 2014-07-01T11:16:00-05:00 Original Data Group with Access UIUC Users [automated] Release Date: none Reason: ETDs are only available to UIUC Users without author permission
-
Layer potential evaluations on distributed memory machines
… investigates evaluating layer potentials on distributed-memory machines. The distributed algorithm introduced in this thesis is based on GIGAQBX and shows GIGAQBX contains plenty of parallelism. We evaluate our algorithm on the Comet supercomputer at the San Diego Supercomputer Center and …
-
Expanding the synthesis of distributed memory implementations
… by the Sketch language to supplement its distributed memory parallelism with shared memory parallelism that uses the popular fork-join model. The primary contribution of this thesis is the means by which the code is assured to be free of race conditions. Sketch uses constraint satisfaction …
-
Fast Fourier Transforms on Distributed Memory Parallel Machines
… in developing a general purpose subroutine on a distributed memory parallel machine is the data distribution. It is possible that users would like to use the subroutine with different data distributions. Thus there is a need to design algorithms on distributed memory parallel machines which can …
-
Mapping numerical software onto distributed memory parallel systems
… the use of parallel computers, in particular distributed memory systems, by proving strategies for parallelisation and developing the core component of tools to aid scalar software porting. The ported code must not only efficiently exploit available parallel processing speed and distributed …
-
Tunable shared-memory abstractions for distributed-memory systems
Distributed memory multiprocessor architectures offer enormous computational power, by exploiting the concurrent execution of many loosely connected processors. Yet, such scalability is not without price. Interface delays and low interconnection bandwidth to the distributed memories make internode …
-
A Reconfigurable, Distributed-Memory Accelerator for Sparse Applications
… On conventional systems, their irregular memory accesses and low arithmetic intensity create challenging memory bandwidth bottlenecks. To overcome such bottlenecks, distributed-SRAM architectures use tiled arrays of high-bandwidth local storage to achieve very high aggregate memory …
-
Vector processor virtualization: distributed memory hierarchy and simultaneous multithreading
… the path for performing data shuffle and memory-indexed accesses from the data path for executing other vector instructions that access the memory. This separation speeds up the most common memory access operations by avoiding extra delays and unnecessary stalls. In this multilane-based VP …
-
Code generation of array constructs for distributed memory systems
… This is particularly evident when programming distributed memory clusters containing multiple NUMA chips and GPUs on each node since it would require a complex combination of MPI, OpenMP, CUDA, OpenCL, etc to achieve high performance even for sequentially simplistic codes. Programs requiring …
-
Compiling reductions in data parallel programs for distributed memory multiprocessors
… code generation become even more important for distributed-memory multiprocessors. As part of the dHPF parallel compiler project, we developed reduction recognition and parallel code generation strategies for distributed-memory multiprocessors. Using combination of dependence analysis and …
-
Compiling for Distributed Memory Multiprocessors Based on Access Region Analysis
… of fully automatic parallelizing techniques for distributed memory multiprocessors. In the research, an ordinary Fortran 77 program is assumed as input, and no information is required from the programmers. Only the abilities of the compiler are used to detect parallelism from the input program, …
-
Run-time thread management for large-scale distributed-memory multiprocessors
Thesis (Ph. D.)--Massachusetts Institute of Technology, Dept. of Electrical Engineering and Computer Science, 1993.
-
Designing a Compiler for a Distributed Memory Parallel Computing System
… integrating multiple processors, a network, and memory onto a single chip. The benefits to this design include a reduction in overhead incurred by synchronization, communication, and memory accesses. To properly determine its effectiveness, the SCMP architecture must be exercised under a wide …
-
Compiler techniques for optimizing communication and data distribution for distributed-memory multicomputers
Distributed-memory multicomputers, such as the Intel iPSC/860, the Intel Paragon, the IBM SP-1 /SP-2, the NCUBE/2, and the Thinking Machines CM-5, offer significant advantages over shared-memory multiprocessors in terms of cost and scalability. However, lacking a global address space, they present …
-
Performance measurement and hardware support for message passing in distributed memory multicomputers
In distributed memory multicomputers, synchronization and data sharing are achieved by explicit message passing. Hence, the speed and efficiency of communication are very important in the overall performance of such machines. The goal of this thesis is to reduce the communication overhead by …
-
Scalable Parallel Delaunay Image-to-Mesh Conversion for Shared and Distributed Memory Architectures
… mesh generation algorithms on both shared and distributed memory architectures. In this thesis we present a novel two-level parallel tetrahedral mesh generation framework capable of delivering and sustaining close to 6000 of concurrent work units (cores). We achieve this by leveraging …
Page 1 of 7