Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 20 of 121 for “"distributed memory"”.

  1. Resource Management for Distributed Memory Multicomputers

    … The remap algorithm provides an efficient, distributed implementation with each processor analyzing the tasks assigned to it. Finally, the template strategy provided a static allocation to dynamic tree-based flow graphs.

    uiuc Repository record for Resource Management for Distributed Memory Multicomputers (opens in a new tab)

  2. Automatic Data Partitioning on Distributed Memory Multicomputers

    Distributed-memory parallel computers are increasingly being used to provide high levels of performance for scientific applications. Unfortunately, such machines are not very easy to program. A number of research efforts seek to alleviate this problem by developing compilers that take over the task …

    uiuc Repository record for Automatic Data Partitioning on Distributed Memory Multicomputers (opens in a new tab)

  3. Reconfiguration and recovery in distributed memory multicomputers

    Restriction data tranferred 2014-07-01T11:16:00-05:00 Original Data Group with Access UIUC Users [automated] Release Date: none Reason: ETDs are only available to UIUC Users without author permission

    uiuc Repository record for Reconfiguration and recovery in distributed memory multicomputers (opens in a new tab)

  4. Layer potential evaluations on distributed memory machines

    … investigates evaluating layer potentials on distributed-memory machines. The distributed algorithm introduced in this thesis is based on GIGAQBX and shows GIGAQBX contains plenty of parallelism. We evaluate our algorithm on the Comet supercomputer at the San Diego Supercomputer Center and …

    uiuc Repository record for Layer potential evaluations on distributed memory machines (opens in a new tab)

  5. Expanding the synthesis of distributed memory implementations

    … by the Sketch language to supplement its distributed memory parallelism with shared memory parallelism that uses the popular fork-join model. The primary contribution of this thesis is the means by which the code is assured to be free of race conditions. Sketch uses constraint satisfaction …

    mit Repository record for Expanding the synthesis of distributed memory implementations (opens in a new tab)

  6. Fast Fourier Transforms on Distributed Memory Parallel Machines

    … in developing a general purpose subroutine on a distributed memory parallel machine is the data distribution. It is possible that users would like to use the subroutine with different data distributions. Thus there is a need to design algorithms on distributed memory parallel machines which can …

    odu Repository record for Fast Fourier Transforms on Distributed Memory Parallel Machines (opens in a new tab)

  7. Mapping numerical software onto distributed memory parallel systems

    … the use of parallel computers, in particular distributed memory systems, by proving strategies for parallelisation and developing the core component of tools to aid scalar software porting. The ported code must not only efficiently exploit available parallel processing speed and distributed

    greenwich Repository record for Mapping numerical software onto distributed memory parallel systems (opens in a new tab)

  8. Tunable shared-memory abstractions for distributed-memory systems

    Distributed memory multiprocessor architectures offer enormous computational power, by exploiting the concurrent execution of many loosely connected processors. Yet, such scalability is not without price. Interface delays and low interconnection bandwidth to the distributed memories make internode …

    uiuc Repository record for Tunable shared-memory abstractions for distributed-memory systems (opens in a new tab)

  9. A Reconfigurable, Distributed-Memory Accelerator for Sparse Applications

    … On conventional systems, their irregular memory accesses and low arithmetic intensity create challenging memory bandwidth bottlenecks. To overcome such bottlenecks, distributed-SRAM architectures use tiled arrays of high-bandwidth local storage to achieve very high aggregate memory

    mit Repository record for A Reconfigurable, Distributed-Memory Accelerator for Sparse Applications (opens in a new tab)

  10. Vector processor virtualization: distributed memory hierarchy and simultaneous multithreading

    … the path for performing data shuffle and memory-indexed accesses from the data path for executing other vector instructions that access the memory. This separation speeds up the most common memory access operations by avoiding extra delays and unnecessary stalls. In this multilane-based VP …

    njit Repository record for Vector processor virtualization: distributed memory hierarchy and simultaneous multithreading (opens in a new tab)

  11. Code generation of array constructs for distributed memory systems

    … This is particularly evident when programming distributed memory clusters containing multiple NUMA chips and GPUs on each node since it would require a complex combination of MPI, OpenMP, CUDA, OpenCL, etc to achieve high performance even for sequentially simplistic codes. Programs requiring …

    uiuc Repository record for Code generation of array constructs for distributed memory systems (opens in a new tab)

  12. Compiling reductions in data parallel programs for distributed memory multiprocessors

    … code generation become even more important for distributed-memory multiprocessors. As part of the dHPF parallel compiler project, we developed reduction recognition and parallel code generation strategies for distributed-memory multiprocessors. Using combination of dependence analysis and …

    rice Repository record for Compiling reductions in data parallel programs for distributed memory multiprocessors (opens in a new tab)

  13. Compiling for Distributed Memory Multiprocessors Based on Access Region Analysis

    … of fully automatic parallelizing techniques for distributed memory multiprocessors. In the research, an ordinary Fortran 77 program is assumed as input, and no information is required from the programmers. Only the abilities of the compiler are used to detect parallelism from the input program, …

    uiuc Repository record for Compiling for Distributed Memory Multiprocessors Based on Access Region Analysis (opens in a new tab)

  14. Run-time thread management for large-scale distributed-memory multiprocessors

    Thesis (Ph. D.)--Massachusetts Institute of Technology, Dept. of Electrical Engineering and Computer Science, 1993.

    mit Repository record for Run-time thread management for large-scale distributed-memory multiprocessors (opens in a new tab)

  15. Designing a Compiler for a Distributed Memory Parallel Computing System

    … integrating multiple processors, a network, and memory onto a single chip. The benefits to this design include a reduction in overhead incurred by synchronization, communication, and memory accesses. To properly determine its effectiveness, the SCMP architecture must be exercised under a wide …

    vt Repository record for Designing a Compiler for a Distributed Memory Parallel Computing System (opens in a new tab)

  16. Compiler techniques for optimizing communication and data distribution for distributed-memory multicomputers

    Distributed-memory multicomputers, such as the Intel iPSC/860, the Intel Paragon, the IBM SP-1 /SP-2, the NCUBE/2, and the Thinking Machines CM-5, offer significant advantages over shared-memory multiprocessors in terms of cost and scalability. However, lacking a global address space, they present …

    uiuc Repository record for Compiler techniques for optimizing communication and data distribution for distributed-memory multicomputers (opens in a new tab)

  17. Performance measurement and hardware support for message passing in distributed memory multicomputers

    In distributed memory multicomputers, synchronization and data sharing are achieved by explicit message passing. Hence, the speed and efficiency of communication are very important in the overall performance of such machines. The goal of this thesis is to reduce the communication overhead by …

    uiuc Repository record for Performance measurement and hardware support for message passing in distributed memory multicomputers (opens in a new tab)

  18. Scalable Parallel Delaunay Image-to-Mesh Conversion for Shared and Distributed Memory Architectures

    … mesh generation algorithms on both shared and distributed memory architectures. In this thesis we present a novel two-level parallel tetrahedral mesh generation framework capable of delivering and sustaining close to 6000 of concurrent work units (cores). We achieve this by leveraging …

    odu Repository record for Scalable Parallel Delaunay Image-to-Mesh Conversion for Shared and Distributed Memory Architectures (opens in a new tab)

Page 1 of 7