Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 20 of 40 for “"parallel performance"”.

  1. Techniques in scalable and effective parallel performance analysis

    Performance analysis tools are essential to the maintenance of efficient parallel execution of scientific applications. As scientific applications are executed on larger and larger parallel supercomputers, it is clear that performance tools must employ more advanced techniques to keep up with the …

    uiuc Repository record for Techniques in scalable and effective parallel performance analysis (opens in a new tab)

  2. A Study of Improving the Parallel Performance of VASP.

    … thesis involves a case study in the use of parallelism to improve the performance of an application for computational research on molecules. The application, VASP, was migrated from a machine with 4 nodes and 16 single-threaded processors to a machine with 60 nodes and 120 dual-threaded …

    etsu Repository record for A Study of Improving the Parallel Performance of VASP. (opens in a new tab)

  3. The Parallel Performance and Implementation of an Adaptive Multigrid Algorithm

    … algorithm has been implemented on shared memory parallel computers to solve large-scale structural mechanics problems. The solution algorithm begins by solving the problem on the initial mesh, refining this mesh as required by the chosen adaptive scheme, and then solving the problem on the new …

    uiuc Repository record for The Parallel Performance and Implementation of an Adaptive Multigrid Algorithm (opens in a new tab)

  4. Optimizing Parallel Performance with Work and Span in the OpenCilk Compiler

    … programming environment designed for high-performance multicore computing. OpenCilk consists of a LLVM fork called Tapir and a runtime scheduler, which, together, allow for OpenCilk’s high performance in practice. However, there are many opportunities to improve on the implementation of …

    mit Repository record for Optimizing Parallel Performance with Work and Span in the OpenCilk Compiler (opens in a new tab)

  5. An efficient parallelization of a real scientific application

    … shows that a substantial improvement in performance can be achieved by the parallelization of a real scientific application for a heterogeneous network of Sun and Silicon Graphics workstations connected by an Ethernet network, but that this is affected by a number of factors. These …

    cape-town Repository record for An efficient parallelization of a real scientific application (opens in a new tab)

  6. Prediction of ship rudder-propeller interaction using parallel computations and wind tunnel measurements

    … between a ship rudder and propeller. A parallel lifting suface panel program (PALISUPAN) has been written in Occam2 which is designed to run across variable sized square arrays of transputers. This program forms the basis of the theoretical method. The rudder and propeller are modelled …

    soton Repository record for Prediction of ship rudder-propeller interaction using parallel computations and wind tunnel measurements (opens in a new tab)

  7. Hypergraph-Based Combinatorial Optimization of Matrix -Vector Multiplication

    The second problem we address is parallel matrix-vector multiplication for large sparse matrices. Parallel sparse matrix-vector multiplication is a particularly important numerical kernel in computational science. We have focused on optimizing the parallel performance of this operation by reducing …

    uiuc Repository record for Hypergraph-Based Combinatorial Optimization of Matrix -Vector Multiplication (opens in a new tab)

  8. The exploitation of pipeline parallelism by compile time dataflow analysis

    … This paper proposes a method which maximizes the parallel performance of an instruction pipeline by detecting and eliminating specific pipeline hazards known as resource conflicts. The detection of resource conflicts is accomplished with data dependence analysis, while the elimination of resource …

    unlv Repository record for The exploitation of pipeline parallelism by compile time dataflow analysis (opens in a new tab)

  9. A strategy for mapping unstructured mesh computational mechanics programs onto distributed memory parallel architectures

    … advantages offered by distributed memory parallel processors. Strategies that successfully map structured mesh codes onto parallel machines have been developed over the previous decade and used to build a toolkit for automation of the parallelisation process. Extension of the capabilities …

    greenwich Repository record for A strategy for mapping unstructured mesh computational mechanics programs onto distributed memory parallel architectures (opens in a new tab)

  10. Performance analysis of cache oblivious Algorithms in the Fresh Breeze memory model

    … for easy, reliable and massively scalable parallel performance. The model achieves these goals by combining a radical memory model with efficient fine-grain parallelsim and managing both in hardware. This presents a unique opportunity for studying program execution in a system whose memory …

    mit Repository record for Performance analysis of cache oblivious Algorithms in the Fresh Breeze memory model (opens in a new tab)

  11. Parallel processing in overlapping tasks: A model and a method

    A dual-bottleneck model for parallel performance of reaction-time tasks in the psychological refractory period paradigm was developed. The first, preparation-based bottleneck prevents perceptual and decisional processes in the second task from operating in parallel with the first task when these …

    uiuc Repository record for Parallel processing in overlapping tasks: A model and a method (opens in a new tab)

  12. High Performance Algorithms for Structural Analysis of Grid Stiffened Panels

    In this research, we apply modern high performance computing techniques to solve an engineering problem, structural analysis of grid stiffened panels. An existing engineering code, SPANDO, is studied and modified to execute more efficiently on high performance workstations and parallel computers. …

    vt Repository record for High Performance Algorithms for Structural Analysis of Grid Stiffened Panels (opens in a new tab)

  13. A Scalable Parallel Multigrid Solver for Three Dimensional Adaptive Mesh Structural Analysis

    A parallel multigrid algorithm for solution of adaptive structural analysis meshes (ParMASA) is described. The user inputs a coarse mesh. The coarse mesh is successively solved, error estimated, and refined. The refinement simultaneously reduces discretization error and creates a hierarchy of …

    uiuc Repository record for A Scalable Parallel Multigrid Solver for Three Dimensional Adaptive Mesh Structural Analysis (opens in a new tab)

  14. Towards automatic migration of sequential kernels : Numba to PyKokkos

    High performance computing (HPC) frameworks have notably been useful for scientific communities, however, for an average user, leveraging their power can prove to be challenging. Using these highly performant frameworks requires knowledge about computer architectures, familiarity with parallel

    texas Repository record for Towards automatic migration of sequential kernels : Numba to PyKokkos (opens in a new tab)

  15. HARP: A MACHINE LEARNING FRAMEWORK ON TOP OF THE COLLECTIVE COMMUNICATION LAYER FOR THE BIG DATA SOFTWARE STACK

    … applications to exploit these new machines’ parallel computing capability. Instead, many efforts focus on specialized ways to speed up individual algorithms. In this thesis, the Harp framework, which uses collective communication techniques, is prototyped to improve the performance of data …

    iu Repository record for HARP: A MACHINE LEARNING FRAMEWORK ON TOP OF THE COLLECTIVE COMMUNICATION LAYER FOR THE BIG DATA SOFTWARE STACK (opens in a new tab)

  16. Multi-rate time integration on overset meshes

    … of overset mesh-described problems using a parallel Fortran code. The thesis focuses on the overarching mathematical theory, implementation via code generation, proof of numerical accuracy and stability, and demonstration of serial and parallel performance capabilities. Specifically, the …

    uiuc Repository record for Multi-rate time integration on overset meshes (opens in a new tab)

  17. Parallel Sparse Linear Algebra for Homotopy Methods

    … structural analysis problems. Developing parallel techniques for robust but expensive sequential computations, such as globally convergent homotopy methods, is important. The design of these techniques encompasses the functionality of the iterative method (adaptive GMRES(k)) implemented …

    vt Repository record for Parallel Sparse Linear Algebra for Homotopy Methods (opens in a new tab)

  18. Performance Portability of CUDA Across NVIDIA GPU Architectures

    … Processing Units (GPUs) provide impressive parallel performance that makes them invaluable to a number of computational workloads such as machine learning, simulations, and many others. NVIDIA GPUs currently outperform all of their competitors and thus make up the lion's share of today's …

    vt Repository record for Performance Portability of CUDA Across NVIDIA GPU Architectures (opens in a new tab)

  19. Modeling and Runtime Systems for Coordinated Power-Performance Management

    Emergent systems in high-performance computing (HPC) expect maximal efficiency to achieve the goal of power budget under 20-40 megawatts for 1 exaflop set by the Department of Energy. To optimize efficiency, emergent systems provide multiple power-performance control techniques to throttle …

    vt Repository record for Modeling and Runtime Systems for Coordinated Power-Performance Management (opens in a new tab)

Page 1 of 2