Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 20 of 40 for “"parallel performance"”.
-
Techniques in scalable and effective parallel performance analysis
Performance analysis tools are essential to the maintenance of efficient parallel execution of scientific applications. As scientific applications are executed on larger and larger parallel supercomputers, it is clear that performance tools must employ more advanced techniques to keep up with the …
-
A Study of Improving the Parallel Performance of VASP.
… thesis involves a case study in the use of parallelism to improve the performance of an application for computational research on molecules. The application, VASP, was migrated from a machine with 4 nodes and 16 single-threaded processors to a machine with 60 nodes and 120 dual-threaded …
-
The Parallel Performance and Implementation of an Adaptive Multigrid Algorithm
… algorithm has been implemented on shared memory parallel computers to solve large-scale structural mechanics problems. The solution algorithm begins by solving the problem on the initial mesh, refining this mesh as required by the chosen adaptive scheme, and then solving the problem on the new …
-
Optimizing Parallel Performance with Work and Span in the OpenCilk Compiler
… programming environment designed for high-performance multicore computing. OpenCilk consists of a LLVM fork called Tapir and a runtime scheduler, which, together, allow for OpenCilk’s high performance in practice. However, there are many opportunities to improve on the implementation of …
-
An efficient parallelization of a real scientific application
… shows that a substantial improvement in performance can be achieved by the parallelization of a real scientific application for a heterogeneous network of Sun and Silicon Graphics workstations connected by an Ethernet network, but that this is affected by a number of factors. These …
-
Prediction of ship rudder-propeller interaction using parallel computations and wind tunnel measurements
… between a ship rudder and propeller. A parallel lifting suface panel program (PALISUPAN) has been written in Occam2 which is designed to run across variable sized square arrays of transputers. This program forms the basis of the theoretical method. The rudder and propeller are modelled …
-
Hypergraph-Based Combinatorial Optimization of Matrix -Vector Multiplication
The second problem we address is parallel matrix-vector multiplication for large sparse matrices. Parallel sparse matrix-vector multiplication is a particularly important numerical kernel in computational science. We have focused on optimizing the parallel performance of this operation by reducing …
-
The exploitation of pipeline parallelism by compile time dataflow analysis
… This paper proposes a method which maximizes the parallel performance of an instruction pipeline by detecting and eliminating specific pipeline hazards known as resource conflicts. The detection of resource conflicts is accomplished with data dependence analysis, while the elimination of resource …
-
A strategy for mapping unstructured mesh computational mechanics programs onto distributed memory parallel architectures
… advantages offered by distributed memory parallel processors. Strategies that successfully map structured mesh codes onto parallel machines have been developed over the previous decade and used to build a toolkit for automation of the parallelisation process. Extension of the capabilities …
-
Performance analysis of cache oblivious Algorithms in the Fresh Breeze memory model
… for easy, reliable and massively scalable parallel performance. The model achieves these goals by combining a radical memory model with efficient fine-grain parallelsim and managing both in hardware. This presents a unique opportunity for studying program execution in a system whose memory …
-
Parallel processing in overlapping tasks: A model and a method
A dual-bottleneck model for parallel performance of reaction-time tasks in the psychological refractory period paradigm was developed. The first, preparation-based bottleneck prevents perceptual and decisional processes in the second task from operating in parallel with the first task when these …
-
High Performance Algorithms for Structural Analysis of Grid Stiffened Panels
In this research, we apply modern high performance computing techniques to solve an engineering problem, structural analysis of grid stiffened panels. An existing engineering code, SPANDO, is studied and modified to execute more efficiently on high performance workstations and parallel computers. …
-
A Scalable Parallel Multigrid Solver for Three Dimensional Adaptive Mesh Structural Analysis
A parallel multigrid algorithm for solution of adaptive structural analysis meshes (ParMASA) is described. The user inputs a coarse mesh. The coarse mesh is successively solved, error estimated, and refined. The refinement simultaneously reduces discretization error and creates a hierarchy of …
-
Towards automatic migration of sequential kernels : Numba to PyKokkos
High performance computing (HPC) frameworks have notably been useful for scientific communities, however, for an average user, leveraging their power can prove to be challenging. Using these highly performant frameworks requires knowledge about computer architectures, familiarity with parallel …
-
HARP: A MACHINE LEARNING FRAMEWORK ON TOP OF THE COLLECTIVE COMMUNICATION LAYER FOR THE BIG DATA SOFTWARE STACK
… applications to exploit these new machines’ parallel computing capability. Instead, many efforts focus on specialized ways to speed up individual algorithms. In this thesis, the Harp framework, which uses collective communication techniques, is prototyped to improve the performance of data …
-
Multi-rate time integration on overset meshes
… of overset mesh-described problems using a parallel Fortran code. The thesis focuses on the overarching mathematical theory, implementation via code generation, proof of numerical accuracy and stability, and demonstration of serial and parallel performance capabilities. Specifically, the …
-
Parallel Sparse Linear Algebra for Homotopy Methods
… structural analysis problems. Developing parallel techniques for robust but expensive sequential computations, such as globally convergent homotopy methods, is important. The design of these techniques encompasses the functionality of the iterative method (adaptive GMRES(k)) implemented …
-
Performance Portability of CUDA Across NVIDIA GPU Architectures
… Processing Units (GPUs) provide impressive parallel performance that makes them invaluable to a number of computational workloads such as machine learning, simulations, and many others. NVIDIA GPUs currently outperform all of their competitors and thus make up the lion's share of today's …
-
Modeling and Runtime Systems for Coordinated Power-Performance Management
Emergent systems in high-performance computing (HPC) expect maximal efficiency to achieve the goal of power budget under 20-40 megawatts for 1 exaflop set by the Department of Energy. To optimize efficiency, emergent systems provide multiple power-performance control techniques to throttle …
Page 1 of 2