Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 8 of 8 for “"Intel Xeon Phi"”.

  1. An Optimized Multiple Right-Hand Side Dslash Kernel for Intel Xeon Phi

    … point coprocessor cards such as GPUs and Intel Xeon Phi Knights Corner (KNC), which boast powerful vector processing units. Most of these efforts in the area of Dslash have focused on single right-hand side solvers. This thesis will present two optimized Dslash kernels which simplify …

    odu Repository record for An Optimized Multiple Right-Hand Side Dslash Kernel for Intel Xeon Phi (opens in a new tab)

  2. Exploring Performance Portability for Accelerators via High-level Parallel Patterns

    … e.g., multi-core CPUs, many-core GPUs (Graphics Processing Units) and Intel Xeon Phi. The performance gains from them can be as high as many orders of magnitude, attracting extensive interest from many scientific domains. However, the gains are closely followed by two main problems: (1) A …

    vt Repository record for Exploring Performance Portability for Accelerators via High-level Parallel Patterns (opens in a new tab)

  3. On-node performance optimization of a Monte-Carlo transport code for leadership architectures

    … profiling analysis was carried out on modern Intel micro-architectures (i.e., Intel Xeon Phi and Intel Xeon Platinum 8180) to understand what hardware and settings configurations were optimal. The specific modules and subroutines that were responsible for the performance drop were also …

    mit Repository record for On-node performance optimization of a Monte-Carlo transport code for leadership architectures (opens in a new tab)

  4. Performance analysis and optimization of a CFD application

    … execution on a dual-socket node fitted with an Intel Xeon Phi MIC card. To reduce the overhead created by host-accelerator copies in heterogeneous execution, the data layout of the halo region was changed from a ''star'' shape to a ''box'' shape to agglomerate small communications and to create …

    uiuc Repository record for Performance analysis and optimization of a CFD application (opens in a new tab)

  5. Modeling and Runtime Systems for Coordinated Power-Performance Management

    … impact on performance and energy consumption on Intel x86 systems with accelerators of Intel Xeon Phi and a Nvidia general-purpose graphics processing unit (GPGPU). We show the trade-offs and potentials for improving efficiency. Furthermore, we propose a parallel performance model for …

    vt Repository record for Modeling and Runtime Systems for Coordinated Power-Performance Management (opens in a new tab)

  6. Directive-Based Data Partitioning and Pipelining and Auto-Tuning for High-Performance GPU Computing

    … potential of accelerators, such as graphics processing units (GPUs), field-programmable gate arrays (FPGAs), and co-processors (e.g., Intel Xeon Phi), due to their increasing use in state-of-the-art supercomputers. Over the past 10 years, we have seen a significant improvement in both …

    vt Repository record for Directive-Based Data Partitioning and Pipelining and Auto-Tuning for High-Performance GPU Computing (opens in a new tab)

  7. Development of the random ray method of neutral particle transport for high-fidelity nuclear reactor simulation

    … computing systems, including CPU, GPU, and Intel Xeon Phi architectures. While 2D MOC has long been used in reactor design and engineering as an efficient simulation method for smaller problems, the transition to 3D has only begun recently, and to our knowledge no 3D MOC based codes are …

    mit Repository record for Development of the random ray method of neutral particle transport for high-fidelity nuclear reactor simulation (opens in a new tab)

  8. Coupled-Cluster Methods for Large Molecular Systems Through Massive Parallelism and Reduced-Scaling Approaches

    … computers equipped with conventional Intel Xeon processors and the Intel Xeon Phi (Knights Landing) processors. With the new implementation, the CCSD(T) energies can be evaluated for systems containing 200 electrons and 1000 basis functions in a few days using a small size commodity …

    vt Repository record for Coupled-Cluster Methods for Large Molecular Systems Through Massive Parallelism and Reduced-Scaling Approaches (opens in a new tab)