Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 14 of 14 for “"Xeon Phi"”.

  1. An Optimized Multiple Right-Hand Side Dslash Kernel for Intel Xeon Phi

    … point coprocessor cards such as GPUs and Intel Xeon Phi Knights Corner (KNC), which boast powerful vector processing units. Most of these efforts in the area of Dslash have focused on single right-hand side solvers. This thesis will present two optimized Dslash kernels which simplify …

    odu Repository record for An Optimized Multiple Right-Hand Side Dslash Kernel for Intel Xeon Phi (opens in a new tab)

  2. Perfilagem do problema de resolução da equação da onda por diferenças finitas em coprocessador Xeon Phi

    Software Optimization is a promising technique to scientific computing applications, which can have calculations that may require several hours or even days. The idea of this work is to validate, using performance metrics, the optimization of a finite differences code that solve the wave equation …

    brazil-uerj Repository record for Perfilagem do problema de resolução da equação da onda por diferenças finitas em coprocessador Xeon Phi (opens in a new tab)

  3. Simulating Nonlinear Neutrino Oscillations on Next-Generation Many-Core Architectures

    … It can run on both the CPU and the Xeon Phi co-processor, the latter of which is based on the Intel Many Integrated Core Architecture (MIC). The performance of XFLAT on various system configurations and physics scenarios has been analyzed. In addition, the impact of I/O and the …

    unm Repository record for Simulating Nonlinear Neutrino Oscillations on Next-Generation Many-Core Architectures (opens in a new tab)

  4. Exploring Performance Portability for Accelerators via High-level Parallel Patterns

    … e.g., multi-core CPUs, many-core GPUs (Graphics Processing Units) and Intel Xeon Phi. The performance gains from them can be as high as many orders of magnitude, attracting extensive interest from many scientific domains. However, the gains are closely followed by two main problems: (1) A …

    vt Repository record for Exploring Performance Portability for Accelerators via High-level Parallel Patterns (opens in a new tab)

  5. On implementing sparse matrix-vector multiplication on intel platform

    … kernel that realizes high performance on Intel Xeon multicore and Phi processors for unstructured matrices. CCF kernel exploits the properties of CCF to enhance load balancing and SIMD efficiency. Moreover, we present the CCF auto-tuner that selects the most effective parameters and the SpMV …

    uiuc Repository record for On implementing sparse matrix-vector multiplication on intel platform (opens in a new tab)

  6. On-node performance optimization of a Monte-Carlo transport code for leadership architectures

    … on modern Intel micro-architectures (i.e., Intel Xeon Phi and Intel Xeon Platinum 8180) to understand what hardware and settings configurations were optimal. The specific modules and subroutines that were responsible for the performance drop were also highlighted. The first round of optimizations …

    mit Repository record for On-node performance optimization of a Monte-Carlo transport code for leadership architectures (opens in a new tab)

  7. A Compiler Framework to Support and Exploit Heterogeneous Overlapping-ISA Multiprocessor Platforms

    … (in our case a platform built with Intel Xeon - Xeon Phi). With the introduced Profiler, Partitioner, and Runtime support, we prove to be able to automatically exploit the heterogeneity in an overlapping-ISA platform, being faster than native execution and other parallelism programming …

    vt Repository record for A Compiler Framework to Support and Exploit Heterogeneous Overlapping-ISA Multiprocessor Platforms (opens in a new tab)

  8. Performance analysis and optimization of a CFD application

    … on a dual-socket node fitted with an Intel Xeon Phi MIC card. To reduce the overhead created by host-accelerator copies in heterogeneous execution, the data layout of the halo region was changed from a ''star'' shape to a ''box'' shape to agglomerate small communications and to create a …

    uiuc Repository record for Performance analysis and optimization of a CFD application (opens in a new tab)

  9. Modeling and Runtime Systems for Coordinated Power-Performance Management

    … on Intel x86 systems with accelerators of Intel Xeon Phi and a Nvidia general-purpose graphics processing unit (GPGPU). We show the trade-offs and potentials for improving efficiency. Furthermore, we propose a parallel performance model for coordinating DVFS, DMT, and DCT simultaneously. We …

    vt Repository record for Modeling and Runtime Systems for Coordinated Power-Performance Management (opens in a new tab)

  10. Directive-Based Data Partitioning and Pipelining and Auto-Tuning for High-Performance GPU Computing

    … potential of accelerators, such as graphics processing units (GPUs), field-programmable gate arrays (FPGAs), and co-processors (e.g., Intel Xeon Phi), due to their increasing use in state-of-the-art supercomputers. Over the past 10 years, we have seen a significant improvement in both …

    vt Repository record for Directive-Based Data Partitioning and Pipelining and Auto-Tuning for High-Performance GPU Computing (opens in a new tab)

  11. Productive Programming Systems for Heterogeneous Supercomputers

    … of NVIDIA Pascal GPUs, Intel Knights Landing Xeon Phi processors, the Epiphany Co-Processor, the Sunway MPP, and other throughput-oriented architectures that enable pre-exascale computing. However, while the majority of the FLOPS in designs for future HPC systems come from throughput-oriented …

    rice Repository record for Productive Programming Systems for Heterogeneous Supercomputers (opens in a new tab)

  12. Development of the random ray method of neutral particle transport for high-fidelity nuclear reactor simulation

    … computing systems, including CPU, GPU, and Intel Xeon Phi architectures. While 2D MOC has long been used in reactor design and engineering as an efficient simulation method for smaller problems, the transition to 3D has only begun recently, and to our knowledge no 3D MOC based codes are currently …

    mit Repository record for Development of the random ray method of neutral particle transport for high-fidelity nuclear reactor simulation (opens in a new tab)

  13. Coupled-Cluster Methods for Large Molecular Systems Through Massive Parallelism and Reduced-Scaling Approaches

    … computers equipped with conventional Intel Xeon processors and the Intel Xeon Phi (Knights Landing) processors. With the new implementation, the CCSD(T) energies can be evaluated for systems containing 200 electrons and 1000 basis functions in a few days using a small size commodity cluster, …

    vt Repository record for Coupled-Cluster Methods for Large Molecular Systems Through Massive Parallelism and Reduced-Scaling Approaches (opens in a new tab)

  14. Neuromorphic Learning Systems for Supervised and Unsupervised Applications

    … the large-scale implementation of neuromorphic learning models and pushed the research on computational intelligence into a new era. Those bio-inspired models are constructed on top of unified building blocks, i.e. neurons, and have revealed potentials for learning of complex information. …

    syracuse-diss Repository record for Neuromorphic Learning Systems for Supervised and Unsupervised Applications (opens in a new tab)