Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 20 of 23 for “"Parallel efficiency"”.

  1. OpenMP-CUDA implementation of the moment method and multilevel fast multipole algorithm on multi-GPU computing systems

    … GPU computation based on the hybrid OpenMP-CUDA parallel programming model. The resultant algorithms are called the OpenMP-CUDA-MoM and the OpenMP-CUDA-MLFMA, respectively. Both of the proposed methods are applied to compute electromagnetic scattering by a three-dimensional conducting object. For …

    uiuc Repository record for OpenMP-CUDA implementation of the moment method and multilevel fast multipole algorithm on multi-GPU computing systems (opens in a new tab)

  2. A meshfree method for the Poisson equation with 3D wall-bounded flow application

    … has been studies based on numerical convergence, parallel efficiency, and computational cost. Practical application of the solver is presented in a simulation of a potential flow field in a wall-bounded domain.

    mit Repository record for A meshfree method for the Poisson equation with 3D wall-bounded flow application (opens in a new tab)

  3. Load Balancing Scientific Applications

    … levels are rapidly increasing. For ideal efficiency, developers of the simulations that run on these machines must ensure that computational work is evenly balanced among processors. Assigning work evenly is challenging because many large modern parallel codes simulate behavior of physical …

    tdl Repository record for Load Balancing Scientific Applications (opens in a new tab)

  4. A strategy for mapping unstructured mesh computational mechanics programs onto distributed memory parallel architectures

    … advantages offered by distributed memory parallel processors. Strategies that successfully map structured mesh codes onto parallel machines have been developed over the previous decade and used to build a toolkit for automation of the parallelisation process. Extension of the capabilities …

    greenwich Repository record for A strategy for mapping unstructured mesh computational mechanics programs onto distributed memory parallel architectures (opens in a new tab)

  5. Parallel algorithms for two-stage stochastic optimization

    … and/or there are integer variables in Stage 1. Parallelization of stochastic integer programs pose very unique characteristics that make them very challenging to parallelize. We develop a Parallel Stochastic Integer Program Solver (PSIPS) that exploits nested parallelism by exploring the …

    uiuc Repository record for Parallel algorithms for two-stage stochastic optimization (opens in a new tab)

  6. Innovative Methods for Solving Multicomponent Biogeochemical Groundwater Transport on Supercomputers

    … Alamos National Laboratory, on which a study of parallel efficiency/scaling using up to 512 processors was conducted. Recently, researchers have been investigating emplacement of carbon sources to create anaerobic conditions and enhance in situ reductive dechlorination of halogenated NAPLs. For …

    uiuc Repository record for Innovative Methods for Solving Multicomponent Biogeochemical Groundwater Transport on Supercomputers (opens in a new tab)

  7. A Study of MLFMA for Large -Scale Scattering Problems

    … of the translation operator in 3-D is shown. A parallel implementation of the multilevel fast multipole algorithm (MLFMA) was studied as far as parallel efficiency and scaling. The large-scale scattering program (LSSP), based on the ScaleME library, was used to solve ultra-large-scale problems …

    uiuc Repository record for A Study of MLFMA for Large -Scale Scattering Problems (opens in a new tab)

  8. The solution of large viscoelastic flow problems using parallel iterative techniques

    Efficient parallel computation of complex viscoelastic flows is essential to bring modem computer power to bear on fluid calculations where complicated constitutive descriptions are required. We have developed an efficient, parallel time integration scheme for calculating viscoelastic flows. The …

    mit Repository record for The solution of large viscoelastic flow problems using parallel iterative techniques (opens in a new tab)

  9. An Asynchronous Ensemble-Averaging Approach to CMFD Source Acceleration: Rearchitecting Monte Carlo Reactor Simulation Paradigms for the Exasacle Computing Age

    … Computing (HPC) systems and significant code parallelization of Monte Carlo algorithms have helped to reduce overall time to solution. This thesis extends these efforts by proposing a novel ensemble averaged Coarse Mesh Finite Difference (CMFD) simulation strategy in order to improve parallel

    mit Repository record for An Asynchronous Ensemble-Averaging Approach to CMFD Source Acceleration: Rearchitecting Monte Carlo Reactor Simulation Paradigms for the Exasacle Computing Age (opens in a new tab)

  10. Development of a Parallel Electrostatic PIC Code for Modeling Electric Propulsion

    This thesis presents the parallel version of Coliseum, the Air Force Research Laboratory plasma simulation framework. The parallel code was designed to run large simulations on the world fastest supercomputers as well as home mode clusters. Plasma simulations are extremely computationally intensive …

    vt Repository record for Development of a Parallel Electrostatic PIC Code for Modeling Electric Propulsion (opens in a new tab)

  11. PICS - a Performance-analysis-based Introspective Control System to steer parallel applications

    Parallel programming has always been difficult due to the complexity of hardware and the diversity of applications. Although significant progress has been achieved over the years, attaining high parallel efficiency on large supercomputers for various applications is still quite challenging. As we …

    uiuc Repository record for PICS - a Performance-analysis-based Introspective Control System to steer parallel applications (opens in a new tab)

  12. Advanced finite-element techniques for simulation of composite materials and large-scale scattering problems

    … of electrically large problems on massive parallelized computers, and efficient broadband analysis of resonant waveguide structures. To these ends, first, an interface-enriched generalized finite-element method (IGFEM) is introduced for electromagnetic analysis of heterogeneous materials. …

    uiuc Repository record for Advanced finite-element techniques for simulation of composite materials and large-scale scattering problems (opens in a new tab)

  13. Parallelization of the Euler Equations on Unstructured Grids

    … are investigated on two distributed-memory parallel computers using an explicit message-passing paradigm: these are classic Euler Explicit, four-stage Jameson-style Runge-Kutta, Block Jacobi, Block Gauss-Seidel, and Block Symmetric Gauss-Seidel. A finite-volume formulation is used for the …

    vt Repository record for Parallelization of the Euler Equations on Unstructured Grids (opens in a new tab)

  14. Development and application of a parallel compositional reservoir simulator

    … are widely available. These systems allow parallel processing, which helps large-scale simulations run faster and more efficient. In this research project, we developed a parallel version of The University of Texas Compositional Simulator (UTCOMP). The parallel UTCOMP is capable of running …

    texas Repository record for Development and application of a parallel compositional reservoir simulator (opens in a new tab)

  15. Advances in an open-source direct simulation Monte Carlo technique for hypersonic rarefied gas flows

    … on the number of particles in a cell. Finally, parallel efficiency tests of dsmcFoam are presented in this thesis along with a new domain decomposition technique for parallel processing. This technique splits up the computational domain based on the number of particles, such that each processor …

    strathclyde Repository record for Advances in an open-source direct simulation Monte Carlo technique for hypersonic rarefied gas flows (opens in a new tab)

  16. Acceleration Techniques for Industrial Large Eddy Simulation with High-Order Methods on CPU-GPU Clusters

    … preconditioning significantly increased the efficiency of the p-multigrid method. For unsteady simulations, the preconditioner helped with the efficiency of the p-multigrid with larger physical time steps. In most steady cases, the preconditioned p-multigrid approach is comparable to or …

    ku Repository record for Acceleration Techniques for Industrial Large Eddy Simulation with High-Order Methods on CPU-GPU Clusters (opens in a new tab)

  17. Numerical Modelling of Sooting Laminar Diffusion Flames at Elevated Pressures and Microgravity

    … second-order accurate finite-volume scheme and a parallel adaptive mesh refinement algorithm on body-fitted, multi-block meshes. An acetylene-based, semi-empirical model was used to predict the nucleation, growth, and oxidation of soot particles. Reasonable agreement with experimental measurements …

    toronto-retro Repository record for Numerical Modelling of Sooting Laminar Diffusion Flames at Elevated Pressures and Microgravity (opens in a new tab)

  18. Performance Optimization and Parallelization of Turbo Decoding for Software-Defined Radio

    … This thesis describes the optimization, parallelization, and simulated performance of a software double-binary turbo decoder implementation supporting the WiMAX standard suitable for SDR. An adapted turbo decoder is implemented in the C language, and numerous software optimizations are …

    queens Repository record for Performance Optimization and Parallelization of Turbo Decoding for Software-Defined Radio (opens in a new tab)

  19. Energetics of neutral interstitials in germanium

    … portions of defect systems to improve DMC efficiency. The best results for the hexagonal structure achieve a 3x speedup over standard total energy calculations. No speedup is obtained for the tetragonal and split interstitial systems, perhaps due to the increased pressure in the fixed …

    uiuc Repository record for Energetics of neutral interstitials in germanium (opens in a new tab)

  20. Domain Decomposition Preconditioners for Hermite Collocation Problems

    … convergence rate of Krylov subspace methods with parallelizable preconditioners is essential for obtaining effective iterative solvers for very large linear systems of equations. Substructuring provides a framework for constructing robust and parallel preconditioners for linear systems arising …

    vt Repository record for Domain Decomposition Preconditioners for Hermite Collocation Problems (opens in a new tab)

Page 1 of 2