Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 8 of 8 for “"CUDA implementation"”.
-
OpenMP-CUDA implementation of the moment method and multilevel fast multipole algorithm on multi-GPU computing systems
… for GPU computation based on the hybrid OpenMP-CUDA parallel programming model. The resultant algorithms are called the OpenMP-CUDA-MoM and the OpenMP-CUDA-MLFMA, respectively. Both of the proposed methods are applied to compute electromagnetic scattering by a three-dimensional conducting …
-
Optimizing Harris Corner Detection on GPGPUs Using CUDA
… Harris Corner Detection on GPGPUs Using CUDA</p> <p>The objective of this thesis is to optimize the Harris corner detection algorithm implementation on NVIDIA GPGPUs using the CUDA software platform and measure the performance benefit. The Harris corner detection algorithm—developed by C. …
-
Real-Time Smoothed Particle Hydrodynamics For CUDA
… NVIDIA's Compute Unifed Device Architecture (CUDA) has enabled developers to harness the massive computational power of the GPU through the CUDA parallel programming model. Smoothed particle hydrodynamics (SPH) simulations are particularly well suited to CUDA implementation due to the high …
-
Real-Time Smoothed Particle Hydrodynamics For CUDA
… NVIDIA's Compute Unifed Device Architecture (CUDA) has enabled developers to harness the massive computational power of the GPU through the CUDA parallel programming model. Smoothed particle hydrodynamics (SPH) simulations are particularly well suited to CUDA implementation due to the high …
-
Efficient Verifiable Computation Made Easy
… on GPU over a highly optimized serial C++ implementation on CPU and an existing multithreaded Rust baseline on CPU, respectively. Compared to our hand-optimized GPU/CUDA implementation requiring an extra 2,000 lines of low-level code (roughly 60 programmer hours), our compiler-generated GPU …
-
GPU-based acceleration of radio interferometry point source visibility simulations in the MeqTrees framework
… as PSV simulations. This thesis presents a GPU/CUDA implementation of the Point Source Visibility calculation within the existing MeqTrees framework. For a large number of sources, this implementation achieves an 18x speed-up over the existing CPU module. With modications to the MeqTrees memory …
-
Hardware Acceleration for Real-Time Compression of 3D Gaussians
… on current hardware, resulting in an optimized CUDA implementation with better than 100× the throughput of prior work and achieving real-time operation on workstation-class hardware. However, after concluding that custom hardware is necessary for further improvement, this thesis also presents a …
-
Improved GPU implementations of the Pair-HMM forward algorithm for DNA sequence alignment
The student, Enliang Li, accepted the attached license on 2021-04-30 at 13:30.