Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 20 of 57 for “"Memory bandwidth"”.
-
Techniques to maximize memory bandwidth on the Rigel compute accelerator
… to this space are performance limited by the memory all, so comparing the memory system performance of Rigel and GPUs is desirable. Memory controllers in GPUs attempt to coalesce memory requests from separate threads to achieve high off-chip bandwidth. This coalescing can be achieved by the …
-
A Reconfigurable, Distributed-Memory Accelerator for Sparse Applications
… On conventional systems, their irregular memory accesses and low arithmetic intensity create challenging memory bandwidth bottlenecks. To overcome such bottlenecks, distributed-SRAM architectures use tiled arrays of high-bandwidth local storage to achieve very high aggregate memory …
-
Architectural techniques to extend multi-core performance scaling
… face problems on two fronts: power and off-chip memory bandwidth. Dennard's scaling is effectively coming to an end which has lead to a gradual increase in chip power dissipation. In addition, sustaining off-chip memory bandwidth has become harder due to the limited space for pins on the die and …
-
Performance Modeling and Prediction for the Scalable Solution of Partial Differential Equations on Unstructured Grids
… kernels) on modern architectures with deep memory hierarchies. We identify that the primary factors responsible for this relatively poor performance are: insufficient available memory bandwidth, low ratio of work to data size (good algorithmic efficiency), and nonscaling cost of …
-
Energy-scalable speech recognition circuits
… to ASR hardware design. We identify external memory bandwidth as the main driver in system power consumption and select algorithms and architectures to minimize it. We evaluate three acoustic modeling approaches-Gaussian mixture models (GMMs), subspace GMMs (SGMMs), and deep neural networks …
-
Off-chip Communications Architectures For High Throughput Network Processors
… to increase the throughput of the currently used memory system. In recent years there is a significant increase in memory bandwidth demand on line cards as a result of higher line rates, an increase in deep packet inspection operations and an unstoppable expansion in lookup tables. As line-rate …
-
Unified RAW Path Oblivious RAM
… is a cryptographic primitive that conceals memory access patterns to untrusted storage. Its applications include oblivious cloud storage, trusted processors, software protection, secure multi-party computation, and so on. This thesis improves the state-of-the-art Path ORAM in several …
-
FleXilicon: a New Coarse-grained Reconfigurable Architecture for Multimedia and Wireless Communications
… through adoption of three schemes, (i) wider memory bandwidth, (ii) adoption of a reconfigurable controller, and (iii) flexible wordlength support. Increased memory bandwidth satisfies memory access requirement in LLP execution. New design of reconfigurable controller minimizes overhead in …
-
Making Computation on Encrypted Data Practical through Hardware Acceleration of Fully Homomorphic Encryption
… architectures cannot handle, as well as extreme memory bandwidth demands. This thesis presents two FHE accelerators that address these challenges: F1 and CraterLake. F1 is the őrst programmable FHE accelerator, i.e., capable of executing full FHE programs. F1 is a wide-vector processor with novel …
-
Energy-efficient smart embedded memory design for IoT and AI
Static Random Access Memory (SRAM) continues to be the embedded memory of choice for modern System-on-a-Chip (SoC) applications, thanks to aggressive CMOS scaling, which keeps on providing higher storage density per unit silicon area. As memory sizes continue to grow, increased bit-cell variation …
-
Micro-architectural analysis of SPACERAM processing element
… an architecture facilitates very high processor-memory bandwidth and hence allows for applications requiring orders of magnitude higher processing and update rates per DRAM than any current hardware. The array of processing elements process data coming simultaneously from several memory blocks by …
-
Data layout transformation through in-place transposition
… architectures due to limited available on-board memory capacity and high throughput. However, direct application of in-place transposition algorithms from CPU lacks the amount of parallelism and locality required by GPU to achieve good performance. In this thesis we present the first known …
-
Garbage Collection Scheduling for Utility Accrual Real-Time Systems
… on CPU overload, this dissertation explores memory overload conditions during which the aggregate memory demand exceeds a system's available memory bandwidth. Real-time systems are typically implemented in C or other languages that use explicit dynamic memory management. Taking advantage of …
-
Gaussian Splatting Device-Architecture Co-Design for Accelerated Physical AI Inference and Reasoning
… that are simultaneously photorealistic, memory-efficient, and computationally lean. Gaussian Splatting has established itself as the state-of-the-art technique for 3D scene representation, delivering photorealistic rendering quality that surpasses traditional methods. However, this …
-
Combustion Simulations Using Graphic Processing Units
… featuring high levels of parallelism and extreme memory bandwidth, which constitute a powerful computing platform to solve complex problems involving chemically reacting flows. In the present study, computer programs for combustion simulations with detailed chemical kinetic mechanisms were …
-
Investigating thermal dependence on monolithically-integrated photonic interconnects
… which has the promising potential to remove memory bandwidth bottleneck in the deep multicore regime. Although with the advantages of high bandwidth-density and energy-efficiency, it comes with design challenges from device, architecture and system perspectives. High thermal sensitivity of …
-
Quantization Methods for Matrix Multiplication and Efficient Transformers
… quantization of LLMs. Beyond reducing the memory footprint, quantization accelerates inference, as the primary bottleneck during autoregressive generation is often the memory bandwidth. NestQuant leverages two nested lattices to construct an efficient vector codebook for quantization, along …
-
Parallel implementations of probabilistic latent semantic analysis on graphic processing units
… because of GPU's multi-core structure and high memory bandwidth, but also because of the recent efforts devoted into building a programming framework to enable developers to easily manipulate GPU's computing power. In this paper, we introduced two methods to parallelize and speed up PLSA via …
-
Algorithms, architectures and circuits for low power HEVC codecs
… on two specific areas: Motion Compensation Bandwidth and Intra Estimation. HEVC uses larger filters for motion compensation leading to a significant increase in decoder bandwidth. We present a novel motion compensation cache that reduces external memory bandwidth by 67% and power by 40%. The …
-
Computing the fast Fourier transform on SIMD microprocessors
… algorithm has advantages in terms of memory bandwidth, and three implementations of this algorithm, which incorporate latency and spatial locality optimizations, are automatically vectorized at the algorithm level of abstraction. Performance results on 2- way, 4-way and 8-way SIMD …
Page 1 of 3