Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 20 of 57 for “"Memory bandwidth"”.

  1. Techniques to maximize memory bandwidth on the Rigel compute accelerator

    … to this space are performance limited by the memory all, so comparing the memory system performance of Rigel and GPUs is desirable. Memory controllers in GPUs attempt to coalesce memory requests from separate threads to achieve high off-chip bandwidth. This coalescing can be achieved by the …

    uiuc Repository record for Techniques to maximize memory bandwidth on the Rigel compute accelerator (opens in a new tab)

  2. A Reconfigurable, Distributed-Memory Accelerator for Sparse Applications

    … On conventional systems, their irregular memory accesses and low arithmetic intensity create challenging memory bandwidth bottlenecks. To overcome such bottlenecks, distributed-SRAM architectures use tiled arrays of high-bandwidth local storage to achieve very high aggregate memory

    mit Repository record for A Reconfigurable, Distributed-Memory Accelerator for Sparse Applications (opens in a new tab)

  3. Architectural techniques to extend multi-core performance scaling

    … face problems on two fronts: power and off-chip memory bandwidth. Dennard's scaling is effectively coming to an end which has lead to a gradual increase in chip power dissipation. In addition, sustaining off-chip memory bandwidth has become harder due to the limited space for pins on the die and …

    purdue-thes Repository record for Architectural techniques to extend multi-core performance scaling (opens in a new tab)

  4. Performance Modeling and Prediction for the Scalable Solution of Partial Differential Equations on Unstructured Grids

    … kernels) on modern architectures with deep memory hierarchies. We identify that the primary factors responsible for this relatively poor performance are: insufficient available memory bandwidth, low ratio of work to data size (good algorithmic efficiency), and nonscaling cost of …

    odu Repository record for Performance Modeling and Prediction for the Scalable Solution of Partial Differential Equations on Unstructured Grids (opens in a new tab)

  5. Energy-scalable speech recognition circuits

    … to ASR hardware design. We identify external memory bandwidth as the main driver in system power consumption and select algorithms and architectures to minimize it. We evaluate three acoustic modeling approaches-Gaussian mixture models (GMMs), subspace GMMs (SGMMs), and deep neural networks …

    mit Repository record for Energy-scalable speech recognition circuits (opens in a new tab)

  6. Off-chip Communications Architectures For High Throughput Network Processors

    … to increase the throughput of the currently used memory system. In recent years there is a significant increase in memory bandwidth demand on line cards as a result of higher line rates, an increase in deep packet inspection operations and an unstoppable expansion in lookup tables. As line-rate …

    ucf

  7. Unified RAW Path Oblivious RAM

    … is a cryptographic primitive that conceals memory access patterns to untrusted storage. Its applications include oblivious cloud storage, trusted processors, software protection, secure multi-party computation, and so on. This thesis improves the state-of-the-art Path ORAM in several …

    mit Repository record for Unified RAW Path Oblivious RAM (opens in a new tab)

  8. FleXilicon: a New Coarse-grained Reconfigurable Architecture for Multimedia and Wireless Communications

    … through adoption of three schemes, (i) wider memory bandwidth, (ii) adoption of a reconfigurable controller, and (iii) flexible wordlength support. Increased memory bandwidth satisfies memory access requirement in LLP execution. New design of reconfigurable controller minimizes overhead in …

    vt Repository record for FleXilicon: a New Coarse-grained Reconfigurable Architecture for Multimedia and Wireless Communications (opens in a new tab)

  9. Making Computation on Encrypted Data Practical through Hardware Acceleration of Fully Homomorphic Encryption

    … architectures cannot handle, as well as extreme memory bandwidth demands. This thesis presents two FHE accelerators that address these challenges: F1 and CraterLake. F1 is the őrst programmable FHE accelerator, i.e., capable of executing full FHE programs. F1 is a wide-vector processor with novel …

    mit Repository record for Making Computation on Encrypted Data Practical through Hardware Acceleration of Fully Homomorphic Encryption (opens in a new tab)

  10. Energy-efficient smart embedded memory design for IoT and AI

    Static Random Access Memory (SRAM) continues to be the embedded memory of choice for modern System-on-a-Chip (SoC) applications, thanks to aggressive CMOS scaling, which keeps on providing higher storage density per unit silicon area. As memory sizes continue to grow, increased bit-cell variation …

    mit Repository record for Energy-efficient smart embedded memory design for IoT and AI (opens in a new tab)

  11. Micro-architectural analysis of SPACERAM processing element

    … an architecture facilitates very high processor-memory bandwidth and hence allows for applications requiring orders of magnitude higher processing and update rates per DRAM than any current hardware. The array of processing elements process data coming simultaneously from several memory blocks by …

    mit Repository record for Micro-architectural analysis of SPACERAM processing element (opens in a new tab)

  12. Data layout transformation through in-place transposition

    … architectures due to limited available on-board memory capacity and high throughput. However, direct application of in-place transposition algorithms from CPU lacks the amount of parallelism and locality required by GPU to achieve good performance. In this thesis we present the first known …

    uiuc Repository record for Data layout transformation through in-place transposition (opens in a new tab)

  13. Garbage Collection Scheduling for Utility Accrual Real-Time Systems

    … on CPU overload, this dissertation explores memory overload conditions during which the aggregate memory demand exceeds a system's available memory bandwidth. Real-time systems are typically implemented in C or other languages that use explicit dynamic memory management. Taking advantage of …

    vt Repository record for Garbage Collection Scheduling for Utility Accrual Real-Time Systems (opens in a new tab)

  14. Gaussian Splatting Device-Architecture Co-Design for Accelerated Physical AI Inference and Reasoning

    … that are simultaneously photorealistic, memory-efficient, and computationally lean. Gaussian Splatting has established itself as the state-of-the-art technique for 3D scene representation, delivering photorealistic rendering quality that surpasses traditional methods. However, this …

    uic

  15. Combustion Simulations Using Graphic Processing Units

    … featuring high levels of parallelism and extreme memory bandwidth, which constitute a powerful computing platform to solve complex problems involving chemically reacting flows. In the present study, computer programs for combustion simulations with detailed chemical kinetic mechanisms were …

    uconn-diss Repository record for Combustion Simulations Using Graphic Processing Units (opens in a new tab)

  16. Investigating thermal dependence on monolithically-integrated photonic interconnects

    … which has the promising potential to remove memory bandwidth bottleneck in the deep multicore regime. Although with the advantages of high bandwidth-density and energy-efficiency, it comes with design challenges from device, architecture and system perspectives. High thermal sensitivity of …

    mit Repository record for Investigating thermal dependence on monolithically-integrated photonic interconnects (opens in a new tab)

  17. Quantization Methods for Matrix Multiplication and Efficient Transformers

    … quantization of LLMs. Beyond reducing the memory footprint, quantization accelerates inference, as the primary bottleneck during autoregressive generation is often the memory bandwidth. NestQuant leverages two nested lattices to construct an efficient vector codebook for quantization, along …

    mit Repository record for Quantization Methods for Matrix Multiplication and Efficient Transformers (opens in a new tab)

  18. Parallel implementations of probabilistic latent semantic analysis on graphic processing units

    … because of GPU's multi-core structure and high memory bandwidth, but also because of the recent efforts devoted into building a programming framework to enable developers to easily manipulate GPU's computing power. In this paper, we introduced two methods to parallelize and speed up PLSA via …

    uiuc Repository record for Parallel implementations of probabilistic latent semantic analysis on graphic processing units (opens in a new tab)

  19. Algorithms, architectures and circuits for low power HEVC codecs

    … on two specific areas: Motion Compensation Bandwidth and Intra Estimation. HEVC uses larger filters for motion compensation leading to a significant increase in decoder bandwidth. We present a novel motion compensation cache that reduces external memory bandwidth by 67% and power by 40%. The …

    mit Repository record for Algorithms, architectures and circuits for low power HEVC codecs (opens in a new tab)

  20. Computing the fast Fourier transform on SIMD microprocessors

    … algorithm has advantages in terms of memory bandwidth, and three implementations of this algorithm, which incorporate latency and spatial locality optimizations, are automatically vectorized at the algorithm level of abstraction. Performance results on 2- way, 4-way and 8-way SIMD …

    waikato-masters Repository record for Computing the fast Fourier transform on SIMD microprocessors (opens in a new tab)

Page 1 of 3