Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 13 of 13 for “"off-chip memory"”.
-
Architectural techniques to extend multi-core performance scaling
… they now face problems on two fronts: power and off-chip memory bandwidth. Dennard's scaling is effectively coming to an end which has lead to a gradual increase in chip power dissipation. In addition, sustaining off-chip memory bandwidth has become harder due to the limited space for pins on the …
-
Data layout transformation through in-place transposition
… architectures due to limited available on-board memory capacity and high throughput. However, direct application of in-place transposition algorithms from CPU lacks the amount of parallelism and locality required by GPU to achieve good performance. In this thesis we present the first known …
-
Energy-efficient video decoding using data statistics
… thesis, we design a hardware-based video decoder chip that exploits the statistics of the video to reduce the energy/pixel cost in several ways. For example, we exploit the sparsity in transform coefficients to reduce the energy/pixel cost of inverse transform by 29%. With the proposed …
-
Hierarchical and scalable bus architecture generation on FPGAs with high-level synthesis
… the system latency model by incorporating off-chip memory communication latency. The system is proved to be light-weight based on post-routing resource reports. Design space exploration among multilevel granularity parallelisms is performed to get the system's best performance, with which a …
-
An optical data receiver for integrated photonic interconnects
… able to communicate with each other as well as off-chip memory, global interconnects have become a major bottleneck. The solution has been proposed through integrated photonic networks, where multiple channels of information can be placed onto a single low-latency waveguide, reducing the number …
-
Balancing Performance, Area, and Power in an On-Chip Network
… the increasing gap between processor and off-chip memory speeds has constrained performance of memory-intensive applications. The Single-Chip Message Passing (SCMP) parallel computer sits at the confluence of these trends. SCMP is a tiled architecture consisting of numerous thread-parallel …
-
POWER-AWARE PERFORMANCE OPTIMIZATION ON MULTICORE ARCHITECTURES
… role to achieve optimal power-performance tradeoff because performance does not necessarily improve with increasing number of cores. Multi- threaded applications suffer due to thread synchronization, negative interference in shared memory including last level cache and main memory. Memory …
-
Low-power techniques for video decoding
… single-core decoder. To reduce the total memory system power, several caching techniques are demonstrated that can dramatically reduce the off-chip memory bandwidth and power at the cost of increased chip area. A 123 kB data-forwarding cache can reduce the read bandwidth from external …
-
Heterogeneous prioritization for network-on-chip based multi-core systems
In chip multi-processor (CMP) systems, communication and memory access both play an important role in influencing the performance achievable by the system. The manner in which the network packets (on-chip cache requests/responses) and off-chip memory bound packets are handled, in multi-core …
-
Thread Scheduling For Chip Multiprocessors
Large, high frequency single-core chip designs are increasingly being replaced with larger chip multiprocessor (CMP) designs that tradeoff frequency for greater numbers of cores. Power has become a first-order design constraint, leading to designs optimized for computing efficiency, defined as the …
-
Lab-to-Fab Monolithic 3D Integrated Carbon Nanotube Transistors: Scaling and Reliability
… is consumed moving data between compute and off-chip memory, which are often physically separate with limited connectivity. This is termed the “memory wall”. A promising solution to this problem is monolithic 3D integration, in which layers of compute and memory are designed and integrated …
-
A pixel-parallel architecture for graph cuts inference
… the physical restraints of hardware resource and memory size, and to accelerate graph cuts on larger images, we developed methods to divide images into virtual regions, which can then be efficiently processed on a physical processor array. We then design a complete memory system for the inference …
-
Improving the Off-chip Bandwidth Utilization and Energy Efficiency in Chip Multiprocessor (CMP) Architectures
This dissertation aims at improving the off-chip bandwidth utilization and energy efficiency in chip multiprocessor (CMP) architectures. This work consists of two main parts. The first part investigates the early write-back technique for a two-level cache hierarchy in a CMP with four processor …