Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 16 of 16 for “"Accelerator architecture"”.
-
Flexible And Efficient Accelerator Architecture For Runtime Monitoring
Secure and reliable computing remains an open problem. Just in the past year, security vulnerabilities continued to be discovered in widely-used applications, and remote attacks that take advantage of these vulnerabilities can cause an enormous amount of damage. Runtime monitoring is a promising …
-
An Energy and Area Estimation Plugin for Accelerator Architecture Simulation
Development of domain-specific hardware accelerators has been an important focus for high performance computing research in recent years, enabling significant gains in a variety of practical applications. Of particular interest is accelerator design for applications involving sparse data. Such …
-
An evaluation framework for massively parallel accelerator processors
… framework for Rigel, a 1024-core single-chip accelerator architecture designed for high throughput on visual computing and scientific workloads. I present an integrated evaluation framework for investigating co-designed architecture, compilers, programming models, and RTL implementation for …
-
Mixed-precision architecture for flexible neural network accelerators
… This thesis suggests a novel neural network accelerator architecture that handles multiple bit precisions for both weights and activations. The architecture is based on a fused spatial and temporal micro-architecture that maximizes both bandwidth eciency and computational ability. …
-
Tradeoffs in Designing Massively Parallel Accelerator Architectures
… an integrated optimization of both the micro-architecture and the physical circuit design of the cores and caches. In our approach, we use statistical sampling of the design space for evaluating the performance of the micro-architecture and RTL synthesis to characterize the area-power-delay of …
-
Achieving performance portability across parallel accelerator architectures
… codes typically only perform well for a single architecture, or even microarchitecture. This thesis focuses on SPMD code written in CUDA, noting that programs must obey a number of constraints to achieve high performance on an NVIDIA GPU. Under such constraints, source-level optimizations can …
-
Spatial Accelerator Generation and Optimization for Tensor Applications
… which increases the demand for flexible accelerator architecture. Existing frameworks suffer from the trade-off between design flexibility and productivity of RTL generation: either limited to very few hand-written templates or cannot automatically generate the RTL. To address this …
-
Accelerated deep learning for the edge-to-cloud continuum: A specialized full stack derived from algorithms
Advances in high-performance computer architecture design have been a major driver for the rapid evolution of Deep Neural Networks (DNN). Due to their insatiable demand for compute power, naturally, both the research community as well the industry have turned to accelerators to accommodate modern …
-
Investigating Opportunities and Challenges in Modeling and Designing Scale-Out DNN Accelerators
… performance vs. generality seen in most CPU architectures today. In most situations, hardware engineers are left to construct systems that serve the needs of various applications, often needing to predict the use-cases of the system. As with any field, the ability to predict and act on the …
-
Efficient Computation of Map-scale Continuous Mutual Information on Chip in Real Time
… In this thesis, we introduce a new hardware accelerator architecture for MI computation that features a high-efficiency MI compute core and an optimized memory subsystem that provides sufficient bandwidth to keep the cores fully utilized. The core employs interleaving to counter the recursive …
-
Design and prototyping of Hardware-Accelerated Locality-aware Memory Compression
… units in the chip. This thesis describes the accelerator architecture, implementation, and prototype for one of such functions namely "Locality-Aware memory compression" which is part of the "OS-controlled memory compression" scheme that has been actively deployed in today's OSes. In brief, …
-
In-Memory Computing Architecture for Deep Learning Acceleration
… multi-core systems that make intensive use of accelerators become the future of computing. Such novel computing systems incurs new challenges including architectural support for model training in the accelerators, large cache demands for multi-core processors, system performance, energy, and …
-
Multithreaded architectures for manycore throughput processors
This dissertation describes work on the architecture of throughput-oriented accelerator processors. First, we examine the limitations of current accelerator processors and identify an opportunity to enable high throughput while also providing a more general-purpose programming model.To address this …
-
Model-Architecture Co-design of Deep Neural Networks for Embedded Systems
… evolving nature of modern deep network architecture aggravates the problem by making it necessary to balance flexibility against specialisation to avoid the inability to adapt. However, much of the baseline architecture of a deep convolutional neural network stayed the same. With careful …
-
Evaluating the OpenACC API for Parallelization of CFD Applications
… simplify the transition of legacy codes to new architectures. This work analyzes the popular OpenACC programming standard, as implemented by the PGI compiler suite, in order to evaluate its utility and performance potential in computational fluid dynamics (CFD) applications. Of particular …
-
Hybrid coherence for scalable multicore architectures
Made available in DSpace on 2011-01-14T22:43:29Z (GMT). No. of bitstreams: 2 Kelm_John.pdf: 2626023 bytes, checksum: 069d102e87802f19ba796da9758e8be9 (MD5) license.txt: 4057 bytes, checksum: 83ce7b997e74195078c3d8fca669fdd2 (MD5)