Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 8 of 8 for “"mixed precision"”.
-
Mixed-precision architecture for flexible neural network accelerators
… architecture that handles multiple bit precisions for both weights and activations. The architecture is based on a fused spatial and temporal micro-architecture that maximizes both bandwidth eciency and computational ability. Furthermore, this thesis presents an FPGA implementation of …
-
Mixed-precision NN accelerator with neural-hardware architecture search
… designs are challenging. We first propose a mixed-precision accelerator, a highly parameterized architecture that can adapt to different bit widths for different quantized layers with significantly reduced overhead. It efficiently provides a vast design space for both neural and hardware …
-
Characterizing the Energy Requirement of Computer Vision
… Findings include that adjusting from single precision model to a mixed precision model can result in energy reductions of nearly 40%. Additionally power capping the GPU can reduce energy cost by an additional 10%.
-
GPU-accelerated Linear Solvers for High-order Finite Element Methods in Poisson Problems
… reducing the number of iterations. An adaptive mixed precision conjugate gradient algorithm is proposed to exploit the superior computational performance of GPUs at lower precisions while maintaining convergence and accuracy. Furthermore, a comprehensive optimization of the PCG algorithm, …
-
Efficient methods for mapping neural machine translator on FPGAs
… beam search algorithm. We quantize the model to mixed-precision representation in which parameters and portions of calculations are in 16-bit half precision, and others remain as 32-bit floating-point. Compared to the float NMT implementation on FPGA, we achieve 13.1x speedup with end-to-end …
-
Efficient Deep Learning Systems for Visual Perception on the Edge
… and is 1.2-1.3× faster than SpConv v2 in mixed precision training across seven representative autonomous driving benchmarks. It also seamlessly supports graph convolutions, achieving 2.6-7.6× faster inference speed compared with state-of-the-art graph deep learning libraries. Furthermore, …
-
Accelerating Quantized DNNs with Dedicated Hardware Accelerators and RISC-V Processors Using Precision-Scalable Multipliers
L'abstract è presente nell'allegato / the abstract is in the attachment
-
Resource-efficient optimizations of 3D vision models for segmentation and detection
Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2027-05-01