Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 8 of 8 for “"mixed precision"”.

  1. Mixed-precision architecture for flexible neural network accelerators

    … architecture that handles multiple bit precisions for both weights and activations. The architecture is based on a fused spatial and temporal micro-architecture that maximizes both bandwidth eciency and computational ability. Furthermore, this thesis presents an FPGA implementation of …

    mit Repository record for Mixed-precision architecture for flexible neural network accelerators (opens in a new tab)

  2. Mixed-precision NN accelerator with neural-hardware architecture search

    … designs are challenging. We first propose a mixed-precision accelerator, a highly parameterized architecture that can adapt to different bit widths for different quantized layers with significantly reduced overhead. It efficiently provides a vast design space for both neural and hardware …

    mit Repository record for Mixed-precision NN accelerator with neural-hardware architecture search (opens in a new tab)

  3. Characterizing the Energy Requirement of Computer Vision

    … Findings include that adjusting from single precision model to a mixed precision model can result in energy reductions of nearly 40%. Additionally power capping the GPU can reduce energy cost by an additional 10%.

    mit Repository record for Characterizing the Energy Requirement of Computer Vision (opens in a new tab)

  4. GPU-accelerated Linear Solvers for High-order Finite Element Methods in Poisson Problems

    … reducing the number of iterations. An adaptive mixed precision conjugate gradient algorithm is proposed to exploit the superior computational performance of GPUs at lower precisions while maintaining convergence and accuracy. Furthermore, a comprehensive optimization of the PCG algorithm, …

    vt Repository record for GPU-accelerated Linear Solvers for High-order Finite Element Methods in Poisson Problems (opens in a new tab)

  5. Efficient methods for mapping neural machine translator on FPGAs

    … beam search algorithm. We quantize the model to mixed-precision representation in which parameters and portions of calculations are in 16-bit half precision, and others remain as 32-bit floating-point. Compared to the float NMT implementation on FPGA, we achieve 13.1x speedup with end-to-end …

    uiuc Repository record for Efficient methods for mapping neural machine translator on FPGAs (opens in a new tab)

  6. Efficient Deep Learning Systems for Visual Perception on the Edge

    … and is 1.2-1.3× faster than SpConv v2 in mixed precision training across seven representative autonomous driving benchmarks. It also seamlessly supports graph convolutions, achieving 2.6-7.6× faster inference speed compared with state-of-the-art graph deep learning libraries. Furthermore, …

    mit Repository record for Efficient Deep Learning Systems for Visual Perception on the Edge (opens in a new tab)

  7. Resource-efficient optimizations of 3D vision models for segmentation and detection

    Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2027-05-01

    uiuc Repository record for Resource-efficient optimizations of 3D vision models for segmentation and detection (opens in a new tab)