Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 5 of 5 for “"FP16"”.

  1. Design and optimization of an embedded machine learning image processing system with a Linux SOC for large-scale livestock tallying

    … these custom-trained architectures for FP32, FP16, and INT8 quantization formats, enabling further performance assessment on the selected NVIDIA Jetson Orin Nano platform. The YOLOv9’s large model variant with FP16 quantization proved superior and achieved a mean average precision (mAP) across …

    stellenbosch Repository record for Design and optimization of an embedded machine learning image processing system with a Linux SOC for large-scale livestock tallying (opens in a new tab)

  2. Efficient Segment Anything on the Edge

    … employing inference engines, and applying FP16 and INT8 quantization, we achieve a 24x speedup relative to the baseline FP32 PyTorch implementation. GazeSAM runs at a speed of over 30 FPS, enabling real-time performance on an RTX 4070 GPU.

    mit Repository record for Efficient Segment Anything on the Edge (opens in a new tab)

  3. Comparing the Performance of Small Word-Size Floating-Point Numerics to Fixed-Point Numerics in Neural Networks

    … floating-point formats (FP8, FP12, FP16) for dynamic, layer-wise quantization during training. APT monitors per-layer Quantization Error Measurement (QEM) to guide precision adjustments and incorporates a novel bit-shuffling mechanism to reallocate bits between exponent and mantissa …

    gatech Repository record for Comparing the Performance of Small Word-Size Floating-Point Numerics to Fixed-Point Numerics in Neural Networks (opens in a new tab)

  4. OS INSPIRED COMPLETE KERNEL FUSION

    … FlashDMoE using FP32 while baselines use FP16.

    cornell Repository record for OS INSPIRED COMPLETE KERNEL FUSION (opens in a new tab)

  5. Efficient Deep Learning Systems for Visual Perception on the Edge

    … offers 2.7-3.7× speedup over the Huggingface FP16 implementation on both desktop and mobile GPUs. It also enables the deployment of the 70B Llama-2 model on mobile GPUs. Together, these techniques significantly reduce the computational and memory costs for deploying deep learning models on …

    mit Repository record for Efficient Deep Learning Systems for Visual Perception on the Edge (opens in a new tab)