Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 5 of 5 for “"FP16"”.
-
Design and optimization of an embedded machine learning image processing system with a Linux SOC for large-scale livestock tallying
… these custom-trained architectures for FP32, FP16, and INT8 quantization formats, enabling further performance assessment on the selected NVIDIA Jetson Orin Nano platform. The YOLOv9’s large model variant with FP16 quantization proved superior and achieved a mean average precision (mAP) across …
-
Efficient Segment Anything on the Edge
… employing inference engines, and applying FP16 and INT8 quantization, we achieve a 24x speedup relative to the baseline FP32 PyTorch implementation. GazeSAM runs at a speed of over 30 FPS, enabling real-time performance on an RTX 4070 GPU.
-
Comparing the Performance of Small Word-Size Floating-Point Numerics to Fixed-Point Numerics in Neural Networks
… floating-point formats (FP8, FP12, FP16) for dynamic, layer-wise quantization during training. APT monitors per-layer Quantization Error Measurement (QEM) to guide precision adjustments and incorporates a novel bit-shuffling mechanism to reallocate bits between exponent and mantissa …
-
OS INSPIRED COMPLETE KERNEL FUSION
… FlashDMoE using FP32 while baselines use FP16.
-
Efficient Deep Learning Systems for Visual Perception on the Edge
… offers 2.7-3.7× speedup over the Huggingface FP16 implementation on both desktop and mobile GPUs. It also enables the deployment of the 70B Llama-2 model on mobile GPUs. Together, these techniques significantly reduce the computational and memory costs for deploying deep learning models on …