Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 1 of 1 for “"FP8"”.

  1. Comparing the Performance of Small Word-Size Floating-Point Numerics to Fixed-Point Numerics in Neural Networks

    … reduced-precision floating-point formats (FP8, FP12, FP16) for dynamic, layer-wise quantization during training. APT monitors per-layer Quantization Error Measurement (QEM) to guide precision adjustments and incorporates a novel bit-shuffling mechanism to reallocate bits between exponent …

    gatech Repository record for Comparing the Performance of Small Word-Size Floating-Point Numerics to Fixed-Point Numerics in Neural Networks (opens in a new tab)