Back to results

Georgia Institute of Technology

Comparing the Performance of Small Word-Size Floating-Point Numerics to Fixed-Point Numerics in Neural Networks

Abstract

dc:description.abstract

As neural network architectures grow in depth and complexity, training them efficiently under hardware constraints has become increasingly important. While fixed-point arithmetic offers resource advantages, it suffers from limited dynamic range and quantization inflexibility. This thesis introduces an alternative approach—Adaptive Precision Training (APT)—which leverages reduced-precision floating-point formats (FP8, FP12, FP16) for dynamic, layer-wise quantization during training. APT monitors per-layer Quantization Error Measurement (QEM) to guide precision adjustments and incorporates a novel bit-shuffling mechanism to reallocate bits between exponent and mantissa fields before escalating to higher-precision formats. This fine-grained control enables minimal precision escalation while preserving numerical fidelity. The APT framework is implemented in software and evaluated using an AlexNet-style model on the SVHN dataset. The experiments compare three training configurations: a fixed-point baseline, adaptive floating-point quantization, and adaptive quantization with bit-shuffling. Results show that the APT models achieve higher validation accuracy and smoother convergence, while maintaining most training in FP8 and FP12. Although memory usage increases due to dynamic quantization emulation, training time per epoch remains competitive. This work demonstrates that dynamic floating-point quantization—augmented with intra-format bit reallocation—offers a scalable and efficient alternative to fixed-point training, particularly for hardware-aware deep learning applications.

Degree

thesis:*
Level thesis:degree_level
Masters
Department dc:contributor.department
Electrical and Computer Engineering
Grantor dc:publisher
Georgia Institute of Technology
Year dc:date.issued
2025

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Vennapusa, Lakshmi Grishma
Advisor dc:contributor.advisor
  • Anderson, David V.
Committee members dc:contributor.committeemember
  • AlRegib, Ghassan
  • Davenport, Mark
  • Coyle, Edward

Subjects

dc:subject × 12

Rights

Language dc:language.iso
en_US

Identifiers

dc:identifier.*
Handle dc:identifier.uri
https://hdl.handle.net/1853/78656
OAI identifier oai:identifier
oai:repository.gatech.edu:1853/78656

Chain of custody

source
Harvested from
Georgia Tech
Base URL
repository.gatech.edu/server/oai/request
Last updated
2026-07-27
Source record
OAI-PMH GetRecord
citation

Vennapusa, Lakshmi Grishma. Comparing the Performance of Small Word-Size Floating-Point Numerics to Fixed-Point Numerics in Neural Networks. Masters thesis, Georgia Institute of Technology, 2025. https://hdl.handle.net/1853/78656