{"id":{"repo_id":"gatech","oai_identifier":"oai:repository.gatech.edu:1853/78656"},"canonical_url":"https://search.dev.ndltd.org/etd/gatech/oai:repository.gatech.edu:1853/78656","repository":{"repo_id":"gatech","name":"Georgia Tech","base_url":"https://repository.gatech.edu/server/oai/request"},"display":{"title":"Comparing the Performance of Small Word-Size Floating-Point Numerics to Fixed-Point Numerics in Neural Networks","abstract":"As neural network architectures grow in depth and complexity, training them efficiently under hardware constraints has become increasingly important. While fixed-point arithmetic offers resource advantages, it suffers from limited dynamic range and quantization inflexibility. This thesis introduces an alternative approach—Adaptive Precision Training (APT)—which leverages reduced-precision floating-point formats (FP8, FP12, FP16) for dynamic, layer-wise quantization during training. APT monitors per-layer Quantization Error Measurement (QEM) to guide precision adjustments and incorporates a novel bit-shuffling mechanism to reallocate bits between exponent and mantissa fields before escalating to higher-precision formats. This fine-grained control enables minimal precision escalation while preserving numerical fidelity. The APT framework is implemented in software and evaluated using an AlexNet-style model on the SVHN dataset. The experiments compare three training configurations: a fixed-point baseline, adaptive floating-point quantization, and adaptive quantization with bit-shuffling. Results show that the APT models achieve higher validation accuracy and smoother convergence, while maintaining most training in FP8 and FP12. Although memory usage increases due to dynamic quantization emulation, training time per epoch remains competitive. This work demonstrates that dynamic floating-point quantization—augmented with intra-format bit reallocation—offers a scalable and efficient alternative to fixed-point training, particularly for hardware-aware deep learning applications.","abstract_html":"As neural network architectures grow in depth and complexity, training them efficiently under hardware constraints has become increasingly important. While fixed-point arithmetic offers resource advantages, it suffers from limited dynamic range and quantization inflexibility. This thesis introduces an alternative approach—Adaptive Precision Training (APT)—which leverages reduced-precision floating-point formats (FP8, FP12, FP16) for dynamic, layer-wise quantization during training. APT monitors per-layer Quantization Error Measurement (QEM) to guide precision adjustments and incorporates a novel bit-shuffling mechanism to reallocate bits between exponent and mantissa fields before escalating to higher-precision formats. This fine-grained control enables minimal precision escalation while preserving numerical fidelity. The APT framework is implemented in software and evaluated using an AlexNet-style model on the SVHN dataset. The experiments compare three training configurations: a fixed-point baseline, adaptive floating-point quantization, and adaptive quantization with bit-shuffling. Results show that the APT models achieve higher validation accuracy and smoother convergence, while maintaining most training in FP8 and FP12. Although memory usage increases due to dynamic quantization emulation, training time per epoch remains competitive. This work demonstrates that dynamic floating-point quantization—augmented with intra-format bit reallocation—offers a scalable and efficient alternative to fixed-point training, particularly for hardware-aware deep learning applications.","abstract_has_math":false,"creators":["Vennapusa, Lakshmi Grishma"],"institution":"Georgia Institute of Technology","degree_name":null,"degree_level":"Masters","degree_discipline":null,"degree_department":"Electrical and Computer Engineering","school":null,"contributors":[],"advisors":["Anderson, David V."],"committee_chairs":[],"committee_members":["AlRegib, Ghassan","Davenport, Mark","Coyle, Edward"],"year":2025,"date_issued":"2025-05-28","date_published":"2025-05-28","updated_at":"2026-07-27T19:49:22Z","subjects":["Adaptive Precision Training","floating-point quantization","FP8","FP12","FP16","bit-shuffling","quantization error measurement (QEM)","neural network training","low-precision arithmetic","fixed-point vs floating-point","quantization-aware training","hardware-efficient deep learning"],"languages":["en_US"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/1853/78656","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Anderson, David V."]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["AlRegib, Ghassan","Davenport, Mark","Coyle, Edward"]},{"key":"dc:contributor.department","label":"Department","values":["Electrical and Computer Engineering"]},{"key":"dc:creator","label":"Author","values":["Vennapusa, Lakshmi Grishma"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2025-08-15T12:40:24Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2025-08-15T12:40:24Z"]},{"key":"dc:date.issued","label":"Date","values":["2025-05-28"]},{"key":"dc:publisher","label":"Institution","values":["Georgia Institute of Technology"]},{"key":"dc:type","label":"Dc Type","values":["Text"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Masters"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Adaptive Precision Training","floating-point quantization","FP8","FP12","FP16","bit-shuffling","quantization error measurement (QEM)","neural network training","low-precision arithmetic","fixed-point vs floating-point","quantization-aware training","hardware-efficient deep learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en_US"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/1853/78656"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["As neural network architectures grow in depth and complexity, training them efficiently under hardware constraints has become increasingly important. While fixed-point arithmetic offers resource advantages, it suffers from limited dynamic range and quantization inflexibility. This thesis introduces an alternative approach—Adaptive Precision Training (APT)—which leverages reduced-precision floating-point formats (FP8, FP12, FP16) for dynamic, layer-wise quantization during training. APT monitors per-layer Quantization Error Measurement (QEM) to guide precision adjustments and incorporates a novel bit-shuffling mechanism to reallocate bits between exponent and mantissa fields before escalating to higher-precision formats. This fine-grained control enables minimal precision escalation while preserving numerical fidelity. The APT framework is implemented in software and evaluated using an AlexNet-style model on the SVHN dataset. The experiments compare three training configurations: a fixed-point baseline, adaptive floating-point quantization, and adaptive quantization with bit-shuffling. Results show that the APT models achieve higher validation accuracy and smoother convergence, while maintaining most training in FP8 and FP12. Although memory usage increases due to dynamic quantization emulation, training time per epoch remains competitive. This work demonstrates that dynamic floating-point quantization—augmented with intra-format bit reallocation—offers a scalable and efficient alternative to fixed-point training, particularly for hardware-aware deep learning applications."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["M.S."]},{"key":"dc:format.mimetype","label":"Dc Format Mimetype","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Comparing the Performance of Small Word-Size Floating-Point Numerics to Fixed-Point Numerics in Neural Networks"]}]}],"canonical_facts":{"dc:contributor.advisor":["Anderson, David V."],"dc:contributor.committeemember":["AlRegib, Ghassan","Davenport, Mark","Coyle, Edward"],"dc:contributor.department":["Electrical and Computer Engineering"],"dc:creator":["Vennapusa, Lakshmi Grishma"],"dc:date.accessioned":["2025-08-15T12:40:24Z"],"dc:date.available":["2025-08-15T12:40:24Z"],"dc:date.issued":["2025-05-28"],"dc:description.abstract":["As neural network architectures grow in depth and complexity, training them efficiently under hardware constraints has become increasingly important. While fixed-point arithmetic offers resource advantages, it suffers from limited dynamic range and quantization inflexibility. This thesis introduces an alternative approach—Adaptive Precision Training (APT)—which leverages reduced-precision floating-point formats (FP8, FP12, FP16) for dynamic, layer-wise quantization during training. APT monitors per-layer Quantization Error Measurement (QEM) to guide precision adjustments and incorporates a novel bit-shuffling mechanism to reallocate bits between exponent and mantissa fields before escalating to higher-precision formats. This fine-grained control enables minimal precision escalation while preserving numerical fidelity. The APT framework is implemented in software and evaluated using an AlexNet-style model on the SVHN dataset. The experiments compare three training configurations: a fixed-point baseline, adaptive floating-point quantization, and adaptive quantization with bit-shuffling. Results show that the APT models achieve higher validation accuracy and smoother convergence, while maintaining most training in FP8 and FP12. Although memory usage increases due to dynamic quantization emulation, training time per epoch remains competitive. This work demonstrates that dynamic floating-point quantization—augmented with intra-format bit reallocation—offers a scalable and efficient alternative to fixed-point training, particularly for hardware-aware deep learning applications."],"dc:description.degree":["M.S."],"dc:format.mimetype":["application/pdf"],"dc:identifier.uri":["https://hdl.handle.net/1853/78656"],"dc:language.iso":["en_US"],"dc:publisher":["Georgia Institute of Technology"],"dc:subject":["Adaptive Precision Training","floating-point quantization","FP8","FP12","FP16","bit-shuffling","quantization error measurement (QEM)","neural network training","low-precision arithmetic","fixed-point vs floating-point","quantization-aware training","hardware-efficient deep learning"],"dc:title":["Comparing the Performance of Small Word-Size Floating-Point Numerics to Fixed-Point Numerics in Neural Networks"],"dc:type":["Text"],"thesis:degree_level":["Masters"]},"updated_at":"2026-07-27T19:49:22Z"}