{"id":{"repo_id":"umn","oai_identifier":"oai:conservancy.umn.edu:11299/275925"},"canonical_url":"https://search.dev.ndltd.org/etd/umn/oai:conservancy.umn.edu:11299/275925","repository":{"repo_id":"umn","name":"University of Minnesota","base_url":"https://conservancy.umn.edu/server/oai/request"},"display":{"title":"Hardware software co-design of machine learning accelerators using univariate functions","abstract":"As neural networks grow in complexity, efficient hardware-software co-design is crucial for balancing performance and resource constraints, especially for edge devices and FPGA accelerators. This thesis investigates optimization techniques for hardware-awaretraining and model compression using univariate functions. First, we optimize hardware using a simple constant coefficient multiplier on Hybrid Binary-Unary (HBUNN) architecture, which offers variable hardware costs for constant coefficients. By applying regularization-based training, we reduce the full model hard- ware area cost by 59.3% while improving accuracy by 0.42% over L2 regularization. Using this optimized model, the fully parallel-pipeline HBUNN-based ResNet-18 has a reduced area cost by 29.6% compared to conventional binary architectures. Next, we introduce Personal Self-Attention (PSA), a novel method for learning non- linear univariate functions. Using PSA with linear transformations, we demonstrate a 2×compression in hidden size of Multi-Layer Perceptrons (MLP), while matching accuracy. Applying this to an MLP-based vision model on CIFAR-10, we cut the number of operations by 45%–28%, boosting hardware efficiency. We validate PSA with two hardware accelerators while maintaining the same reference accuracy. First, an unrolled streaming design that reduces LUT + DSP usage by 25% while doubling throughput to 32kFPS. Second, a fixed-size SIMD accelerator that improves throughput by 62.1% while using only 3.5% additional LUTs.","abstract_html":"As neural networks grow in complexity, efficient hardware-software co-design is crucial for balancing performance and resource constraints, especially for edge devices and FPGA accelerators. This thesis investigates optimization techniques for hardware-awaretraining and model compression using univariate functions. First, we optimize hardware using a simple constant coefficient multiplier on Hybrid Binary-Unary (HBUNN) architecture, which offers variable hardware costs for constant coefficients. By applying regularization-based training, we reduce the full model hard- ware area cost by 59.3% while improving accuracy by 0.42% over L2 regularization. Using this optimized model, the fully parallel-pipeline HBUNN-based ResNet-18 has a reduced area cost by 29.6% compared to conventional binary architectures. Next, we introduce Personal Self-Attention (PSA), a novel method for learning non- linear univariate functions. Using PSA with linear transformations, we demonstrate a 2×compression in hidden size of Multi-Layer Perceptrons (MLP), while matching accuracy. Applying this to an MLP-based vision model on CIFAR-10, we cut the number of operations by 45%–28%, boosting hardware efficiency. We validate PSA with two hardware accelerators while maintaining the same reference accuracy. First, an unrolled streaming design that reduces LUT + DSP usage by 25% while doubling throughput to 32kFPS. Second, a fixed-size SIMD accelerator that improves throughput by 62.1% while using only 3.5% additional LUTs.","abstract_has_math":false,"creators":["Singh, Gaurav"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-05","date_published":"2025-05","updated_at":"2026-07-24T05:19:44Z","subjects":["Accelerator","Co-Design","Compression","Deep Learning","FPGA","Learnable Activation"],"languages":["en"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/11299/275925","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["Singh, Gaurav"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2025-08-25T12:14:23Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2025-08-25T12:14:23Z"]},{"key":"dc:date.issued","label":"Date","values":["2025-05"]},{"key":"dc:type","label":"Dc Type","values":["Thesis or Dissertation"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Accelerator","Co-Design","Compression","Deep Learning","FPGA","Learnable Activation"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/11299/275925"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["University of Minnesota Ph.D. dissertation. May 2025. Major: Electrical Engineering. Advisor: Kia Bazargan. 1 computer file (PDF); xiii, 109 pages."]},{"key":"dc:description.abstract","label":"Abstract","values":["As neural networks grow in complexity, efficient hardware-software co-design is crucial for balancing performance and resource constraints, especially for edge devices and FPGA accelerators. This thesis investigates optimization techniques for hardware-awaretraining and model compression using univariate functions. First, we optimize hardware using a simple constant coefficient multiplier on Hybrid Binary-Unary (HBUNN) architecture, which offers variable hardware costs for constant coefficients. By applying regularization-based training, we reduce the full model hard- ware area cost by 59.3% while improving accuracy by 0.42% over L2 regularization. Using this optimized model, the fully parallel-pipeline HBUNN-based ResNet-18 has a reduced area cost by 29.6% compared to conventional binary architectures. Next, we introduce Personal Self-Attention (PSA), a novel method for learning non- linear univariate functions. Using PSA with linear transformations, we demonstrate a 2×compression in hidden size of Multi-Layer Perceptrons (MLP), while matching accuracy. Applying this to an MLP-based vision model on CIFAR-10, we cut the number of operations by 45%–28%, boosting hardware efficiency. We validate PSA with two hardware accelerators while maintaining the same reference accuracy. First, an unrolled streaming design that reduces LUT + DSP usage by 25% while doubling throughput to 32kFPS. Second, a fixed-size SIMD accelerator that improves throughput by 62.1% while using only 3.5% additional LUTs."]},{"key":"dc:title","label":"Title","values":["Hardware software co-design of machine learning accelerators using univariate functions"]}]}],"canonical_facts":{"dc:creator":["Singh, Gaurav"],"dc:date.accessioned":["2025-08-25T12:14:23Z"],"dc:date.available":["2025-08-25T12:14:23Z"],"dc:date.issued":["2025-05"],"dc:description":["University of Minnesota Ph.D. dissertation. May 2025. Major: Electrical Engineering. Advisor: Kia Bazargan. 1 computer file (PDF); xiii, 109 pages."],"dc:description.abstract":["As neural networks grow in complexity, efficient hardware-software co-design is crucial for balancing performance and resource constraints, especially for edge devices and FPGA accelerators. This thesis investigates optimization techniques for hardware-awaretraining and model compression using univariate functions. First, we optimize hardware using a simple constant coefficient multiplier on Hybrid Binary-Unary (HBUNN) architecture, which offers variable hardware costs for constant coefficients. By applying regularization-based training, we reduce the full model hard- ware area cost by 59.3% while improving accuracy by 0.42% over L2 regularization. Using this optimized model, the fully parallel-pipeline HBUNN-based ResNet-18 has a reduced area cost by 29.6% compared to conventional binary architectures. Next, we introduce Personal Self-Attention (PSA), a novel method for learning non- linear univariate functions. Using PSA with linear transformations, we demonstrate a 2×compression in hidden size of Multi-Layer Perceptrons (MLP), while matching accuracy. Applying this to an MLP-based vision model on CIFAR-10, we cut the number of operations by 45%–28%, boosting hardware efficiency. We validate PSA with two hardware accelerators while maintaining the same reference accuracy. First, an unrolled streaming design that reduces LUT + DSP usage by 25% while doubling throughput to 32kFPS. Second, a fixed-size SIMD accelerator that improves throughput by 62.1% while using only 3.5% additional LUTs."],"dc:identifier.uri":["https://hdl.handle.net/11299/275925"],"dc:language.iso":["en"],"dc:subject":["Accelerator","Co-Design","Compression","Deep Learning","FPGA","Learnable Activation"],"dc:title":["Hardware software co-design of machine learning accelerators using univariate functions"],"dc:type":["Thesis or Dissertation"]},"updated_at":"2026-07-24T05:19:44Z"}