{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/116129"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/116129","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Efficient convolutional neural network inference on microcontrollers","abstract":"Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2024-08-01","abstract_html":"Submission published under a 24 month embargo labeled &#x27;U of I Access&#x27;, the embargo will last until 2024-08-01","abstract_has_math":false,"creators":["Tuttle, Michael"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Shanbhag, Naresh R"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-08","date_published":"2022-08","updated_at":"2026-07-22T22:24:55Z","subjects":["Machine Learning","TinyML","Microcontrollers","Convolutional neural networks","CNN","Pruning"],"languages":["en","eng"],"rights":["Copyright 2022 Michael Tuttle"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/116129","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Shanbhag, Naresh R"]},{"key":"dc:creator","label":"Author","values":["Tuttle, Michael"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022-08","2022-07-21"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Machine Learning","TinyML","Microcontrollers","Convolutional neural networks","CNN","Pruning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2022 Michael Tuttle"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/116129"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2024-08-01","The student, Michael Tuttle, accepted the attached license on 2022-07-21 at 01:46.","The student, Michael Tuttle, submitted this Thesis for approval on 2022-07-21 at 02:08.","This Thesis was approved for publication on 2022-07-21 at 16:57.","DSpace SAF Submission Ingestion Package generated from Vireo submission #18404 on 2022-11-15 at 21:40:41","Convolutional Neural Networks provide state-of-the-art performance on a wide variety of computer vision tasks. However, the large size and computational complexity of these models makes their deployment on resource-constrained edge devices difficult. To remedy this, efficient versions of these layers such as depthwise-separable convolutions and sparse convolutions have been proposed which dramatically reduce the number of parameters and operations required for accurate inference. This work explores various optimizations for these layers on Cortex-M4 MCUs. Memory optimizations such as in-place DWS convolutions and patch-based inference reduce MobileNetV1 peak memory usage by $3.75\\times$. The typically inefficient DW layers are sped up by $2\\times$ over the CMSIS-NN reference kernel, and memory-aware sparsity enables up to $5.7\\times$ speed-up over dense convolutional layers at $95\\%$ sparsity."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Efficient convolutional neural network inference on microcontrollers"]}]}],"canonical_facts":{"dc:contributor":["Shanbhag, Naresh R"],"dc:creator":["Tuttle, Michael"],"dc:date":["2022-08","2022-07-21"],"dc:description":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2024-08-01","The student, Michael Tuttle, accepted the attached license on 2022-07-21 at 01:46.","The student, Michael Tuttle, submitted this Thesis for approval on 2022-07-21 at 02:08.","This Thesis was approved for publication on 2022-07-21 at 16:57.","DSpace SAF Submission Ingestion Package generated from Vireo submission #18404 on 2022-11-15 at 21:40:41","Convolutional Neural Networks provide state-of-the-art performance on a wide variety of computer vision tasks. However, the large size and computational complexity of these models makes their deployment on resource-constrained edge devices difficult. To remedy this, efficient versions of these layers such as depthwise-separable convolutions and sparse convolutions have been proposed which dramatically reduce the number of parameters and operations required for accurate inference. This work explores various optimizations for these layers on Cortex-M4 MCUs. Memory optimizations such as in-place DWS convolutions and patch-based inference reduce MobileNetV1 peak memory usage by $3.75\\times$. The typically inefficient DW layers are sped up by $2\\times$ over the CMSIS-NN reference kernel, and memory-aware sparsity enables up to $5.7\\times$ speed-up over dense convolutional layers at $95\\%$ sparsity."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/116129"],"dc:language":["en","eng"],"dc:rights":["Copyright 2022 Michael Tuttle"],"dc:subject":["Machine Learning","TinyML","Microcontrollers","Convolutional neural networks","CNN","Pruning"],"dc:title":["Efficient convolutional neural network inference on microcontrollers"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:55Z"}