University of Illinois at Urbana-Champaign
Efficient convolutional neural network inference on microcontrollers
Abstract
dc:descriptionConvolutional Neural Networks provide state-of-the-art performance on a wide variety of computer vision tasks. However, the large size and computational complexity of these models makes their deployment on resource-constrained edge devices difficult. To remedy this, efficient versions of these layers such as depthwise-separable convolutions and sparse convolutions have been proposed which dramatically reduce the number of parameters and operations required for accurate inference. This work explores various optimizations for these layers on Cortex-M4 MCUs. Memory optimizations such as in-place DWS convolutions and patch-based inference reduce MobileNetV1 peak memory usage by $3.75\times$. The typically inefficient DW layers are sped up by $2\times$ over the CMSIS-NN reference kernel, and memory-aware sparsity enables up to $5.7\times$ speed-up over dense convolutional layers at $95\%$ sparsity.
Degree
thesis:*- Name thesis:degree_name
- M.S.
- Level thesis:degree_level
- Thesis
- Discipline thesis:degree_discipline
- Electrical & Computer Engr
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2022
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Tuttle, Michael
- Contributors dc:contributor
-
- Shanbhag, Naresh R
Subjects
dc:subject × 6Rights
dc:rights- Statement dc:rights
-
- Copyright 2022 Michael Tuttle
- Language dc:language
- en, eng
Identifiers
dc:identifier.*- Handle dc:identifier
- https://hdl.handle.net/2142/116129