University of Cambridge
Taming TinyML: deep learning inference at computational extremes
Abstract
dc:description.abstractThe advanced data modelling capabilities of neural networks allowed deep learning to become a cornerstone of many applications of artificial intelligence (AI). AI can be brought into our environments by deploying neural models to ubiquitous Internet-of-Things (IoT), wearable and embedded devices to enhance tasks like voice assistance, home security, or health and fitness monitoring. However, such devices are typically powered by microcontroller units (MCUs), which have orders of magnitude fewer computational resources than what is typically required to execute neural networks. This resource scarcity presents a significant obstacle for on-device deep learning, forcing applications to offload the computation to a remote device or server, sacrificing user data privacy and the autonomy of the system. The challenge of bringing neural networks to microcontroller-powered devices, potentially the smallest-scale viable computing platform for deep learning, is tackled by the emerging research field called *TinyML*. In this thesis, I develop model discovery and compression methodology whose common threads are automation and holistic optimisation of network architectures and their execution software, informed by the computational limitations of microcontroller hardware. This unlocks previously-impossible applications of deep learning on MCUs and advances the predictive performance vs resource usage trade-off of neural networks compared to prior work. The novel methodology spans the domains of model compression, neural architecture search (NAS) and software-level execution optimisation. Firstly, I present μNAS, an evolutionary architecture search algorithm that uses a highly granular search space to find performant models with an extremely low (<32 kB) memory footprint. Secondly, I devise a microcontroller-specific network pruning algorithm, which leverages differentiable optimisation, providing drop-in compression within existing network training pipelines with negligible computational overhead. Finally, I develop a model compiler, called Pex, which automatically reduces the memory usage of neural networks by creating partial layer execution schedules. The proposed methodologies optimise accurate resource usage objectives, identified separately by analysing neural network execution bottlenecks on microcontroller hardware. This includes a novel analytical peak memory usage metric, which optimises the order of network layer execution. Overall, the original contributions of the thesis address the challenge of bringing computationally-expensive deep learning modelling to resource-constrained microcontroller hardware, with a focus on practicality and applicability to ubiquitous computing applications.
Degree
thesis:*- Name dc:type.qualificationname
- Doctor of Philosophy (PhD)
- Level dc:type.qualificationlevel
- Doctoral
- Grantor dc:publisher.institution
- University of Cambridge
- Year dc:date.issued
- 2022
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Liberis, Edgaras
- Advisor dc:contributor.advisor
-
- Lane, Nicholas
Subjects
dc:subject × 3Rights
dc:rightsIdentifiers
dc:identifier.*- DOI dc:identifier.doi
- https://doi.org/10.17863/CAM.96947
- OAI identifier oai:identifier
- oai:www.repository.cam.ac.uk:1810/350413