University of Cambridge
Near-Memory Processing for Low-precision Deep Neural Networks
Abstract
dc:description.abstractDeep neural networks (DNNs) provide many application domains with state-of-the-art performance and accuracy. However, they are compute-heavy and data-intensive which makes deploying them on resource-constrained edge devices challenging. Transforming the real-valued parameters of a DNN into a minimised-bit-width approximated version through quantising or binarising lowers their accuracy to an extent but significantly reduces their computation complexity and storage space. Moreover, it makes them good candidates to be considered for near-memory processing due to these criteria. On the other hand, when it comes to designing a hardware module (either for near-memory processing or not), the more specialised the hardware the better the performance. The downside with over-specialised modules then becomes their inability to adapt, which can be problematic in a fast-evolving field such as deep learning. To this end, designing a module with reasonable performance, area overhead and energy consumption while maintaining a good balance between the physical limitations of the design and its flexibility is a challenge. The contribution of this thesis is to introduce a processing-near-memory module (PNM) for low-precision convolution neural networks. The placement of this module is near the main memory (DDR4 DRAM) and the memory controller, with a data layout that does not require rearrangement either before or after the convolution operations. Here, two distinct design modes for the PNM module are presented: mode S, which focuses on minimizing area overhead, and mode T, aimed at optimizing data transfers between the module and DRAM. The performance, area, and energy costs of these designs were thoroughly assessed through both analytical and practical analysis. These evaluations highlighted the impact of varying filters, hardware replicas, and bit-widths for each model on the stated criteria, leading to recommended design choices tailored to specific use cases' demands and constraints. An evaluation using a BCNN based on AlexNet indicates that for mode S, a configuration of one hardware replica with 16 filters per replica offers an optimal balance between area, runtime, and energy. In contrast, for mode T, the best configuration comprises 16 replicas and 32 filters. Comparatively, mode T surpasses mode S by a factor of 6.39 in performance, while mode S achieves greater savings in area and energy, by factors of 3.56 and 15.99, respectively.
Degree
thesis:*- Name dc:type.qualificationname
- Doctor of Philosophy (PhD)
- Level dc:type.qualificationlevel
- Doctoral
- Grantor dc:publisher.institution
- University of Cambridge
- Year dc:date.issued
- 2024
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Miralaei, Aida
- Advisor dc:contributor.advisor
-
- Jones, Timothy M
Subjects
dc:subject × 6Rights
dc:rightsIdentifiers
dc:identifier.*- DOI dc:identifier.doi
- https://doi.org/10.17863/CAM.108441
- OAI identifier oai:identifier
- oai:www.repository.cam.ac.uk:1810/368112