University of Illinois at Urbana-Champaign
Profiling and characterization of deep learning model inference on CPU
Abstract
dc:descriptionWith the rapid growth of deep learning models and higher expectations for their accuracy and throughput in real-world applications, the demand for profiling and characterizing model inference on different hardware/software stacks is significantly increased. As the model inference characterization on GPU has already been extensively studied, it is worth exploring how performance-enhancing libraries like Intel MKL-DNN help to boost the performance on Intel CPU. We develop a profiling mechanism to capture the MKL-DNN operation calls and formulate the tracing timeline with spans on the server. Through profiling and characterization that give insights into Intel MKL-DNN, we evaluate and demonstrate that the optimization techniques, including blocked memory layout, layers fusion, and low precision operation used in deep learning model inference, have accelerated the performance on the Intel CPU.
Degree
thesis:*- Name thesis:degree_name
- M.S.
- Level thesis:degree_level
- Thesis
- Discipline thesis:degree_discipline
- Electrical & Computer Engr
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2020
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Qian, Yanli
- Contributors dc:contributor
-
- Hwu, Wen-Mei
Subjects
dc:subject × 4Rights
dc:rights- Statement dc:rights
-
- Copyright 2020 Yanli Qian
- Language dc:language
- en
Identifiers
dc:identifier.*- Handle dc:identifier
- http://hdl.handle.net/2142/108281
- OAI identifier oai:identifier
- oai:www.ideals.illinois.edu:2142/108281