{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/108281"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/108281","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Profiling and characterization of deep learning model inference on CPU","abstract":"With the rapid growth of deep learning models and higher expectations for their accuracy and throughput in real-world applications, the demand for proﬁling and characterizing model inference on diﬀerent hardware/software stacks is signiﬁcantly increased. As the model inference characterization on GPU has already been extensively studied, it is worth exploring how performance-enhancing libraries like Intel MKL-DNN help to boost the performance on Intel CPU. We develop a proﬁling mechanism to capture the MKL-DNN operation calls and formulate the tracing timeline with spans on the server. Through proﬁling and characterization that give insights into Intel MKL-DNN, we evaluate and demonstrate that the optimization techniques, including blocked memory layout, layers fusion, and low precision operation used in deep learning model inference, have accelerated the performance on the Intel CPU.","abstract_html":"With the rapid growth of deep learning models and higher expectations for their accuracy and throughput in real-world applications, the demand for proﬁling and characterizing model inference on diﬀerent hardware/software stacks is signiﬁcantly increased. As the model inference characterization on GPU has already been extensively studied, it is worth exploring how performance-enhancing libraries like Intel MKL-DNN help to boost the performance on Intel CPU. We develop a proﬁling mechanism to capture the MKL-DNN operation calls and formulate the tracing timeline with spans on the server. Through proﬁling and characterization that give insights into Intel MKL-DNN, we evaluate and demonstrate that the optimization techniques, including blocked memory layout, layers fusion, and low precision operation used in deep learning model inference, have accelerated the performance on the Intel CPU.","abstract_has_math":false,"creators":["Qian, Yanli"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Hwu, Wen-Mei"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2020,"date_issued":"2020-08-27T00:50:04Z","date_published":"2020-08-27T00:50:04Z","updated_at":"2026-07-22T22:24:48Z","subjects":["Deep learning","Machine learning","Profiling","Performance-library"],"languages":["en"],"rights":["Copyright 2020 Yanli Qian"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/108281","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Hwu, Wen-Mei"]},{"key":"dc:creator","label":"Author","values":["Qian, Yanli"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2020-08-27T00:50:04Z","2022-08-27T00:51:40Z","2020-04-28","2020-05"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Deep learning","Machine learning","Profiling","Performance-library"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2020 Yanli Qian"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/108281"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["With the rapid growth of deep learning models and higher expectations for their accuracy and throughput in real-world applications, the demand for proﬁling and characterizing model inference on diﬀerent hardware/software stacks is signiﬁcantly increased. As the model inference characterization on GPU has already been extensively studied, it is worth exploring how performance-enhancing libraries like Intel MKL-DNN help to boost the performance on Intel CPU. We develop a proﬁling mechanism to capture the MKL-DNN operation calls and formulate the tracing timeline with spans on the server. Through proﬁling and characterization that give insights into Intel MKL-DNN, we evaluate and demonstrate that the optimization techniques, including blocked memory layout, layers fusion, and low precision operation used in deep learning model inference, have accelerated the performance on the Intel CPU.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2022-05-01","The student, Yanli Qian, accepted the attached license on 2020-04-28 at 07:31.","The student, Yanli Qian, submitted this Thesis for approval on 2020-04-28 at 07:33.","This Thesis was approved for publication on 2020-04-28 at 10:42.","DSpace SAF Submission Ingestion Package generated from Vireo submission #15088 on 2020-08-25 at 17:41:19","Made available in DSpace on 2020-08-27T00:50:04Z (GMT). No. of bitstreams: 2 QIAN-THESIS-2020.pdf: 3770926 bytes, checksum: 3d88f9152e41b4e964fe192b3915ab25 (MD5) LICENSE.txt: 4207 bytes, checksum: 4ca1a48a73edc66d21a9078eeaef766b (MD5) Previous issue date: 2020-04-28","Embargo set by: Seth Robbins for item 115895 Lift date: 2022-08-27T00:50:22Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 115895 Lift date: 2022-08-27T00:51:40Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Profiling and characterization of deep learning model inference on CPU"]}]}],"canonical_facts":{"dc:contributor":["Hwu, Wen-Mei"],"dc:creator":["Qian, Yanli"],"dc:date":["2020-08-27T00:50:04Z","2022-08-27T00:51:40Z","2020-04-28","2020-05"],"dc:description":["With the rapid growth of deep learning models and higher expectations for their accuracy and throughput in real-world applications, the demand for proﬁling and characterizing model inference on diﬀerent hardware/software stacks is signiﬁcantly increased. As the model inference characterization on GPU has already been extensively studied, it is worth exploring how performance-enhancing libraries like Intel MKL-DNN help to boost the performance on Intel CPU. We develop a proﬁling mechanism to capture the MKL-DNN operation calls and formulate the tracing timeline with spans on the server. Through proﬁling and characterization that give insights into Intel MKL-DNN, we evaluate and demonstrate that the optimization techniques, including blocked memory layout, layers fusion, and low precision operation used in deep learning model inference, have accelerated the performance on the Intel CPU.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2022-05-01","The student, Yanli Qian, accepted the attached license on 2020-04-28 at 07:31.","The student, Yanli Qian, submitted this Thesis for approval on 2020-04-28 at 07:33.","This Thesis was approved for publication on 2020-04-28 at 10:42.","DSpace SAF Submission Ingestion Package generated from Vireo submission #15088 on 2020-08-25 at 17:41:19","Made available in DSpace on 2020-08-27T00:50:04Z (GMT). No. of bitstreams: 2 QIAN-THESIS-2020.pdf: 3770926 bytes, checksum: 3d88f9152e41b4e964fe192b3915ab25 (MD5) LICENSE.txt: 4207 bytes, checksum: 4ca1a48a73edc66d21a9078eeaef766b (MD5) Previous issue date: 2020-04-28","Embargo set by: Seth Robbins for item 115895 Lift date: 2022-08-27T00:50:22Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 115895 Lift date: 2022-08-27T00:51:40Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/108281"],"dc:language":["en"],"dc:rights":["Copyright 2020 Yanli Qian"],"dc:subject":["Deep learning","Machine learning","Profiling","Performance-library"],"dc:title":["Profiling and characterization of deep learning model inference on CPU"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:48Z"}