{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/105251"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/105251","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Efficient inference of convolutional neural networks on general purpose hardware using weight repetition","abstract":"Deep Neural Networks (DNNs) have begun to permeate all corners of electronic society due to their high accuracy and machine efficiency per operation. Recent work has shown how weights within and across DNN filters have large degrees of repetition due to the pigeonhole principle and modern weight quantization schemes, and that this weight repetition can be harnessed improve DNN inference efficiency in an accelerator/ASIC context. This thesis develops new techniques so that weight repetition leads to an efficiency gain on general-purpose and programmable SIMD-based architectures such as CPUs equipped with vector extensions. We show how to write high-performance software that does not require hardware modifications and can cope with the irregularity introduced by weight repetition schemes. Overall, our highly parallel software kernel achieves up to 1:51 speedup in runtime of inference over state-of-the-art baseline.","abstract_html":"Deep Neural Networks (DNNs) have begun to permeate all corners of electronic society due to their high accuracy and machine efficiency per operation. Recent work has shown how weights within and across DNN filters have large degrees of repetition due to the pigeonhole principle and modern weight quantization schemes, and that this weight repetition can be harnessed improve DNN inference efficiency in an accelerator/ASIC context. This thesis develops new techniques so that weight repetition leads to an efficiency gain on general-purpose and programmable SIMD-based architectures such as CPUs equipped with vector extensions. We show how to write high-performance software that does not require hardware modifications and can cope with the irregularity introduced by weight repetition schemes. Overall, our highly parallel software kernel achieves up to 1:51 speedup in runtime of inference over state-of-the-art baseline.","abstract_has_math":false,"creators":["Agrawal, Rohit"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Fletcher, Christopher W."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019-08-23T20:48:24Z","date_published":"2019-08-23T20:48:24Z","updated_at":"2026-07-22T22:24:44Z","subjects":["Deep Neural Networks","Convolutional Neural Networks","Accelerator","CPU","GPU","Deep Learning Hardware","CNN Inference"],"languages":["en"],"rights":["Copyright 2019 Rohit Agrawal"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/105251","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Fletcher, Christopher W."]},{"key":"dc:creator","label":"Author","values":["Agrawal, Rohit"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2019-08-23T20:48:24Z","2021-08-24T09:15:24Z","2019-04-24","2019-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Deep Neural Networks","Convolutional Neural Networks","Accelerator","CPU","GPU","Deep Learning Hardware","CNN Inference"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2019 Rohit Agrawal"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/105251"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Deep Neural Networks (DNNs) have begun to permeate all corners of electronic society due to their high accuracy and machine efficiency per operation. Recent work has shown how weights within and across DNN filters have large degrees of repetition due to the pigeonhole principle and modern weight quantization schemes, and that this weight repetition can be harnessed improve DNN inference efficiency in an accelerator/ASIC context. This thesis develops new techniques so that weight repetition leads to an efficiency gain on general-purpose and programmable SIMD-based architectures such as CPUs equipped with vector extensions. We show how to write high-performance software that does not require hardware modifications and can cope with the irregularity introduced by weight repetition schemes. Overall, our highly parallel software kernel achieves up to 1:51 speedup in runtime of inference over state-of-the-art baseline.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2021-05-01","The student, Rohit Agrawal, accepted the attached license on 2019-04-23 at 21:40.","The student, Rohit Agrawal, submitted this Thesis for approval on 2019-04-23 at 21:43.","This Thesis was approved for publication on 2019-04-24 at 15:53.","DSpace SAF Submission Ingestion Package generated from Vireo submission #13851 on 2019-08-22 at 16:23:40","Made available in DSpace on 2019-08-23T20:48:24Z (GMT). No. of bitstreams: 2 AGRAWAL-THESIS-2019.pdf: 1182971 bytes, checksum: 87ecc270d1f189389584f14dd0439fb5 (MD5) LICENSE.txt: 4210 bytes, checksum: b4c74e22275709262713de9137b1212d (MD5) Previous issue date: 2019-04-24","Embargo set by: Seth Robbins for item 112373 Lift date: 2021-08-23T20:48:32Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 112373 on 2021-08-24T09:15:24Z."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Efficient inference of convolutional neural networks on general purpose hardware using weight repetition"]}]}],"canonical_facts":{"dc:contributor":["Fletcher, Christopher W."],"dc:creator":["Agrawal, Rohit"],"dc:date":["2019-08-23T20:48:24Z","2021-08-24T09:15:24Z","2019-04-24","2019-05"],"dc:description":["Deep Neural Networks (DNNs) have begun to permeate all corners of electronic society due to their high accuracy and machine efficiency per operation. Recent work has shown how weights within and across DNN filters have large degrees of repetition due to the pigeonhole principle and modern weight quantization schemes, and that this weight repetition can be harnessed improve DNN inference efficiency in an accelerator/ASIC context. This thesis develops new techniques so that weight repetition leads to an efficiency gain on general-purpose and programmable SIMD-based architectures such as CPUs equipped with vector extensions. We show how to write high-performance software that does not require hardware modifications and can cope with the irregularity introduced by weight repetition schemes. Overall, our highly parallel software kernel achieves up to 1:51 speedup in runtime of inference over state-of-the-art baseline.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2021-05-01","The student, Rohit Agrawal, accepted the attached license on 2019-04-23 at 21:40.","The student, Rohit Agrawal, submitted this Thesis for approval on 2019-04-23 at 21:43.","This Thesis was approved for publication on 2019-04-24 at 15:53.","DSpace SAF Submission Ingestion Package generated from Vireo submission #13851 on 2019-08-22 at 16:23:40","Made available in DSpace on 2019-08-23T20:48:24Z (GMT). No. of bitstreams: 2 AGRAWAL-THESIS-2019.pdf: 1182971 bytes, checksum: 87ecc270d1f189389584f14dd0439fb5 (MD5) LICENSE.txt: 4210 bytes, checksum: b4c74e22275709262713de9137b1212d (MD5) Previous issue date: 2019-04-24","Embargo set by: Seth Robbins for item 112373 Lift date: 2021-08-23T20:48:32Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 112373 on 2021-08-24T09:15:24Z."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/105251"],"dc:language":["en"],"dc:rights":["Copyright 2019 Rohit Agrawal"],"dc:subject":["Deep Neural Networks","Convolutional Neural Networks","Accelerator","CPU","GPU","Deep Learning Hardware","CNN Inference"],"dc:title":["Efficient inference of convolutional neural networks on general purpose hardware using weight repetition"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:44Z"}