{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/108349"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/108349","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Utilizing GPU tensor cores for algorithmic acceleration","abstract":"There has been a surge in the demand for a Domain Specific Architecture due to wide ranging deep learning applications like Image classification, speech recognition, in healthcare, self-driving cars etc. Matrix Multiplication acceleration has been a popular design choice when creating these specialized units to boost deep learning training and inference. Nvidia's Volta architecture introduced Tensor Cores which promised a 3 times speedup over their Pascal architecture. Despite the favorable performance gains, these accelerators have not been applied extensively to a wider class of algorithms. Through this thesis we introduce novel ways of mapping various algorithms on the Tensor Cores. We implemented Tensor Core based reduction, power iteration and Fast Fourier Transform (FFT) and show that effectively utilizing GPU compute resources would result in substantial gains in performance. Our reduction gave a 1.5 times speedup against CUB API; power iteration gave on average 2 times the speedup against Thrust and cuBLAS based implementation while our FFT implementation was able to outperform cuFFT with up to 8 times the speedup.","abstract_html":"There has been a surge in the demand for a Domain Specific Architecture due to wide ranging deep learning applications like Image classification, speech recognition, in healthcare, self-driving cars etc. Matrix Multiplication acceleration has been a popular design choice when creating these specialized units to boost deep learning training and inference. Nvidia&#x27;s Volta architecture introduced Tensor Cores which promised a 3 times speedup over their Pascal architecture. Despite the favorable performance gains, these accelerators have not been applied extensively to a wider class of algorithms. Through this thesis we introduce novel ways of mapping various algorithms on the Tensor Cores. We implemented Tensor Core based reduction, power iteration and Fast Fourier Transform (FFT) and show that effectively utilizing GPU compute resources would result in substantial gains in performance. Our reduction gave a 1.5 times speedup against CUB API; power iteration gave on average 2 times the speedup against Thrust and cuBLAS based implementation while our FFT implementation was able to outperform cuFFT with up to 8 times the speedup.","abstract_has_math":false,"creators":["Durrani, Sultan Hayat Khan"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Hwu, Wen-Mei W"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2020,"date_issued":"2020-08-27T00:51:33Z","date_published":"2020-08-27T00:51:33Z","updated_at":"2026-07-22T22:24:48Z","subjects":["Computer Architecture","GPU","Tensor Cores","CUDA","FFT"],"languages":["en"],"rights":["Copyright 2020 Sultan Hayat Khan Durrani"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/108349","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Hwu, Wen-Mei W"]},{"key":"dc:creator","label":"Author","values":["Durrani, Sultan Hayat Khan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2020-08-27T00:51:33Z","2022-08-27T00:51:40Z","2020-05-13","2020-05"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer Architecture","GPU","Tensor Cores","CUDA","FFT"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2020 Sultan Hayat Khan Durrani"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/108349"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["There has been a surge in the demand for a Domain Specific Architecture due to wide ranging deep learning applications like Image classification, speech recognition, in healthcare, self-driving cars etc. Matrix Multiplication acceleration has been a popular design choice when creating these specialized units to boost deep learning training and inference. Nvidia's Volta architecture introduced Tensor Cores which promised a 3 times speedup over their Pascal architecture. Despite the favorable performance gains, these accelerators have not been applied extensively to a wider class of algorithms. Through this thesis we introduce novel ways of mapping various algorithms on the Tensor Cores. We implemented Tensor Core based reduction, power iteration and Fast Fourier Transform (FFT) and show that effectively utilizing GPU compute resources would result in substantial gains in performance. Our reduction gave a 1.5 times speedup against CUB API; power iteration gave on average 2 times the speedup against Thrust and cuBLAS based implementation while our FFT implementation was able to outperform cuFFT with up to 8 times the speedup.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2022-05-01","The student, Sultan Durrani, accepted the attached license on 2020-05-12 at 13:44.","The student, Sultan Durrani, submitted this Thesis for approval on 2020-05-12 at 14:03.","This Thesis was approved for publication on 2020-05-13 at 11:11.","DSpace SAF Submission Ingestion Package generated from Vireo submission #15357 on 2020-08-25 at 17:44:25","Made available in DSpace on 2020-08-27T00:51:33Z (GMT). No. of bitstreams: 2 DURRANI-THESIS-2020.pdf: 1000595 bytes, checksum: 5b6eb05e70aa304a693e486c6775096c (MD5) LICENSE.txt: 4211 bytes, checksum: a77a36712856784deef4c5a2e33d1940 (MD5) Previous issue date: 2020-05-13","Embargo set by: Seth Robbins for item 115964 Lift date: 2022-08-27T00:51:40Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Utilizing GPU tensor cores for algorithmic acceleration"]}]}],"canonical_facts":{"dc:contributor":["Hwu, Wen-Mei W"],"dc:creator":["Durrani, Sultan Hayat Khan"],"dc:date":["2020-08-27T00:51:33Z","2022-08-27T00:51:40Z","2020-05-13","2020-05"],"dc:description":["There has been a surge in the demand for a Domain Specific Architecture due to wide ranging deep learning applications like Image classification, speech recognition, in healthcare, self-driving cars etc. Matrix Multiplication acceleration has been a popular design choice when creating these specialized units to boost deep learning training and inference. Nvidia's Volta architecture introduced Tensor Cores which promised a 3 times speedup over their Pascal architecture. Despite the favorable performance gains, these accelerators have not been applied extensively to a wider class of algorithms. Through this thesis we introduce novel ways of mapping various algorithms on the Tensor Cores. We implemented Tensor Core based reduction, power iteration and Fast Fourier Transform (FFT) and show that effectively utilizing GPU compute resources would result in substantial gains in performance. Our reduction gave a 1.5 times speedup against CUB API; power iteration gave on average 2 times the speedup against Thrust and cuBLAS based implementation while our FFT implementation was able to outperform cuFFT with up to 8 times the speedup.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2022-05-01","The student, Sultan Durrani, accepted the attached license on 2020-05-12 at 13:44.","The student, Sultan Durrani, submitted this Thesis for approval on 2020-05-12 at 14:03.","This Thesis was approved for publication on 2020-05-13 at 11:11.","DSpace SAF Submission Ingestion Package generated from Vireo submission #15357 on 2020-08-25 at 17:44:25","Made available in DSpace on 2020-08-27T00:51:33Z (GMT). No. of bitstreams: 2 DURRANI-THESIS-2020.pdf: 1000595 bytes, checksum: 5b6eb05e70aa304a693e486c6775096c (MD5) LICENSE.txt: 4211 bytes, checksum: a77a36712856784deef4c5a2e33d1940 (MD5) Previous issue date: 2020-05-13","Embargo set by: Seth Robbins for item 115964 Lift date: 2022-08-27T00:51:40Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/108349"],"dc:language":["en"],"dc:rights":["Copyright 2020 Sultan Hayat Khan Durrani"],"dc:subject":["Computer Architecture","GPU","Tensor Cores","CUDA","FFT"],"dc:title":["Utilizing GPU tensor cores for algorithmic acceleration"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:48Z"}