{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/99536"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/99536","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"FCUDA: Efficient high-level automation CUDA-to-FPGA compilation","abstract":"The demand for high-performance computing has been growing significantly in the past decade. The bottleneck of Moore's law and the increasing power consumption in the traditional computing industry have stimulated the popularity of parallel computing. GPUs and FPGAs became popular and played very important roles in heterogeneous systems for accelerating various compute intensive tasks in different areas. Modern GPUs can execute more than thousands of threads, providing strong parallelism. FPGAs, however, provide highly customized concurrency for parallel kernels. The current version of source-to-source compiler FCUDA, which transforms CUDA kernel code into synthesizable C code, exploits the parallelism in different applications with the help of the manually inserted pragmas by the programmers. The additional effort to tweak the code to enable efficient mapping of the tasks across the heterogeneous architectures cannot be ignored. In this thesis, a new code optimization flow is proposed. The flow will restructure and analyze the CUDA kernel code, optimizing the performance by extracting the parallelism in GPU devices. The generated C code will further be synthesized and programmed on FPGAs. With help of the new flow, there is no need for programmers to manually annotate and tweak the source code, making the whole process a push-button one.","abstract_html":"The demand for high-performance computing has been growing significantly in the past decade. The bottleneck of Moore&#x27;s law and the increasing power consumption in the traditional computing industry have stimulated the popularity of parallel computing. GPUs and FPGAs became popular and played very important roles in heterogeneous systems for accelerating various compute intensive tasks in different areas. Modern GPUs can execute more than thousands of threads, providing strong parallelism. FPGAs, however, provide highly customized concurrency for parallel kernels. The current version of source-to-source compiler FCUDA, which transforms CUDA kernel code into synthesizable C code, exploits the parallelism in different applications with the help of the manually inserted pragmas by the programmers. The additional effort to tweak the code to enable efficient mapping of the tasks across the heterogeneous architectures cannot be ignored. In this thesis, a new code optimization flow is proposed. The flow will restructure and analyze the CUDA kernel code, optimizing the performance by extracting the parallelism in GPU devices. The generated C code will further be synthesized and programmed on FPGAs. With help of the new flow, there is no need for programmers to manually annotate and tweak the source code, making the whole process a push-button one.","abstract_has_math":false,"creators":["Xu, Zhangqi"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Electrical and Computer Engineering","degree_department":null,"school":null,"contributors":["Chen, Deming"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2018,"date_issued":"2018-03-13T17:35:57Z","date_published":"2018-03-13T17:35:57Z","updated_at":"2026-07-22T22:24:37Z","subjects":["Field programmable gate array (FPGA)","High-level synthesis (HLS)","Compute unified device architecture (CUDA)"],"languages":["en"],"rights":["Copyright 2017 Zhangqi Xu"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/99536","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Chen, Deming"]},{"key":"dc:creator","label":"Author","values":["Xu, Zhangqi"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2018-03-13T17:35:57Z","2020-03-14T09:15:12Z","2017-12-14","2017-12"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical and Computer Engineering"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Field programmable gate array (FPGA)","High-level synthesis (HLS)","Compute unified device architecture (CUDA)"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2017 Zhangqi Xu"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/99536"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The demand for high-performance computing has been growing significantly in the past decade. The bottleneck of Moore's law and the increasing power consumption in the traditional computing industry have stimulated the popularity of parallel computing. GPUs and FPGAs became popular and played very important roles in heterogeneous systems for accelerating various compute intensive tasks in different areas. Modern GPUs can execute more than thousands of threads, providing strong parallelism. FPGAs, however, provide highly customized concurrency for parallel kernels. The current version of source-to-source compiler FCUDA, which transforms CUDA kernel code into synthesizable C code, exploits the parallelism in different applications with the help of the manually inserted pragmas by the programmers. The additional effort to tweak the code to enable efficient mapping of the tasks across the heterogeneous architectures cannot be ignored. In this thesis, a new code optimization flow is proposed. The flow will restructure and analyze the CUDA kernel code, optimizing the performance by extracting the parallelism in GPU devices. The generated C code will further be synthesized and programmed on FPGAs. With help of the new flow, there is no need for programmers to manually annotate and tweak the source code, making the whole process a push-button one.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2019-12-01","The student, Zhangqi Xu, accepted the attached license on 2017-12-13 at 22:29.","The student, Zhangqi Xu, submitted this Thesis for approval on 2017-12-13 at 22:42.","This Thesis was approved for publication on 2017-12-14 at 09:42.","DSpace SAF Submission Ingestion Package generated from Vireo submission #11977 on 2018-03-13 at 10:38:26","Made available in DSpace on 2018-03-13T17:35:57Z (GMT). No. of bitstreams: 2 XU-THESIS-2017.pdf: 1069692 bytes, checksum: 589309adf9eec53dd9da0088aa6cd368 (MD5) LICENSE.txt: 4207 bytes, checksum: 1e4bcfed14fc9b633451fe9b9990c3b4 (MD5) Previous issue date: 2017-12-14","Embargo set by: Seth Robbins for item 105505 Lift date: 2020-03-13T17:36:05Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 105505 on 2020-03-14T09:15:12Z."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["FCUDA: Efficient high-level automation CUDA-to-FPGA compilation"]}]}],"canonical_facts":{"dc:contributor":["Chen, Deming"],"dc:creator":["Xu, Zhangqi"],"dc:date":["2018-03-13T17:35:57Z","2020-03-14T09:15:12Z","2017-12-14","2017-12"],"dc:description":["The demand for high-performance computing has been growing significantly in the past decade. The bottleneck of Moore's law and the increasing power consumption in the traditional computing industry have stimulated the popularity of parallel computing. GPUs and FPGAs became popular and played very important roles in heterogeneous systems for accelerating various compute intensive tasks in different areas. Modern GPUs can execute more than thousands of threads, providing strong parallelism. FPGAs, however, provide highly customized concurrency for parallel kernels. The current version of source-to-source compiler FCUDA, which transforms CUDA kernel code into synthesizable C code, exploits the parallelism in different applications with the help of the manually inserted pragmas by the programmers. The additional effort to tweak the code to enable efficient mapping of the tasks across the heterogeneous architectures cannot be ignored. In this thesis, a new code optimization flow is proposed. The flow will restructure and analyze the CUDA kernel code, optimizing the performance by extracting the parallelism in GPU devices. The generated C code will further be synthesized and programmed on FPGAs. With help of the new flow, there is no need for programmers to manually annotate and tweak the source code, making the whole process a push-button one.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2019-12-01","The student, Zhangqi Xu, accepted the attached license on 2017-12-13 at 22:29.","The student, Zhangqi Xu, submitted this Thesis for approval on 2017-12-13 at 22:42.","This Thesis was approved for publication on 2017-12-14 at 09:42.","DSpace SAF Submission Ingestion Package generated from Vireo submission #11977 on 2018-03-13 at 10:38:26","Made available in DSpace on 2018-03-13T17:35:57Z (GMT). No. of bitstreams: 2 XU-THESIS-2017.pdf: 1069692 bytes, checksum: 589309adf9eec53dd9da0088aa6cd368 (MD5) LICENSE.txt: 4207 bytes, checksum: 1e4bcfed14fc9b633451fe9b9990c3b4 (MD5) Previous issue date: 2017-12-14","Embargo set by: Seth Robbins for item 105505 Lift date: 2020-03-13T17:36:05Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 105505 on 2020-03-14T09:15:12Z."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/99536"],"dc:language":["en"],"dc:rights":["Copyright 2017 Zhangqi Xu"],"dc:subject":["Field programmable gate array (FPGA)","High-level synthesis (HLS)","Compute unified device architecture (CUDA)"],"dc:title":["FCUDA: Efficient high-level automation CUDA-to-FPGA compilation"],"dc:type":["text"],"thesis:degree_discipline":["Electrical and Computer Engineering"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:37Z"}