{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/108177"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/108177","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"FOpenCL: An optimized OpenCL-to-FPGA design flow for Xilinx SDx","abstract":"The semiconductor industry has been working constantly to reduce transistor size and thereby to get the last bit of performance out of single-core processors. With miniaturization becoming more difficult due to the physical limit and fabrication technique limit, alternative techniques to increase computing performance, like parallelism, have become more popular. GPU and FPGA are two significant examples of accelerators that provide high-degree parallelism and can yield significant speed-up if the developer properly understands the parallelism. However, in reality, it is very difficult to explore all the fine and coarse grained parallelism in the specific applications. Previous work has explored the adaptation of CUDA code to FPGA synthesizable C code. In our work, we discover and explain why OpenCL kernels designed for GPU will have poor performance on FPGA if the kernels are copied over without transformation and optimization. Specifically, we uncover an implicit limitation in Xilinx SDx, which claims to support OpenCL. Then, we explore the possibility of transforming OpenCL kernels to achieve better performance and hardware usage on FPGAs. Our proposed FOpenCL flow works around the limitation of Xilinx SDx, and gives GPU developers the ability to easily migrate their OpenCL kernels to FPGA and achieve superior performance. Our results show that FOpenCL transformed OpenCL kernels can achieve significant speed-up on Xilinx SDx with reasonable overhead on hardware resources. We believe FOpenCL is the first flow that targets automatically migrating GPU OpenCL kernels to optimized FPGA OpenCL kernels.","abstract_html":"The semiconductor industry has been working constantly to reduce transistor size and thereby to get the last bit of performance out of single-core processors. With miniaturization becoming more difficult due to the physical limit and fabrication technique limit, alternative techniques to increase computing performance, like parallelism, have become more popular. GPU and FPGA are two significant examples of accelerators that provide high-degree parallelism and can yield significant speed-up if the developer properly understands the parallelism. However, in reality, it is very difficult to explore all the fine and coarse grained parallelism in the specific applications. Previous work has explored the adaptation of CUDA code to FPGA synthesizable C code. In our work, we discover and explain why OpenCL kernels designed for GPU will have poor performance on FPGA if the kernels are copied over without transformation and optimization. Specifically, we uncover an implicit limitation in Xilinx SDx, which claims to support OpenCL. Then, we explore the possibility of transforming OpenCL kernels to achieve better performance and hardware usage on FPGAs. Our proposed FOpenCL flow works around the limitation of Xilinx SDx, and gives GPU developers the ability to easily migrate their OpenCL kernels to FPGA and achieve superior performance. Our results show that FOpenCL transformed OpenCL kernels can achieve significant speed-up on Xilinx SDx with reasonable overhead on hardware resources. We believe FOpenCL is the first flow that targets automatically migrating GPU OpenCL kernels to optimized FPGA OpenCL kernels.","abstract_has_math":false,"creators":["Zhou, Xingkai"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Chen, Deming"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2020,"date_issued":"2020-08-26T23:58:45Z","date_published":"2020-08-26T23:58:45Z","updated_at":"2026-07-22T22:24:47Z","subjects":["OpenCL","FPGA","Source-to-source compilation"],"languages":["en"],"rights":["Copyright 2020 Xingkai Zhou"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/108177","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Chen, Deming"]},{"key":"dc:creator","label":"Author","values":["Zhou, Xingkai"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2020-08-26T23:58:45Z","2022-08-26T23:58:55Z","2020-05-14","2020-05"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["OpenCL","FPGA","Source-to-source compilation"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2020 Xingkai Zhou"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/108177"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The semiconductor industry has been working constantly to reduce transistor size and thereby to get the last bit of performance out of single-core processors. With miniaturization becoming more difficult due to the physical limit and fabrication technique limit, alternative techniques to increase computing performance, like parallelism, have become more popular. GPU and FPGA are two significant examples of accelerators that provide high-degree parallelism and can yield significant speed-up if the developer properly understands the parallelism. However, in reality, it is very difficult to explore all the fine and coarse grained parallelism in the specific applications. Previous work has explored the adaptation of CUDA code to FPGA synthesizable C code. In our work, we discover and explain why OpenCL kernels designed for GPU will have poor performance on FPGA if the kernels are copied over without transformation and optimization. Specifically, we uncover an implicit limitation in Xilinx SDx, which claims to support OpenCL. Then, we explore the possibility of transforming OpenCL kernels to achieve better performance and hardware usage on FPGAs. Our proposed FOpenCL flow works around the limitation of Xilinx SDx, and gives GPU developers the ability to easily migrate their OpenCL kernels to FPGA and achieve superior performance. Our results show that FOpenCL transformed OpenCL kernels can achieve significant speed-up on Xilinx SDx with reasonable overhead on hardware resources. We believe FOpenCL is the first flow that targets automatically migrating GPU OpenCL kernels to optimized FPGA OpenCL kernels.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2022-05-01","The student, Xingkai Zhou, accepted the attached license on 2020-05-13 at 10:14.","The student, Xingkai Zhou, submitted this Thesis for approval on 2020-05-13 at 10:32.","This Thesis was approved for publication on 2020-05-14 at 09:43.","DSpace SAF Submission Ingestion Package generated from Vireo submission #15313 on 2020-08-25 at 17:30:48","Made available in DSpace on 2020-08-26T23:58:45Z (GMT). No. of bitstreams: 2 ZHOU-THESIS-2020.pdf: 995105 bytes, checksum: a62ed771a9c6a13f48ac3306f53a25d4 (MD5) LICENSE.txt: 4209 bytes, checksum: efe83858f69f78aedd45dcc25f6c592e (MD5) Previous issue date: 2020-05-14","Embargo set by: Seth Robbins for item 115790 Lift date: 2022-08-26T23:58:55Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["FOpenCL: An optimized OpenCL-to-FPGA design flow for Xilinx SDx"]}]}],"canonical_facts":{"dc:contributor":["Chen, Deming"],"dc:creator":["Zhou, Xingkai"],"dc:date":["2020-08-26T23:58:45Z","2022-08-26T23:58:55Z","2020-05-14","2020-05"],"dc:description":["The semiconductor industry has been working constantly to reduce transistor size and thereby to get the last bit of performance out of single-core processors. With miniaturization becoming more difficult due to the physical limit and fabrication technique limit, alternative techniques to increase computing performance, like parallelism, have become more popular. GPU and FPGA are two significant examples of accelerators that provide high-degree parallelism and can yield significant speed-up if the developer properly understands the parallelism. However, in reality, it is very difficult to explore all the fine and coarse grained parallelism in the specific applications. Previous work has explored the adaptation of CUDA code to FPGA synthesizable C code. In our work, we discover and explain why OpenCL kernels designed for GPU will have poor performance on FPGA if the kernels are copied over without transformation and optimization. Specifically, we uncover an implicit limitation in Xilinx SDx, which claims to support OpenCL. Then, we explore the possibility of transforming OpenCL kernels to achieve better performance and hardware usage on FPGAs. Our proposed FOpenCL flow works around the limitation of Xilinx SDx, and gives GPU developers the ability to easily migrate their OpenCL kernels to FPGA and achieve superior performance. Our results show that FOpenCL transformed OpenCL kernels can achieve significant speed-up on Xilinx SDx with reasonable overhead on hardware resources. We believe FOpenCL is the first flow that targets automatically migrating GPU OpenCL kernels to optimized FPGA OpenCL kernels.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2022-05-01","The student, Xingkai Zhou, accepted the attached license on 2020-05-13 at 10:14.","The student, Xingkai Zhou, submitted this Thesis for approval on 2020-05-13 at 10:32.","This Thesis was approved for publication on 2020-05-14 at 09:43.","DSpace SAF Submission Ingestion Package generated from Vireo submission #15313 on 2020-08-25 at 17:30:48","Made available in DSpace on 2020-08-26T23:58:45Z (GMT). No. of bitstreams: 2 ZHOU-THESIS-2020.pdf: 995105 bytes, checksum: a62ed771a9c6a13f48ac3306f53a25d4 (MD5) LICENSE.txt: 4209 bytes, checksum: efe83858f69f78aedd45dcc25f6c592e (MD5) Previous issue date: 2020-05-14","Embargo set by: Seth Robbins for item 115790 Lift date: 2022-08-26T23:58:55Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/108177"],"dc:language":["en"],"dc:rights":["Copyright 2020 Xingkai Zhou"],"dc:subject":["OpenCL","FPGA","Source-to-source compilation"],"dc:title":["FOpenCL: An optimized OpenCL-to-FPGA design flow for Xilinx SDx"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:47Z"}