{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/78686"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/78686","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Hierarchical and scalable bus architecture generation on FPGAs with high-level synthesis","abstract":"This thesis presents and evaluates a bus-based system for FCUDA, a translation tool enabling CUDA code to be run on FPGAs. With the goal of constructing a solid light-weight back-end with optimized performance, we choose AXI4 as the communication protocol and plug in all necessary components on a hierarchical bus system. Then, FCUDA cores are added in the back-end and the comprehensive system is automated into a single tool chain. Several optimizations are added in this automated FCUDA bus system for the delivery of better performance. For example, FCUDA cores are tiled into clusters based on configuration inputs, and clock domains are separated to reduce long wires. For the experiments, this work adjusts the existing resources and period models and enhances the system latency model by incorporating off-chip memory communication latency. The system is proved to be light-weight based on post-routing resource reports. Design space exploration among multilevel granularity parallelisms is performed to get the system's best performance, with which a comparison with GPU is made. Our system can achieve at most 2.08 performance improvement when compared with the execution latency on GPU.","abstract_html":"This thesis presents and evaluates a bus-based system for FCUDA, a translation tool enabling CUDA code to be run on FPGAs. With the goal of constructing a solid light-weight back-end with optimized performance, we choose AXI4 as the communication protocol and plug in all necessary components on a hierarchical bus system. Then, FCUDA cores are added in the back-end and the comprehensive system is automated into a single tool chain. Several optimizations are added in this automated FCUDA bus system for the delivery of better performance. For example, FCUDA cores are tiled into clusters based on configuration inputs, and clock domains are separated to reduce long wires. For the experiments, this work adjusts the existing resources and period models and enhances the system latency model by incorporating off-chip memory communication latency. The system is proved to be light-weight based on post-routing resource reports. Design space exploration among multilevel granularity parallelisms is performed to get the system&#x27;s best performance, with which a comparison with GPU is made. Our system can achieve at most 2.08 performance improvement when compared with the execution latency on GPU.","abstract_has_math":false,"creators":["Chen, Ying"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2015,"date_issued":"2015-07-22T22:34:01Z","date_published":"2015-07-22T22:34:01Z","updated_at":"2026-07-22T22:26:12Z","subjects":["FCUDA","Advanced eXtensible Interface (AXI)","BUS","High-Level Synthesis","Compute Unified Device Architecture (CUDA)"],"languages":["en"],"rights":["Copyright 2015 Ying Chen"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/78686","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["Chen, Ying"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2015-07-22T22:34:01Z","2017-07-23T09:15:36Z","2015-05","2015-04-30","2015-5"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["FCUDA","Advanced eXtensible Interface (AXI)","BUS","High-Level Synthesis","Compute Unified Device Architecture (CUDA)"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2015 Ying Chen"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/78686"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["This thesis presents and evaluates a bus-based system for FCUDA, a translation tool enabling CUDA code to be run on FPGAs. With the goal of constructing a solid light-weight back-end with optimized performance, we choose AXI4 as the communication protocol and plug in all necessary components on a hierarchical bus system. Then, FCUDA cores are added in the back-end and the comprehensive system is automated into a single tool chain. Several optimizations are added in this automated FCUDA bus system for the delivery of better performance. For example, FCUDA cores are tiled into clusters based on configuration inputs, and clock domains are separated to reduce long wires. For the experiments, this work adjusts the existing resources and period models and enhances the system latency model by incorporating off-chip memory communication latency. The system is proved to be light-weight based on post-routing resource reports. Design space exploration among multilevel granularity parallelisms is performed to get the system's best performance, with which a comparison with GPU is made. Our system can achieve at most 2.08 performance improvement when compared with the execution latency on GPU.","Submission published under a 24 month embargo labeled 'U of I only', the embargo will last until 2017-05-01","The student, Ying Chen, accepted the attached license on 2015-04-28 at 22:43.","The student, Ying Chen, submitted this Thesis for approval on 2015-04-28 at 23:01.","This Thesis was approved for publication on 2015-04-30 at 10:02.","DSpace SAF Submission Ingestion Package generated from Vireo submission #8173 on 2015-07-22 at 14:19:00","Made available in DSpace on 2015-07-22T22:34:01Z (GMT). No. of bitstreams: 2 CHEN-THESIS-2015.pdf: 890535 bytes, checksum: af07fa90834a8f25563a48e21add1450 (MD5) LICENSE.txt: 4206 bytes, checksum: f52bd2542e3f37df23ab3cb8f94171f2 (MD5) Previous issue date: 2015-04-30","Embargo set by: Seth Robbins for item 79927 Lift date: 2017-07-22T22:34:16Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 79927 on 2017-07-23T09:15:36Z."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Hierarchical and scalable bus architecture generation on FPGAs with high-level synthesis"]}]}],"canonical_facts":{"dc:creator":["Chen, Ying"],"dc:date":["2015-07-22T22:34:01Z","2017-07-23T09:15:36Z","2015-05","2015-04-30","2015-5"],"dc:description":["This thesis presents and evaluates a bus-based system for FCUDA, a translation tool enabling CUDA code to be run on FPGAs. With the goal of constructing a solid light-weight back-end with optimized performance, we choose AXI4 as the communication protocol and plug in all necessary components on a hierarchical bus system. Then, FCUDA cores are added in the back-end and the comprehensive system is automated into a single tool chain. Several optimizations are added in this automated FCUDA bus system for the delivery of better performance. For example, FCUDA cores are tiled into clusters based on configuration inputs, and clock domains are separated to reduce long wires. For the experiments, this work adjusts the existing resources and period models and enhances the system latency model by incorporating off-chip memory communication latency. The system is proved to be light-weight based on post-routing resource reports. Design space exploration among multilevel granularity parallelisms is performed to get the system's best performance, with which a comparison with GPU is made. Our system can achieve at most 2.08 performance improvement when compared with the execution latency on GPU.","Submission published under a 24 month embargo labeled 'U of I only', the embargo will last until 2017-05-01","The student, Ying Chen, accepted the attached license on 2015-04-28 at 22:43.","The student, Ying Chen, submitted this Thesis for approval on 2015-04-28 at 23:01.","This Thesis was approved for publication on 2015-04-30 at 10:02.","DSpace SAF Submission Ingestion Package generated from Vireo submission #8173 on 2015-07-22 at 14:19:00","Made available in DSpace on 2015-07-22T22:34:01Z (GMT). No. of bitstreams: 2 CHEN-THESIS-2015.pdf: 890535 bytes, checksum: af07fa90834a8f25563a48e21add1450 (MD5) LICENSE.txt: 4206 bytes, checksum: f52bd2542e3f37df23ab3cb8f94171f2 (MD5) Previous issue date: 2015-04-30","Embargo set by: Seth Robbins for item 79927 Lift date: 2017-07-22T22:34:16Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 79927 on 2017-07-23T09:15:36Z."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/78686"],"dc:language":["en"],"dc:rights":["Copyright 2015 Ying Chen"],"dc:subject":["FCUDA","Advanced eXtensible Interface (AXI)","BUS","High-Level Synthesis","Compute Unified Device Architecture (CUDA)"],"dc:title":["Hierarchical and scalable bus architecture generation on FPGAs with high-level synthesis"],"dc:type":["text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:26:12Z"}