{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/110732"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/110732","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Reconfigurable and heterogeneous architectures for efficient computing","abstract":"The saturation of single-thread performance, along with the advent of the power wall, has resulted in the need for efficient use of area and power budgets. With the end of Dennard scaling, and the slow down of Moore's law, scaling from one process node to another no longer delivers gains in performance or power for general-purpose computing. Thus, there is an increase in the adoption of specialized hardware, tuned to the requirements of the application or domain. These accelerators promise high performance and energy efficiency. However, with the increasing complexity and resource requirements of applications and algorithms, there is also a need for more flexibility in these accelerator platforms. Along with high performance and energy efficiency, they must be able to cope with changes at an application and algorithmic level. In the face of these challenges, this dissertation explores the use of reconfiguration to balance flexibility, performance, and energy efficiency. We begin by presenting three novel approaches that explore the use of reconfiguration in the three dominant computing devices -- CPUs, GPUs, and FPGAs. First, we consider general-purpose GPU (GPGPU) computing and highlight the inefficiencies in GPGPU, and identify opportunities to leverage reconfiguration to address these inefficiencies. Our solution is novel reconfigurable GPU architecture that can adapt to the needs of GPUs by dynamically allocating computational and memory resources among GPU cores (SMs). Second, we consider the limitations of dynamic partial reconfiguration (DPR) in modern FPGAs. We observe that while DPR is a potentially powerful technique, it is difficult to leverage. Thus, we propose an end-to-end methodology to leverage dynamic partial reconfiguration in FPGAs. The approach scales from edge to cloud devices, and presents an overlay architecture and an integer linear programming (ILP) based scheduler and mapper. We also demonstrate the ability to simultaneously map multiple applications to one FPGA, and explore different scheduling and sharing strategies. Third, we attempt to bridge the gap between the efficiency of reconfigurable computing and near-memory computing for general-purpose computing. Thus, we consider a modern multi-core CPU, and propose a novel architecture that uses SRAM arrays in the last level cache to create a reconfigurable computing fabric. Our approach is cheap, fast, energy-efficient, non-invasive, and flexible. Finally, this dissertation concludes by considering the lessons learned from exploiting reconfiguration on CPUs, GPUs, and FPGAs, and asks how a modern reconfigurable computing device should be designed. With the explosion of data, large computational workloads, and increasing demands of efficiency, we propose a new memory-centric reconfigurable architecture, capable of fast dynamic reconfiguration and altering its compute to memory ratio and organization. We demonstrate significantly higher performance, density, and memory capacity than modern FPGAs.","abstract_html":"The saturation of single-thread performance, along with the advent of the power wall, has resulted in the need for efficient use of area and power budgets. With the end of Dennard scaling, and the slow down of Moore&#x27;s law, scaling from one process node to another no longer delivers gains in performance or power for general-purpose computing. Thus, there is an increase in the adoption of specialized hardware, tuned to the requirements of the application or domain. These accelerators promise high performance and energy efficiency. However, with the increasing complexity and resource requirements of applications and algorithms, there is also a need for more flexibility in these accelerator platforms. Along with high performance and energy efficiency, they must be able to cope with changes at an application and algorithmic level. In the face of these challenges, this dissertation explores the use of reconfiguration to balance flexibility, performance, and energy efficiency. We begin by presenting three novel approaches that explore the use of reconfiguration in the three dominant computing devices -- CPUs, GPUs, and FPGAs. First, we consider general-purpose GPU (GPGPU) computing and highlight the inefficiencies in GPGPU, and identify opportunities to leverage reconfiguration to address these inefficiencies. Our solution is novel reconfigurable GPU architecture that can adapt to the needs of GPUs by dynamically allocating computational and memory resources among GPU cores (SMs). Second, we consider the limitations of dynamic partial reconfiguration (DPR) in modern FPGAs. We observe that while DPR is a potentially powerful technique, it is difficult to leverage. Thus, we propose an end-to-end methodology to leverage dynamic partial reconfiguration in FPGAs. The approach scales from edge to cloud devices, and presents an overlay architecture and an integer linear programming (ILP) based scheduler and mapper. We also demonstrate the ability to simultaneously map multiple applications to one FPGA, and explore different scheduling and sharing strategies. Third, we attempt to bridge the gap between the efficiency of reconfigurable computing and near-memory computing for general-purpose computing. Thus, we consider a modern multi-core CPU, and propose a novel architecture that uses SRAM arrays in the last level cache to create a reconfigurable computing fabric. Our approach is cheap, fast, energy-efficient, non-invasive, and flexible. Finally, this dissertation concludes by considering the lessons learned from exploiting reconfiguration on CPUs, GPUs, and FPGAs, and asks how a modern reconfigurable computing device should be designed. With the explosion of data, large computational workloads, and increasing demands of efficiency, we propose a new memory-centric reconfigurable architecture, capable of fast dynamic reconfiguration and altering its compute to memory ratio and organization. We demonstrate significantly higher performance, density, and memory capacity than modern FPGAs.","abstract_has_math":false,"creators":["Dhar, Ashutosh"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Chen, Deming","Hwu, Wen-mei","Torrellas, Josep","Xiong, Jinjun","Huang, Jian"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2021,"date_issued":"2021-09-17T02:34:45Z","date_published":"2021-09-17T02:34:45Z","updated_at":"2026-07-22T22:24:52Z","subjects":["Computer Architecture","Reconfigurable Computing","GPUs","FPGAs","CPUs"],"languages":["en"],"rights":["Copyright 2021 Ashutosh Dhar"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/110732","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Chen, Deming","Hwu, Wen-mei","Torrellas, Josep","Xiong, Jinjun","Huang, Jian"]},{"key":"dc:creator","label":"Author","values":["Dhar, Ashutosh"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2021-09-17T02:34:45Z","2023-09-17T02:34:57Z","2021-04-23","2021-05"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer Architecture","Reconfigurable Computing","GPUs","FPGAs","CPUs"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2021 Ashutosh Dhar"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/110732"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The saturation of single-thread performance, along with the advent of the power wall, has resulted in the need for efficient use of area and power budgets. With the end of Dennard scaling, and the slow down of Moore's law, scaling from one process node to another no longer delivers gains in performance or power for general-purpose computing. Thus, there is an increase in the adoption of specialized hardware, tuned to the requirements of the application or domain. These accelerators promise high performance and energy efficiency. However, with the increasing complexity and resource requirements of applications and algorithms, there is also a need for more flexibility in these accelerator platforms. Along with high performance and energy efficiency, they must be able to cope with changes at an application and algorithmic level. In the face of these challenges, this dissertation explores the use of reconfiguration to balance flexibility, performance, and energy efficiency. We begin by presenting three novel approaches that explore the use of reconfiguration in the three dominant computing devices -- CPUs, GPUs, and FPGAs. First, we consider general-purpose GPU (GPGPU) computing and highlight the inefficiencies in GPGPU, and identify opportunities to leverage reconfiguration to address these inefficiencies. Our solution is novel reconfigurable GPU architecture that can adapt to the needs of GPUs by dynamically allocating computational and memory resources among GPU cores (SMs). Second, we consider the limitations of dynamic partial reconfiguration (DPR) in modern FPGAs. We observe that while DPR is a potentially powerful technique, it is difficult to leverage. Thus, we propose an end-to-end methodology to leverage dynamic partial reconfiguration in FPGAs. The approach scales from edge to cloud devices, and presents an overlay architecture and an integer linear programming (ILP) based scheduler and mapper. We also demonstrate the ability to simultaneously map multiple applications to one FPGA, and explore different scheduling and sharing strategies. Third, we attempt to bridge the gap between the efficiency of reconfigurable computing and near-memory computing for general-purpose computing. Thus, we consider a modern multi-core CPU, and propose a novel architecture that uses SRAM arrays in the last level cache to create a reconfigurable computing fabric. Our approach is cheap, fast, energy-efficient, non-invasive, and flexible. Finally, this dissertation concludes by considering the lessons learned from exploiting reconfiguration on CPUs, GPUs, and FPGAs, and asks how a modern reconfigurable computing device should be designed. With the explosion of data, large computational workloads, and increasing demands of efficiency, we propose a new memory-centric reconfigurable architecture, capable of fast dynamic reconfiguration and altering its compute to memory ratio and organization. We demonstrate significantly higher performance, density, and memory capacity than modern FPGAs.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2023-05-01","The student, Ashutosh Dhar, accepted the attached license on 2021-04-22 at 19:11.","The student, Ashutosh Dhar, submitted this Dissertation for approval on 2021-04-22 at 19:22.","This Dissertation was approved for publication on 2021-04-23 at 08:48.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16518 on 2021-09-16 at 17:05:00","Made available in DSpace on 2021-09-17T02:34:45Z (GMT). No. of bitstreams: 2 DHAR-DISSERTATION-2021.pdf: 3961974 bytes, checksum: 3ede57178748228b2b08944c761b3fcf (MD5) LICENSE.txt: 4210 bytes, checksum: 16c27b2fae4dabc5e151a01dea7680c0 (MD5) Previous issue date: 2021-04-23","Embargo set by: Seth Robbins for item 118575 Lift date: 2023-09-17T02:34:57Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Reconfigurable and heterogeneous architectures for efficient computing"]}]}],"canonical_facts":{"dc:contributor":["Chen, Deming","Hwu, Wen-mei","Torrellas, Josep","Xiong, Jinjun","Huang, Jian"],"dc:creator":["Dhar, Ashutosh"],"dc:date":["2021-09-17T02:34:45Z","2023-09-17T02:34:57Z","2021-04-23","2021-05"],"dc:description":["The saturation of single-thread performance, along with the advent of the power wall, has resulted in the need for efficient use of area and power budgets. With the end of Dennard scaling, and the slow down of Moore's law, scaling from one process node to another no longer delivers gains in performance or power for general-purpose computing. Thus, there is an increase in the adoption of specialized hardware, tuned to the requirements of the application or domain. These accelerators promise high performance and energy efficiency. However, with the increasing complexity and resource requirements of applications and algorithms, there is also a need for more flexibility in these accelerator platforms. Along with high performance and energy efficiency, they must be able to cope with changes at an application and algorithmic level. In the face of these challenges, this dissertation explores the use of reconfiguration to balance flexibility, performance, and energy efficiency. We begin by presenting three novel approaches that explore the use of reconfiguration in the three dominant computing devices -- CPUs, GPUs, and FPGAs. First, we consider general-purpose GPU (GPGPU) computing and highlight the inefficiencies in GPGPU, and identify opportunities to leverage reconfiguration to address these inefficiencies. Our solution is novel reconfigurable GPU architecture that can adapt to the needs of GPUs by dynamically allocating computational and memory resources among GPU cores (SMs). Second, we consider the limitations of dynamic partial reconfiguration (DPR) in modern FPGAs. We observe that while DPR is a potentially powerful technique, it is difficult to leverage. Thus, we propose an end-to-end methodology to leverage dynamic partial reconfiguration in FPGAs. The approach scales from edge to cloud devices, and presents an overlay architecture and an integer linear programming (ILP) based scheduler and mapper. We also demonstrate the ability to simultaneously map multiple applications to one FPGA, and explore different scheduling and sharing strategies. Third, we attempt to bridge the gap between the efficiency of reconfigurable computing and near-memory computing for general-purpose computing. Thus, we consider a modern multi-core CPU, and propose a novel architecture that uses SRAM arrays in the last level cache to create a reconfigurable computing fabric. Our approach is cheap, fast, energy-efficient, non-invasive, and flexible. Finally, this dissertation concludes by considering the lessons learned from exploiting reconfiguration on CPUs, GPUs, and FPGAs, and asks how a modern reconfigurable computing device should be designed. With the explosion of data, large computational workloads, and increasing demands of efficiency, we propose a new memory-centric reconfigurable architecture, capable of fast dynamic reconfiguration and altering its compute to memory ratio and organization. We demonstrate significantly higher performance, density, and memory capacity than modern FPGAs.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2023-05-01","The student, Ashutosh Dhar, accepted the attached license on 2021-04-22 at 19:11.","The student, Ashutosh Dhar, submitted this Dissertation for approval on 2021-04-22 at 19:22.","This Dissertation was approved for publication on 2021-04-23 at 08:48.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16518 on 2021-09-16 at 17:05:00","Made available in DSpace on 2021-09-17T02:34:45Z (GMT). No. of bitstreams: 2 DHAR-DISSERTATION-2021.pdf: 3961974 bytes, checksum: 3ede57178748228b2b08944c761b3fcf (MD5) LICENSE.txt: 4210 bytes, checksum: 16c27b2fae4dabc5e151a01dea7680c0 (MD5) Previous issue date: 2021-04-23","Embargo set by: Seth Robbins for item 118575 Lift date: 2023-09-17T02:34:57Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/110732"],"dc:language":["en"],"dc:rights":["Copyright 2021 Ashutosh Dhar"],"dc:subject":["Computer Architecture","Reconfigurable Computing","GPUs","FPGAs","CPUs"],"dc:title":["Reconfigurable and heterogeneous architectures for efficient computing"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:52Z"}