{"id":{"repo_id":"wustl","oai_identifier":"oai:openscholarship.wustl.edu:eng_etds-1764"},"canonical_url":"https://search.dev.ndltd.org/etd/wustl/oai:openscholarship.wustl.edu:eng_etds-1764","repository":{"repo_id":"wustl","name":"Washington University in St. Louis","base_url":"https://openscholarship.wustl.edu/do/oai/"},"display":{"title":"A Reconfigurable FPGA Overlay Architecture for Matrix-Matrix Multiplication","abstract":"<p>The increasing popularity of deep learning in workloads across vision, speech, and language has inspired many attempts to develop hardware accelerators for matrix-matrix multiplication. Both application-specific integrated circuits (ASICs), and field-programmable arrays (FPGAs) are used for this purpose. However, a trade-off between the two platforms is that ASICs provide little flexibility after they are manufactured while designs on FPGAs are flexible but application development on FPGAs is more time-consuming. In this work, we aim to find the balance between reconfigurability and development efficiency by designing a reconfigurable systolic architecture as an overlay on the FPGA. Our contribution to the reconfigurable systolic architectures is a multiplexer-based crossbar network that interconnects every processing element in the network. The crossbar network grants user run-time reconfigurability of the topology of the systolic array, enabling the user to specify the shape and size of the systolic architecture on-the-fly. The proposed overlay architecture achieves similar computational hardware resource usage and maximum clock frequency compared to the baseline designs.</p>","abstract_html":"&lt;p&gt;The increasing popularity of deep learning in workloads across vision, speech, and language has inspired many attempts to develop hardware accelerators for matrix-matrix multiplication. Both application-specific integrated circuits (ASICs), and field-programmable arrays (FPGAs) are used for this purpose. However, a trade-off between the two platforms is that ASICs provide little flexibility after they are manufactured while designs on FPGAs are flexible but application development on FPGAs is more time-consuming. In this work, we aim to find the balance between reconfigurability and development efficiency by designing a reconfigurable systolic architecture as an overlay on the FPGA. Our contribution to the reconfigurable systolic architectures is a multiplexer-based crossbar network that interconnects every processing element in the network. The crossbar network grants user run-time reconfigurability of the topology of the systolic array, enabling the user to specify the shape and size of the systolic architecture on-the-fly. The proposed overlay architecture achieves similar computational hardware resource usage and maximum clock frequency compared to the baseline designs.&lt;/p&gt;","abstract_has_math":false,"creators":["Chen, Zihao"],"institution":null,"degree_name":"Master of Science (MS)","degree_level":"Thesis","degree_discipline":"Computer Science & Engineering","degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-04-25T07:00:00Z","date_published":"2022-04-25T07:00:00Z","updated_at":"2026-07-24T06:12:58Z","subjects":["FPGA overlay","Engineering"],"languages":["English (en)"],"rights":["I have not registered my thesis with the U.S. Copyright Office, and do not intend to."],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["https://openscholarship.wustl.edu/eng_etds/710"],"render_values":[{"text":"https://openscholarship.wustl.edu/eng_etds/710","href":"https://openscholarship.wustl.edu/eng_etds/710","code":true}]}]},"links":{"outbound_url":"https://doi.org/10.7936/76vz-dv86","outbound_label":"DOI","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["Chen, Zihao"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2022-04-25T07:00:00Z"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science & Engineering","McKelvey School of Engineering"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science (MS)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["FPGA overlay","Engineering"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["English (en)"]},{"key":"dc:rights","label":"Dc Rights","values":["I have not registered my thesis with the U.S. Copyright Office, and do not intend to."]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://doi.org/10.7936/76vz-dv86","https://openscholarship.wustl.edu/eng_etds/710"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>The increasing popularity of deep learning in workloads across vision, speech, and language has inspired many attempts to develop hardware accelerators for matrix-matrix multiplication. Both application-specific integrated circuits (ASICs), and field-programmable arrays (FPGAs) are used for this purpose. However, a trade-off between the two platforms is that ASICs provide little flexibility after they are manufactured while designs on FPGAs are flexible but application development on FPGAs is more time-consuming. In this work, we aim to find the balance between reconfigurability and development efficiency by designing a reconfigurable systolic architecture as an overlay on the FPGA. Our contribution to the reconfigurable systolic architectures is a multiplexer-based crossbar network that interconnects every processing element in the network. The crossbar network grants user run-time reconfigurability of the topology of the systolic array, enabling the user to specify the shape and size of the systolic architecture on-the-fly. The proposed overlay architecture achieves similar computational hardware resource usage and maximum clock frequency compared to the baseline designs.</p>"]},{"key":"dc:title","label":"Title","values":["A Reconfigurable FPGA Overlay Architecture for Matrix-Matrix Multiplication"]}]}],"canonical_facts":{"dc:creator":["Chen, Zihao"],"dc:date.available":["2022-04-25T07:00:00Z"],"dc:description.abstract":["<p>The increasing popularity of deep learning in workloads across vision, speech, and language has inspired many attempts to develop hardware accelerators for matrix-matrix multiplication. Both application-specific integrated circuits (ASICs), and field-programmable arrays (FPGAs) are used for this purpose. However, a trade-off between the two platforms is that ASICs provide little flexibility after they are manufactured while designs on FPGAs are flexible but application development on FPGAs is more time-consuming. In this work, we aim to find the balance between reconfigurability and development efficiency by designing a reconfigurable systolic architecture as an overlay on the FPGA. Our contribution to the reconfigurable systolic architectures is a multiplexer-based crossbar network that interconnects every processing element in the network. The crossbar network grants user run-time reconfigurability of the topology of the systolic array, enabling the user to specify the shape and size of the systolic architecture on-the-fly. The proposed overlay architecture achieves similar computational hardware resource usage and maximum clock frequency compared to the baseline designs.</p>"],"dc:identifier":["https://doi.org/10.7936/76vz-dv86","https://openscholarship.wustl.edu/eng_etds/710"],"dc:language":["English (en)"],"dc:rights":["I have not registered my thesis with the U.S. Copyright Office, and do not intend to."],"dc:subject":["FPGA overlay","Engineering"],"dc:title":["A Reconfigurable FPGA Overlay Architecture for Matrix-Matrix Multiplication"],"thesis:degree_discipline":["Computer Science & Engineering","McKelvey School of Engineering"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["Master of Science (MS)"]},"updated_at":"2026-07-24T06:12:58Z"}