{"id":{"repo_id":"wustl","oai_identifier":"oai:openscholarship.wustl.edu:eng_etds-1592"},"canonical_url":"https://search.dev.ndltd.org/etd/wustl/oai:openscholarship.wustl.edu:eng_etds-1592","repository":{"repo_id":"wustl","name":"Washington University in St. Louis","base_url":"https://openscholarship.wustl.edu/do/oai/"},"display":{"title":"Investigating Single Precision Floating General Matrix Multiply in Heterogeneous Hardware","abstract":"<p>The fundamental operation of matrix multiplication is ubiquitous across a myriad of disciplines. Yet, the identification of new optimizations for matrix multiplication remains relevant for emerging hardware architectures and heterogeneous systems. Frameworks such as OpenCL enable computation orchestration on existing systems, and its availability using the Intel High Level Synthesis compiler allows users to architect new designs for reconfigurable hardware using C/C++. Using the HARPv2 as a vehicle for exploration, we investigate the utility of several of the most notable matrix multiplication optimizations to better understand the performance portability of OpenCL and the implications for such optimizations on this and future heterogeneous architectures. Our results give targeted insights into the applicability of best practices that were for existing architectures when used on emerging heterogeneous systems.</p>","abstract_html":"&lt;p&gt;The fundamental operation of matrix multiplication is ubiquitous across a myriad of disciplines. Yet, the identification of new optimizations for matrix multiplication remains relevant for emerging hardware architectures and heterogeneous systems. Frameworks such as OpenCL enable computation orchestration on existing systems, and its availability using the Intel High Level Synthesis compiler allows users to architect new designs for reconfigurable hardware using C/C++. Using the HARPv2 as a vehicle for exploration, we investigate the utility of several of the most notable matrix multiplication optimizations to better understand the performance portability of OpenCL and the implications for such optimizations on this and future heterogeneous architectures. Our results give targeted insights into the applicability of best practices that were for existing architectures when used on emerging heterogeneous systems.&lt;/p&gt;","abstract_has_math":false,"creators":["Harris, Steven"],"institution":null,"degree_name":"Master of Science (MS)","degree_level":"Thesis","degree_discipline":"Computer Science & Engineering","degree_department":null,"school":null,"contributors":["guerin@wustl.edu","Professor Roger Chamberlain, Professor Christopher Gill, Professor Ning Zhang"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2020,"date_issued":"2020-08-01T07:00:00Z","date_published":"2020-08-01T07:00:00Z","updated_at":"2026-07-24T06:13:31Z","subjects":["Design space search","HARP","High-level synthesis","OpenCL","SGEMM","Field programmable gate array","heterogeneous architecture","Computer and Systems Architecture","Engineering","Hardware Systems","Numerical Analysis and Scientific Computing","Software Engineering","Systems Architecture","Theory and Algorithms","VLSI and Circuits, Embedded and Hardware Systems"],"languages":["English (en)"],"rights":["I have not registered my thesis with the U.S. Copyright Office, and do not intend to."],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["https://openscholarship.wustl.edu/eng_etds/536"],"render_values":[{"text":"https://openscholarship.wustl.edu/eng_etds/536","href":"https://openscholarship.wustl.edu/eng_etds/536","code":true}]}]},"links":{"outbound_url":"https://doi.org/10.7936/dbf4-6j92","outbound_label":"DOI","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["guerin@wustl.edu","Professor Roger Chamberlain, Professor Christopher Gill, Professor Ning Zhang"]},{"key":"dc:creator","label":"Author","values":["Harris, Steven"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2020-05-29T07:00:00Z"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science & Engineering","McKelvey School of Engineering"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science (MS)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Design space search","HARP","High-level synthesis","OpenCL","SGEMM","Field programmable gate array","heterogeneous architecture","Computer and Systems Architecture","Engineering","Hardware Systems","Numerical Analysis and Scientific Computing","Software Engineering","Systems Architecture","Theory and Algorithms","VLSI and Circuits, Embedded and Hardware Systems"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["English (en)"]},{"key":"dc:rights","label":"Dc Rights","values":["I have not registered my thesis with the U.S. Copyright Office, and do not intend to."]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://doi.org/10.7936/dbf4-6j92","https://openscholarship.wustl.edu/eng_etds/536"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["<p>Permanent URL: https://doi.org/10.7936/dbf4-6j92</p>"]},{"key":"dc:description.abstract","label":"Abstract","values":["<p>The fundamental operation of matrix multiplication is ubiquitous across a myriad of disciplines. Yet, the identification of new optimizations for matrix multiplication remains relevant for emerging hardware architectures and heterogeneous systems. Frameworks such as OpenCL enable computation orchestration on existing systems, and its availability using the Intel High Level Synthesis compiler allows users to architect new designs for reconfigurable hardware using C/C++. Using the HARPv2 as a vehicle for exploration, we investigate the utility of several of the most notable matrix multiplication optimizations to better understand the performance portability of OpenCL and the implications for such optimizations on this and future heterogeneous architectures. Our results give targeted insights into the applicability of best practices that were for existing architectures when used on emerging heterogeneous systems.</p>"]},{"key":"dc:title","label":"Title","values":["Investigating Single Precision Floating General Matrix Multiply in Heterogeneous Hardware"]}]}],"canonical_facts":{"dc:contributor":["guerin@wustl.edu","Professor Roger Chamberlain, Professor Christopher Gill, Professor Ning Zhang"],"dc:creator":["Harris, Steven"],"dc:date.available":["2020-05-29T07:00:00Z"],"dc:description":["<p>Permanent URL: https://doi.org/10.7936/dbf4-6j92</p>"],"dc:description.abstract":["<p>The fundamental operation of matrix multiplication is ubiquitous across a myriad of disciplines. Yet, the identification of new optimizations for matrix multiplication remains relevant for emerging hardware architectures and heterogeneous systems. Frameworks such as OpenCL enable computation orchestration on existing systems, and its availability using the Intel High Level Synthesis compiler allows users to architect new designs for reconfigurable hardware using C/C++. Using the HARPv2 as a vehicle for exploration, we investigate the utility of several of the most notable matrix multiplication optimizations to better understand the performance portability of OpenCL and the implications for such optimizations on this and future heterogeneous architectures. Our results give targeted insights into the applicability of best practices that were for existing architectures when used on emerging heterogeneous systems.</p>"],"dc:identifier":["https://doi.org/10.7936/dbf4-6j92","https://openscholarship.wustl.edu/eng_etds/536"],"dc:language":["English (en)"],"dc:rights":["I have not registered my thesis with the U.S. Copyright Office, and do not intend to."],"dc:subject":["Design space search","HARP","High-level synthesis","OpenCL","SGEMM","Field programmable gate array","heterogeneous architecture","Computer and Systems Architecture","Engineering","Hardware Systems","Numerical Analysis and Scientific Computing","Software Engineering","Systems Architecture","Theory and Algorithms","VLSI and Circuits, Embedded and Hardware Systems"],"dc:title":["Investigating Single Precision Floating General Matrix Multiply in Heterogeneous Hardware"],"thesis:degree_discipline":["Computer Science & Engineering","McKelvey School of Engineering"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["Master of Science (MS)"]},"updated_at":"2026-07-24T06:13:31Z"}