Back to results

NJIT

High-performance matrix multiplication on Intel and FGPA platforms

Abstract

dc:description.abstract

Matrix multiplication is at the core of high-performance numerical computation. Software methods of accelerating matrix multiplication fall into two categories. One is based on calculation simplification. The other one is based on increasing the memory access efficiency. Also matrix multiplication can be accelerated using vector processors. In this investigation, various matrix multiplication algorithms and the vector-based hardware acceleration method are analyzed and compared in terms of performance and memory requirements. Results are shown for Intel and Xilinx FPGA platforms. They show that when the CPU is fast, Goto's algorithm runs faster than Strassen's algorithm because the data access speed is the bottleneck in this case. On the contrary, when the CPU is slow, Strassen's algorithm runs faster because the computation complexity becomes the key factor in this case. Also, the results show that SIMD platforms, such as Intel Xeon and SIMD extensions and an in-house developed VP (Vector co-Processor), for an FPGA, can accelerate matrix multiplication substantially. It is even shown that the VP runs faster than MKL (Intel's optimized Math Kernel Library). This is because not only can the VP take advantage of larger vector lengths but it also minimizes inherent hardware overheads.

Degree

thesis:*
Name thesis:degree_name
Master of Science in Computer Engineering - (M.S.)
Discipline thesis:degree_discipline
Electrical and Computer Engineering
Year
2012

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Li, Gang
Contributors dc:contributor
  • Sotirios Ziavras
  • Roberto Rojas-Cessa
  • Edwin Hou

Subjects

dc:subject × 3

Identifiers

dc:identifier.*
Repository record dc:identifier
https://digitalcommons.njit.edu/theses/136
OAI identifier oai:identifier
oai:digitalcommons.njit.edu:theses-1135

Chain of custody

source
Harvested from
NJIT
Base URL
digitalcommons.njit.edu/do/oai/
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Li, Gang. High-performance matrix multiplication on Intel and FGPA platforms. 2012. https://digitalcommons.njit.edu/theses/136