Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 20 of 78 for “"matrix multiplication"”.

  1. FPGA Based Matrix Multiplication Accelerator

    The demand for fast matrix multiplication continues to increase due to recent advances in image processing, graphics processing, digital signal processing, and communication over the wireless network. Therefore, the development of a hardware-based matrix multiplication is essential, which is …

    texas-state Repository record for FPGA Based Matrix Multiplication Accelerator (opens in a new tab)

  2. Matrix multiplication with Asynchronous Logic Automata

    … to achieve performance. In this thesis I use matrix multiplication as a case study to investigate how numerical computation can be performed in this substrate, and how the potential benefits play out in terms of hardware performance estimates. First we take a brief tour of supercomputing …

    mit Repository record for Matrix multiplication with Asynchronous Logic Automata (opens in a new tab)

  3. Second-generation polyalgorithms for parallel dense-matrix multiplication

    … and Anthony Skjellum, includes fourteen dense matrix multiplication algorithms mapped onto two-dimensional process grids using the Message Passing Interface (MPI). This thesis' goal is to achieve optimized performance of parallel, dense linear algebra algorithms by varying the algorithm as a …

    utc Repository record for Second-generation polyalgorithms for parallel dense-matrix multiplication (opens in a new tab)

  4. Quantization Methods for Matrix Multiplication and Efficient Transformers

    … NestQuant — a technique for quantization of matrix products and post-training quantization of LLMs. Beyond reducing the memory footprint, quantization accelerates inference, as the primary bottleneck during autoregressive generation is often the memory bandwidth. NestQuant leverages two …

    mit Repository record for Quantization Methods for Matrix Multiplication and Efficient Transformers (opens in a new tab)

  5. Optimizing Out-Of-Memory Sparse-Dense Matrix Multiplication

    … state-of-the-art approaches for sparse-dense matrix multiplication (SpMDM), with a focused application on graph machine learning workloads, such as graph neural networks (GNNs), though this work is general enough such that it should apply to any application tailored for running matrix

    mit Repository record for Optimizing Out-Of-Memory Sparse-Dense Matrix Multiplication (opens in a new tab)

  6. A Reconfigurable FPGA Overlay Architecture for Matrix-Matrix Multiplication

    … attempts to develop hardware accelerators for matrix-matrix multiplication. Both application-specific integrated circuits (ASICs), and field-programmable arrays (FPGAs) are used for this purpose. However, a trade-off between the two platforms is that ASICs provide little flexibility after they …

    wustl Repository record for A Reconfigurable FPGA Overlay Architecture for Matrix-Matrix Multiplication (opens in a new tab)

  7. High-performance matrix multiplication on Intel and FGPA platforms

    Matrix multiplication is at the core of high-performance numerical computation. Software methods of accelerating matrix multiplication fall into two categories. One is based on calculation simplification. The other one is based on increasing the memory access efficiency. Also matrix multiplication

    njit Repository record for High-performance matrix multiplication on Intel and FGPA platforms (opens in a new tab)

  8. Algebraic geometry for tensor networks, matrix multiplication, and flag matroids

    … tensor networks, and more specifically: uniform matrix product states. We use methods from nonlinear algebra and algebraic geometry to answer questions about topology, defining equations, and identifiability of uniform matrix product states. By an interplay of theorems from algebra, geometry, and …

    qucosa-diss

  9. Tardigrade: A Hardware Accelerator for Sparse Matrix Multiplication and Sparse Convolution

    Sparse matrix-sparse matrix multiplication (SpMSpM) and sparse convolution are critical primitive operations for scientific computing and deep learning. Prior work has proposed accelerators for each of these primitives, but these systems are often specialized to run either SpMSpM or sparse …

    mit Repository record for Tardigrade: A Hardware Accelerator for Sparse Matrix Multiplication and Sparse Convolution (opens in a new tab)

  10. Linear Exact Repair Schemes for Distributed Storage and Secure Distributed Matrix Multiplication

    … of distributed storage and secure distributed matrix multiplication. We develop the (Λ, Γ, W, ⊙)-exact repair scheme framework for discussing both of these contexts and develop a multitude of explicit exact repair schemes utilizing decreasing monomial-Cartesian codes (DMC codes). Specifically, …

    vt Repository record for Linear Exact Repair Schemes for Distributed Storage and Secure Distributed Matrix Multiplication (opens in a new tab)

  11. A hybrid communication pattern and algorithm for distributed sparse-times-dense matrix multiplication

    Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-09-16 without embargo terms

    uiuc Repository record for A hybrid communication pattern and algorithm for distributed sparse-times-dense matrix multiplication (opens in a new tab)

  12. A hybrid communication pattern and algorithm for distributed sparse-times-dense matrix multiplication

    Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-09-16 without embargo terms

    uiuc Repository record for A hybrid communication pattern and algorithm for distributed sparse-times-dense matrix multiplication (opens in a new tab)

  13. Investigating Single Precision Floating General Matrix Multiply in Heterogeneous Hardware

    <p>The fundamental operation of matrix multiplication is ubiquitous across a myriad of disciplines. Yet, the identification of new optimizations for matrix multiplication remains relevant for emerging hardware architectures and heterogeneous systems. Frameworks such as OpenCL enable computation …

    wustl Repository record for Investigating Single Precision Floating General Matrix Multiply in Heterogeneous Hardware (opens in a new tab)

  14. Machine Learning Techniques for Code Generation and Optimization

    … to generate high performance libraries for matrix-matrix multiplication. Our library generator produces matrix multiplication routines that use recursive layouts and several levels of tiling. Our approach is to use a classifier learning system to search in the space of the different ways to …

    uiuc Repository record for Machine Learning Techniques for Code Generation and Optimization (opens in a new tab)

  15. New results in canonical polyadic decomposition overfinite fields

    … rank of the tensor. CPD is at the core of fast matrix multiplication, a computational problem with widespread implications across several seemingly unrelated problems in computer science. Much recent progress in this field has used randomized heuristic search to find new CPDs, often over a …

    mit Repository record for New results in canonical polyadic decomposition overfinite fields (opens in a new tab)

  16. Linear algebraic techniques in algorithms and complexity

    … of different problems. We focus in particular on matrix multiplication algorithms, which have surprisingly fast running times and can hence be used to design fast algorithms in many settings, and matrix rank methods, which can be used to design algorithms or prove lower bounds by analyzing the …

    mit Repository record for Linear algebraic techniques in algorithms and complexity (opens in a new tab)

  17. Inference neural network hardware acceleration techniques

    … Most of these algorithms heavily involve matrix multiplication. As a result, building a neural processing unit (NPU) beside the CPU to accelerate matrix multiplication is a popular approach. The NPU helps reduce the work done by the CPU, and often operates in parallel with the CPU, so in …

    uiuc Repository record for Inference neural network hardware acceleration techniques (opens in a new tab)

  18. Efficient parallel computation on multiprocessors with optical interconnection networks

    … sorting, merging, and selection; Boolean matrix multiplication, transitive closure and their applications to connected component problems. We implement an optimal sorting algorithm on an n-processor LARPBS. With this optimal sorting algorithm at disposal, we study the sorting problem for …

    lsu-thes Repository record for Efficient parallel computation on multiprocessors with optical interconnection networks (opens in a new tab)

  19. Coded Distributed Function Computation

    … data. Of particular interest is the operation of matrix multiplication, a fundamental operation in many big data/machine learning algorithms, which is the main focus of this dissertation. Two serendipitous consequences of the (bi-)linear nature of matrix multiplication is that it is both highly …

    cuny-grad Repository record for Coded Distributed Function Computation (opens in a new tab)

  20. Efficient Algorithms for Graph-Theoretic and Geometric Problems

    … range from circuit design optimization to fast matrix multiplication. First, we study a graph-theoretical model of the so called ''firefighter problem''. The objective is to save as much as possible of an area by appropriately placing firefighters. We provide both new exact algorithms for the …

    lund Repository record for Efficient Algorithms for Graph-Theoretic and Geometric Problems (opens in a new tab)

Page 1 of 4