Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 9 of 9 for “"micro-benchmark"”.

  1. Performance Evaluation of Blocking and Non-Blocking Concurrent Queues on GPUs

    … their performance and behavior using both micro-benchmark and real-world application. We provide a complete evaluation and analysis of our implementations on an AMD Radeon R7 GPU. Our experiment shows that non-blocking approach outperforms blocking approach by up to 15.1 times when …

    mississippi Repository record for Performance Evaluation of Blocking and Non-Blocking Concurrent Queues on GPUs (opens in a new tab)

  2. Integration and Evaluation of Cache Coherence Protocols for Multiprocessor SoCs

    … memory controller. The simulations based on micro-benchmark and RTOS kernel showed the benefits of my methodologies over a generic software solution. This thesis also evaluated and quantified the efficiency of coherence traffic based on a novel emulation platform using FPGA. The proposed …

    gatech Repository record for Integration and Evaluation of Cache Coherence Protocols for Multiprocessor SoCs (opens in a new tab)

  3. Stela: on-demand elasticity in distributed data stream processing systems

    … We conducted experiments on Stela using a set of micro benchmark topologies as well as two topologies from Yahoo! Inc. Our experiment results shows Stela achieves 5% to 120% higher post scale throughput comparing to default Storm scheduler performing scale out operations, and 40% to 500% of …

    uiuc Repository record for Stela: on-demand elasticity in distributed data stream processing systems (opens in a new tab)

  4. Accelerating MPI collective communications through hierarchical algorithms with flexible inter-node communication and imbalance awareness

    … times are common in these applications. A micro-benchmark is used to investigate the nature of process imbalance with perfectly balanced workloads, and understand the nature of inter- versus intra-node imbalance. These insights are then used to develop imbalance tolerant reduction, …

    purdue-thes Repository record for Accelerating MPI collective communications through hierarchical algorithms with flexible inter-node communication and imbalance awareness (opens in a new tab)

  5. High-Performance Network- and GPU-Aware Communication for MPI Partitioned and MPI Neighbourhoods

    … workloads. We addressed the lack of open-source micro-benchmarks for MPI Partitioned communication. We designed a micro-benchmark suite that allows users to search the parameter space for optimal partition communication usage for their application. We provided benchmarks for halo exchanges and …

    queens Repository record for High-Performance Network- and GPU-Aware Communication for MPI Partitioned and MPI Neighbourhoods (opens in a new tab)

  6. Techniques for communication optimization of parallel programs in an adaptive runtime system

    … We extend these techniques through micro-benchmark studies and integration into the production scale Charm++ runtime. We also turn our attention from internode communication optimization to apply these same techniques to intranode communication between various hardware devices, i.e. …

    uiuc Repository record for Techniques for communication optimization of parallel programs in an adaptive runtime system (opens in a new tab)

  7. dPool : a distributed data structure for factored operating systems

    … to achieve scalability across even and uneven micro-benchmark workloads. This thesis shows that common parallel and distributed programming techniques apply to the creation of dPool and that background threads within a dPool can increase performance. Finally, this thesis evaluates different …

    mit Repository record for dPool : a distributed data structure for factored operating systems (opens in a new tab)

  8. Automated performance characterization of applications using hardware monitoring events

    … machine learning mechanism, decision tree on our micro-benchmarks in order to fingerprint the performance problems. Decision trees are trained on a sampled set of hardware events to fingerprint the architectural hardware bottlenecks. Our system divides a profiled application into functions using …

    uiuc Repository record for Automated performance characterization of applications using hardware monitoring events (opens in a new tab)

  9. Scheduling Memory Transactions in Distributed Systems

    … closed nesting) by up to 3.5x and 4.5x on a micro-benchmark (bank) and the TPC-C transactional benchmark, respectively. CTS considers the replicated DTM model: object replicas are distributed across clusters of nodes, where clusters are determined based on inter-node distance, to maximize …

    vt Repository record for Scheduling Memory Transactions in Distributed Systems (opens in a new tab)