Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 20 of 35 for “"Data partitioning"”.

  1. Automatic Data Partitioning on Distributed Memory Multicomputers

    … are determined largely by the manner in which data is partitioned across different processors of the machine. Most of the compilers provide no assistance to the programmer in the crucial task of determining a good data partitioning scheme.

    uiuc Repository record for Automatic Data Partitioning on Distributed Memory Multicomputers (opens in a new tab)

  2. Robust data partitioning for ad-hoc query processing

    Data partitioning can significantly improve query performance in distributed database systems. Most proposed data partitioning techniques choose the partitioning based on a particular expected query workload or use a simple upfront scheme, such as uniform range partitioning or hash partitioning on …

    mit Repository record for Robust data partitioning for ad-hoc query processing (opens in a new tab)

  3. Directive-Based Data Partitioning and Pipelining and Auto-Tuning for High-Performance GPU Computing

    … generally have their own discrete memory space, data needs to be copied from the CPU host memory to the accelerator (device) memory before computation starts on the accelerator. Moreover, programming models like CUDA, OpenMP, OpenACC, and OpenCL can efficiently offload compute-intensive workloads …

    vt Repository record for Directive-Based Data Partitioning and Pipelining and Auto-Tuning for High-Performance GPU Computing (opens in a new tab)

  4. Development of new data partitioning and allocation algorithms for query optimization of distributed data warehouse systems

    Distributed databases and in particular distributed data warehousing are becoming an increasingly important technology for information integration and data analysis. Data Warehouse (DW) systems are used by decision makers for performance measurement and decision support. However, although data

    london-metro Repository record for Development of new data partitioning and allocation algorithms for query optimization of distributed data warehouse systems (opens in a new tab)

  5. An adaptive partitioning scheme for ad-hoc and time-varying database analytics

    Data partitioning significantly improves query performance in distributed database systems. A large number of techniques have been proposed to efficiently partition a dataset, often focusing on finding the best partitioning for a particular query workload. However, many modern analytic applications …

    mit Repository record for An adaptive partitioning scheme for ad-hoc and time-varying database analytics (opens in a new tab)

  6. A Reconfigurable, Distributed-Memory Accelerator for Sparse Applications

    … Quartz, a new architecture that uses short dataflow tasks and reconfigurable compute in a distributed-SRAM system to deliver both high performance and high programmability. Unlike traditional sparse CGRAs or on-die reconfigurable engines, Quartz allows reconfigurable compute to be highly …

    mit Repository record for A Reconfigurable, Distributed-Memory Accelerator for Sparse Applications (opens in a new tab)

  7. A distributed workload-aware approach to partitioning geospatial big data for cybergis analytics

    … contributed to tremendous growth of geospatial data during the past several decades. To resolve the volume and velocity of such big data, distributed system approaches have been extensively studied to partition data for scalable analytics and associated applications. However, previous work on …

    uiuc Repository record for A distributed workload-aware approach to partitioning geospatial big data for cybergis analytics (opens in a new tab)

  8. A statistical approach towards performance analysis of multimodal biometrics systems

    … 126 experiments are conducted with the BSSR1 dataset. The proposed approach helps to examine the performance of typical fusion methods that use different normalization and data partitioning techniques. Experiment results demonstrate that the Simple Sum fusion method working with the Min-Max …

    windsor Repository record for A statistical approach towards performance analysis of multimodal biometrics systems (opens in a new tab)

  9. Modifications To The Fuzzy-ARTMAP Algorithm For Distributed Learning In Large Data Sets

    … of thousands. In this dissertation we apply data partitioning and network partitioning to the FAM algorithm in a sequential and parallel setting to achieve better convergence time and to efficiently train with large databases (hundreds of thousands of patterns). We implement our …

    ucf

  10. AdaptDB : adaptive partitioning for distributed joins

    Big data analytics often involves complex join queries over two or more tables. Such join processing is expensive in a distributed setting both because large amounts of data must be read from disk, and because of data shuffling across the network. Many techniques based on data partitioning have …

    mit Repository record for AdaptDB : adaptive partitioning for distributed joins (opens in a new tab)

  11. Database partitioning strategies for social network data

    … prototyped and benchmarked two different data partitioning strategies for social network type workloads. The first strategy takes advantage of the heavy-tailed degree distributions of social networks to optimize the latency of vertex neighborhood queries. The second strategy takes …

    mit Repository record for Database partitioning strategies for social network data (opens in a new tab)

  12. MPI-based scalable computing platform for parallel numerical application

    … involves a variety of challenges in dealing with data partitioning, workload balancing, data dependencies, and synchronization. Many numerical applications share the need for an underlying parallel framework for parallelization on multi-core/multi-machine hardware. In this thesis, a computing …

    mit Repository record for MPI-based scalable computing platform for parallel numerical application (opens in a new tab)

  13. Improving the Programmability of A Distributed Hardware Accelerator

    … C++ code for (Ōmeteōtl. Lapis abstracts away data partitioning and task orchestration, reducing implementation complexity: for example, it lowers lines of code by 30× for conjugate gradients and 46× for power iteration. Despite this abstraction, generated code achieves 75.7% to 92.6% of the …

    mit Repository record for Improving the Programmability of A Distributed Hardware Accelerator (opens in a new tab)

  14. Transforming and Optimizing Irregular Applications for Parallel Architectures

    … many approaches have been proposed to improve data locality and reduce irregularities through computational and data transformations. However, there are two major drawbacks in these existing approaches that prevent them from achieving optimal performance. First, these approaches use local …

    vt Repository record for Transforming and Optimizing Irregular Applications for Parallel Architectures (opens in a new tab)

  15. ParaView: Performance debugging through visualization of shared data

    … correct performance problems relating to poor data partitioning, false sharing, contention for shared data constructs, and unnecessary synchronization.

    rice Repository record for ParaView: Performance debugging through visualization of shared data (opens in a new tab)

  16. Parallel load and query processing in a distributed array database

    … collect large amounts of multi-dimensional data in their day to day work. They require high performance, scalable systems to manage and process their data. Oftentimes, the underlying distribution of these types of data is skewed and sparse, rather than dense and uniform. As input data sizes …

    mit Repository record for Parallel load and query processing in a distributed array database (opens in a new tab)

  17. Beyond music sharing: an evaluation of peer-to-peer data dissemination techniques in large scientific collaborations

    The avalanche of data from scientific instruments and the ensuing interest from geographically distributed users to analyze and interpret it accentuates the need for efficient data dissemination. An optimal data distribution scheme will find the delicate balance between conflicting requirements of …

    ubc Repository record for Beyond music sharing: an evaluation of peer-to-peer data dissemination techniques in large scientific collaborations (opens in a new tab)

  18. Optimal round and sample-size complexity for partitioning in parallel sorting

    Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2024-05-01

    uiuc Repository record for Optimal round and sample-size complexity for partitioning in parallel sorting (opens in a new tab)

  19. Model of PAH and PCB bioaccumulation in Mya arenaria and application for site assessment in conjunction with sediment quality screening criteria

    … should use a combination of the Equilibrium Partitioning modeling approach and the Threshold Effects Level/Probable Effects Level correlative approach to SQC derivation. Criteria should then be applied as screening values for evaluation of sediment toxicity. One significant component of …

    woods-hole Repository record for Model of PAH and PCB bioaccumulation in Mya arenaria and application for site assessment in conjunction with sediment quality screening criteria (opens in a new tab)

  20. Harnessing the power of intersection for data disaggregation: a novel similarity measure and unsupervised data-driven classification method applied to financial contagion

    … unsupervised classification methods to economic data partitioning, a nonparametric, deterministic, and robust data-driven partitional hard cluster analysis, derived from a novel similarity measure, is introduced. The main objective of the proposed method in the present study is to support …

    cambridge Repository record for Harnessing the power of intersection for data disaggregation: a novel similarity measure and unsupervised data-driven classification method applied to financial contagion (opens in a new tab)

Page 1 of 2