Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 20 of 309 for “"large datasets"”.

  1. Scaling data mining activities on very large datasets

    This thesis addresses the issue of enhancing the scalability of data mining techniques, with specific emphasis on association rule and frequent itemset mining. In particular, it proposes a scalable itemset mining approach relying on (i) a persistent (disk-based) representation of the transactional …

    poli-torino Repository record for Scaling data mining activities on very large datasets (opens in a new tab)

  2. Multi-armed bandits and applications to large datasets

    … ABE with other bandit algorithms on real world datasets.

    uiuc Repository record for Multi-armed bandits and applications to large datasets (opens in a new tab)

  3. Extracting and utilizing hidden structures in large datasets

    The hidden structure within datasets --- capturing the inherent structure within the data not explicitly captured or encoded in the data format --- can often be automatically extracted and used to improve various data processing applications. Utilizing such hidden structure enables us to …

    uiuc Repository record for Extracting and utilizing hidden structures in large datasets (opens in a new tab)

  4. Recovering signals in physiological systems with large datasets

    … is an extension of Tikhonov regularization to large-scale problems, using a sequential update approach. In the second method, we improve the conditioning of the problem by assuming that the input is uniform over a known time interval, and then we use a least-squares method to estimate the …

    vt Repository record for Recovering signals in physiological systems with large datasets (opens in a new tab)

  5. Efficient Data Modelling, Indexing and Processing in Large Datasets

    Many devices and applications in social networks and on-line services are producing, storing, and using description, location, and occurrence time of objects. There are various systems to study, model, index, and process a huge amount of data. In this thesis, we study graphs and publish/subscribe …

    unsw Repository record for Efficient Data Modelling, Indexing and Processing in Large Datasets (opens in a new tab)

  6. Adapting ADTrees for Improved Performance on Large Datasets with High Arity Features

    … both space usage and query time, particularly on datasets with very high dimensionality and with high arity features. We propose five modifications to the ADtree, each of which can be used to improve size and query time under specific types of datasets and features. These modifications also …

    byu Repository record for Adapting ADTrees for Improved Performance on Large Datasets with High Arity Features (opens in a new tab)

  7. Random sampling as a clutter reduction technique to facilitate interactive visualisation of large datasets

    … of sizeable data collections. Exploring these large datasets for patterns or trends is a difficult and complex task, especially when users do not always know what they are looking for. Information visualisation can facilitate this task through an interactive visual representation, thus making …

    lancaster Repository record for Random sampling as a clutter reduction technique to facilitate interactive visualisation of large datasets (opens in a new tab)

  8. Smooth Interactive Visualization

    … is a powerful tool for understanding large datasets. However, many commonly-used techniques in information visualization are not C^1 smooth, i.e. when represented as a function, they are either discontinuous or have a discontinuous first derivative. For example, histograms are a …

    vt Repository record for Smooth Interactive Visualization (opens in a new tab)

  9. Intuitive, interactive, and scalable multi-resolution interfaces for accelerating data exploration

    … However, with the availability of increasingly large datasets, data analysts are often faced with two types of scalability challenges—perceptual and interactive scalability. Perceptual scalability stems from the increasing complexity and volume of the underlying data being presented or analyzed …

    uiuc Repository record for Intuitive, interactive, and scalable multi-resolution interfaces for accelerating data exploration (opens in a new tab)

  10. Machine learning ensemble method for discovering knowledge from big data

    … approach in dealing with the problem of mining large datasets because of their accuracy and ability of utilizing the divide-and-conquer mechanism in parallel computing environments. This research proposes a machine learning ensemble framework and implements it in a high performance computing …

    east-anglia Repository record for Machine learning ensemble method for discovering knowledge from big data (opens in a new tab)

  11. Towards unifying spreadsheets with databases for ad-hoc interactive data management at scale

    … interactive ad-hoc management of very large datasets presents a host of challenges, ranging from performance to interface usability. This thesis introduces a new research direction of manipulation of large datasets using an interactive interface and makes several steps towards this …

    uiuc Repository record for Towards unifying spreadsheets with databases for ad-hoc interactive data management at scale (opens in a new tab)

  12. A scalable direct manipulation engine for position-aware presentational data management

    With the explosion of data, large datasets become more common for data analysis. How- ever, existing analytic tools are lack of scalability and large-scale data management tools are lack of interactivity. A lot of data analysis tasks are based on the order of data, we are proposing the very first …

    uiuc Repository record for A scalable direct manipulation engine for position-aware presentational data management (opens in a new tab)

  13. Modeling and Computation of Complex Interventions in Large-scale Epidemiological Simulations using SQL and Distributed Database

    … simulate complex intervention scenarios over large datasets. Indemics is one such interactive data intensive framework for High-performance computing (HPC) based large-scale epidemic simulations. In the Indemics framework, interventions are supplied from an external, standalone database which …

    vt Repository record for Modeling and Computation of Complex Interventions in Large-scale Epidemiological Simulations using SQL and Distributed Database (opens in a new tab)

  14. Variational Approximation for Complex Regression Models

    The trend towards collecting large datasets has resulted in the need for more flexible models and fast computational approximations. My thesis reflects these themes by considering some very flexible regression models and developing fast variational approximation methods for fitting them under a …

    nus Repository record for Variational Approximation for Complex Regression Models (opens in a new tab)

  15. Democratizing Details-on-demand Data Visualizations at Scale

    … and zoom to help the viewer navigate through a large data space. In the past years, we have witnessed an increasing amount of data visualization applications that embrace this paradigm to facilitate data exploration and analysis. Web maps are a clear example. However, due to the highly …

    mit Repository record for Democratizing Details-on-demand Data Visualizations at Scale (opens in a new tab)

  16. Data visualization in the first person

    … how it can be used for navigating and presenting large datasets. Recent years have seen rapid growth in Big Data methodologies throughout scientific research, business analytics, and online services. The datasets used in these areas are not only growing exponentially larger, but also more complex, …

    mit Repository record for Data visualization in the first person (opens in a new tab)

  17. A Fast Minimal Infrequent Itemset Mining Algorithm

    … fast algorithm for finding quasi identifiers in large datasets is presented. Performance measurements on a broad range of datasets demonstrate substantial reductions in run-time relative to the state of the art and the scalability of the algorithm to realistically-sized datasets up to several …

    maynooth Repository record for A Fast Minimal Infrequent Itemset Mining Algorithm (opens in a new tab)

  18. Parallel and distributed MCMC inference using Julia

    … often computationally intensive and operate on large datasets. Being able to eciently learn models on large datasets holds the future of machine learning. As the speed of serial computation stalls, it is necessary to utilize the power of parallel computing in order to better scale with the …

    mit Repository record for Parallel and distributed MCMC inference using Julia (opens in a new tab)

  19. Optical Property Prediction and Molecular Discovery through Multi-Fidelity Deep Learning and Computational Chemistry

    … approaches and statistical modeling. Recently, large datasets of both computed and experimental optical properties have become available, along with the advent of powerful deep learning approaches cable of learning meaningful representations from these large datasets. This thesis presents new …

    mit Repository record for Optical Property Prediction and Molecular Discovery through Multi-Fidelity Deep Learning and Computational Chemistry (opens in a new tab)

Page 1 of 16