Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 20 of 309 for “"Large datasets"”.
-
Scaling data mining activities on very large datasets
This thesis addresses the issue of enhancing the scalability of data mining techniques, with specific emphasis on association rule and frequent itemset mining. In particular, it proposes a scalable itemset mining approach relying on (i) a persistent (disk-based) representation of the transactional …
-
Multi-armed bandits and applications to large datasets
… ABE with other bandit algorithms on real world datasets.
-
Extracting and utilizing hidden structures in large datasets
The hidden structure within datasets --- capturing the inherent structure within the data not explicitly captured or encoded in the data format --- can often be automatically extracted and used to improve various data processing applications. Utilizing such hidden structure enables us to …
-
Recovering signals in physiological systems with large datasets
… is an extension of Tikhonov regularization to large-scale problems, using a sequential update approach. In the second method, we improve the conditioning of the problem by assuming that the input is uniform over a known time interval, and then we use a least-squares method to estimate the …
-
Efficient Data Modelling, Indexing and Processing in Large Datasets
Many devices and applications in social networks and on-line services are producing, storing, and using description, location, and occurrence time of objects. There are various systems to study, model, index, and process a huge amount of data. In this thesis, we study graphs and publish/subscribe …
-
Adapting ADTrees for Improved Performance on Large Datasets with High Arity Features
… both space usage and query time, particularly on datasets with very high dimensionality and with high arity features. We propose five modifications to the ADtree, each of which can be used to improve size and query time under specific types of datasets and features. These modifications also …
-
Random sampling as a clutter reduction technique to facilitate interactive visualisation of large datasets
… of sizeable data collections. Exploring these large datasets for patterns or trends is a difficult and complex task, especially when users do not always know what they are looking for. Information visualisation can facilitate this task through an interactive visual representation, thus making …
-
Investigating interactive visualisation in a Cloud computing environment : a dissertation submitted in partial fulfilment of the requirements for the degree of Bachelor of Software and Information Technology with Honours at Lincoln University
Scientists working with large datasets without a desktop with advanced capacity may not be able to visualise the simulation output efficiently. This is because visualisation of large datasets is computationally intensive in terms of filtering, mapping, and rendering the datasets. The time taken to …
-
Smooth Interactive Visualization
… is a powerful tool for understanding large datasets. However, many commonly-used techniques in information visualization are not C^1 smooth, i.e. when represented as a function, they are either discontinuous or have a discontinuous first derivative. For example, histograms are a …
-
Intuitive, interactive, and scalable multi-resolution interfaces for accelerating data exploration
… However, with the availability of increasingly large datasets, data analysts are often faced with two types of scalability challenges—perceptual and interactive scalability. Perceptual scalability stems from the increasing complexity and volume of the underlying data being presented or analyzed …
-
Machine learning ensemble method for discovering knowledge from big data
… approach in dealing with the problem of mining large datasets because of their accuracy and ability of utilizing the divide-and-conquer mechanism in parallel computing environments. This research proposes a machine learning ensemble framework and implements it in a high performance computing …
-
Towards unifying spreadsheets with databases for ad-hoc interactive data management at scale
… interactive ad-hoc management of very large datasets presents a host of challenges, ranging from performance to interface usability. This thesis introduces a new research direction of manipulation of large datasets using an interactive interface and makes several steps towards this …
-
A scalable direct manipulation engine for position-aware presentational data management
With the explosion of data, large datasets become more common for data analysis. How- ever, existing analytic tools are lack of scalability and large-scale data management tools are lack of interactivity. A lot of data analysis tasks are based on the order of data, we are proposing the very first …
-
Modeling and Computation of Complex Interventions in Large-scale Epidemiological Simulations using SQL and Distributed Database
… simulate complex intervention scenarios over large datasets. Indemics is one such interactive data intensive framework for High-performance computing (HPC) based large-scale epidemic simulations. In the Indemics framework, interventions are supplied from an external, standalone database which …
-
Variational Approximation for Complex Regression Models
The trend towards collecting large datasets has resulted in the need for more flexible models and fast computational approximations. My thesis reflects these themes by considering some very flexible regression models and developing fast variational approximation methods for fitting them under a …
-
Democratizing Details-on-demand Data Visualizations at Scale
… and zoom to help the viewer navigate through a large data space. In the past years, we have witnessed an increasing amount of data visualization applications that embrace this paradigm to facilitate data exploration and analysis. Web maps are a clear example. However, due to the highly …
-
Data visualization in the first person
… how it can be used for navigating and presenting large datasets. Recent years have seen rapid growth in Big Data methodologies throughout scientific research, business analytics, and online services. The datasets used in these areas are not only growing exponentially larger, but also more complex, …
-
A Fast Minimal Infrequent Itemset Mining Algorithm
… fast algorithm for finding quasi identifiers in large datasets is presented. Performance measurements on a broad range of datasets demonstrate substantial reductions in run-time relative to the state of the art and the scalability of the algorithm to realistically-sized datasets up to several …
-
Parallel and distributed MCMC inference using Julia
… often computationally intensive and operate on large datasets. Being able to eciently learn models on large datasets holds the future of machine learning. As the speed of serial computation stalls, it is necessary to utilize the power of parallel computing in order to better scale with the …
-
Optical Property Prediction and Molecular Discovery through Multi-Fidelity Deep Learning and Computational Chemistry
… approaches and statistical modeling. Recently, large datasets of both computed and experimental optical properties have become available, along with the advent of powerful deep learning approaches cable of learning meaningful representations from these large datasets. This thesis presents new …
Page 1 of 16