Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 20 of 63 for “"Data Preprocessing"”.
-
Revolutionizing Time Series Data Preprocessing with a Novel Cycling Layer in Self-Attention Mechanisms
… a novel method for improving time series data preprocessing by incorporating a cycling layer into self-attention mechanisms. Traditional techniques often struggle to capture the cyclical nature of time series data, impacting predictive model accuracy. By integrating a cycling layer, this …
-
Improving the Prediction Accuracy of Text Data and Attribute Data Mining with Data Preprocessing
<p>Data Mining is the extraction of valuable information from the patterns of data and turning it into useful knowledge. Data preprocessing is an important step in the data mining process. The quality of the data affects the result and accuracy of the data mining results. Hence, Data preprocessing …
-
Hydraulic Data Preprocessing for Anomaly Based Intrusion Detection on SCADA Level of Water Treatment Systems
… attack in the system. The ability to uncover data patterns and gather knowledge from data is a significant benefit of machine learning (ML), however factors such as noise, missing values, excessive features, and inconsistent and redundant data negatively affects the performance of the model, …
-
Automatic Modulation Classification Using Grey Relational Analysis
… of multiple signal features. An evaluation of data preprocessing methods is conducted and performance of the classifier is investigated with the addition of each new signal feature used for classification.
-
Spatio-temporal comparative analysis of scooter share in Washington D.C.
Geospatial-temporal data for different e-scooter firms was collected and investigated for differences in e-scooter usage patterns among customers of the firms. Computational analysis using predictive algorithms and correlation analysis was done to find co-relationally important features for …
-
Hybrid Methods for Feature Selection
<p>Feature selection is one of the important data preprocessing steps in data mining. The feature selection problem involves finding a feature subset such that a classification model built only with this subset would have better predictive accuracy than model built with a complete set of features. …
-
Variable Selection for DNA Methylation Data Using Model-Based Clustering
… simulations and applied to DNA methylation data generated from the Isle of Wight birth cohort study. The approach featured the use of clustering with a penalty function to select informative variables. We evaluated the method by conducting simulations for a variety of scenarios. For …
-
DEVELOPMENT OF COMPUTATIONAL METHODS FOR MASS SPECTROMETRY-BASED UNTARGETED METABOLOMICS DATA ANALYSIS
… there are still limitations in the current data processing pipelines. In this thesis, we first addressed the limitations in MS1 based analysis and developed a software package “MetTailor” containing two novel post-alignment data preprocessing functions: 1) re-align the potential misaligned …
-
Uma abordagem para geração e visualização de regras de associação de acesso a conteúdos de portal de notícias
… rules obtained from the content access history data of a Brazilian journal. The proposed approach is composed of four phases: exploratory data analysis (EDA), data preprocessing, generation of association and sequence rules, and visualization of results. The algorithms Apriori and FP-Growth were …
-
An analytics approach to hypertension treatment
… to existing electronic health record data to (1) find conjectures parallel and potentially orthogonal to guidelines, (2) hasten response time to therapy, and/or (3) optimize therapy selection. This thesis presents work toward these goals including data preprocessing and exploration, …
-
Real-time aerial vehicle detection and tracking with depth-aided vision sensing
… learning approach by utilizing depth data from depth vision sensor to achieve much faster detection speed while maintain high detection accuracy. We revised some of algorithms presented in part-based representation method to get marginally better performance. Then we invented a novel …
-
An end-to-end online quality prediction system for ultrasonic metal welding based on deep learning
… tool conditions), and not involving tedious data preprocessing and feature engineering. The effectiveness of the proposed method is shown using real-world data generated from a UMW process. A comparative case study is presented to compare three data fusion strategies (early fusion, middle …
-
Credit card fraud detection using machine learning with integration of contextual knowledge
… multi-perspective approach allows automated data pre-processing to model time correlations to complement and eventually replace transaction aggregation strategies to improve detection efficiency. Experiments carried out on a large set of credit card transaction data from the real world (46 …
-
Accelerating human-in-the-loop machine learning
… succinct syntax defining unified processes for data preprocessing, model specification, and learning. We demonstrate that the reuse problem can be cast as a Max-Flow problem, while the caching problem is NP-Hard. We develop effective lightweight heuristics for the latter. Empirical evaluation …
-
Scalable Model for Reaction Outcome Prediction and One-step Retrosynthesis with a Graph-to-Sequence Architecture
… organic chemistry for which a variety of data-driven approaches have emerged. Natural language approaches that model each problem as a SMILESto-SMILES translation lead to a simple end-to-end formulation, reduce the need for data preprocessing, and enable the use of well-optimized machine …
-
Health-AIM: An artificial intelligence approach for inference with clinical health datasets
… artificial intelligence (AI) models need to use data to make reliable decisions regarding patient trajectories and treatments. Incorrect decisions can lead to strain on both caregiver and patient, and clinical datasets often have low volume, high dimensionality, and many missing values. This …
-
Fast Partitioning for Distributed Graph Learning using Multi-level Label Propagation
… techniques to perform inference on unstructured data. However, when graphs become too large, partitioning becomes necessary to allow for distributed computation. Standard graph partitioning methods for GNNsinclude Random partitioning and the state-of-the-art METIS. Whereas METIS produces …
-
Democratizing data science through interactive curation of ML pipelines
… are key to extract actionable insights out of data, yet such skills rarely coexist together. In Machine Learning, high-quality results are only attainable via mindful data preprocessing, hyperparameter tuning and model selection. Domain experts are often overwhelmed by such complexity, de-facto …
-
Optimizing Data Compression via Data Reordering Strategies
… and cost-effectiveness of handling large tabular datasets stored in databases, a range of data compression techniques are employed. Among these, dictionary-based compression methods such as Lz4, Gzip, and Zstandard are commonly utilized to decrease data size. However, while these traditional …
-
Deep Learning Approach for Cell Nuclear Pore Detection and Quantification over High Resolution 3D Data
… nuclear pores in high-resolution 3D microscopy data is critical for cellular biology and disease research. This thesis introduces a deep learning pipeline crafted to automate the segmentation and quantification of nuclear pores from high-resolution 3D cell organelle images. Our aim is to refine …
Page 1 of 4