Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 20 of 122 for “"training dataset"”.
-
Influence of training dataset selection on the performance of a machine learning model
… images based on the learning from a given set of training images, called ‘ground-truths’. This work proposes to compose a good training dataset that would give good accuracy with a robust object detection model by using different training and testing combinations. Various evaluation techniques …
-
Data redundancy reduction using sensitivity analysis method for machine-learning-based battery management system
… and may be excluded in the formation of the training dataset. The newly discovered finding is applied to the ML-based BMS. In the development of the training dataset, a reduced-sized dataset is formed by excluding current, discharge time, and power from the training dataset for the real-time …
-
Defining Safe Training Datasets for Machine Learning Models Using Ontologies
… the safety and completeness of the model’s training dataset. Since most of the complexity of the model is built through training, ensuring the safety of the training dataset could help to increase the trust in the safety of the model. The method proposed in this research uses a domain …
-
A Transformer-Based Foundation Model for Human Microbiome Analysis
… samples in 10-fold cross-validation on the training dataset with an accuracy of 83.7%. On an external validation dataset of 927 samples, our model had an accuracy of 74.9%. Notably, our model performed even better at differentiating diseases from one another. On the diseased samples in the …
-
Inferring properties of neural networks with intelligent designs
… these networks are functions derived from their training data and thorough analysis of these networks reveals information about the training dataset. This could be dire in many scenarios such as network log anomaly classifiers leaking data about the network they were trained on, disease detectors …
-
Spectrum Awareness: Deep Learning and Isolation Forest Approaches for Open-set Identification of Signals
… within the spectrum environment are within the training dataset. To account for this, we have proposed a novel classifier design for detection of unknown signals outside of the training dataset. This two-classifier system forms an open-set recognition (OSR) system that is used to provide more …
-
Implementation of gaussian process models for non-linear system identification
… effective in identifying models from sparse datasets. Therefore, the GP model has been proposed for the identification of models in off-equilibrium regions of operating space, where more established methods might struggle due to a lack of data. The majority of the existing research into …
-
Three Essays in Econometrics
… Finally, these estimators are applied to a job training dataset.
-
Robust Domain Adaptation Using Active Learning
Traditional machine learning algorithms assume training and test datasets are generated from the same underlying distribution, which is not true for most real-world datasets. As a result, a model trained on the training dataset fails to produce good classification accuracy on the test dataset. One …
-
A Data-Based Perspective on Model Reliability
… corrupted, or underrepresented during training. In such settings, the set of features that a model relies on, or its feature prior, often determines the model’s ultimate reliability. While many factors contribute to a model’s feature prior, recent evidence indicates that the training …
-
A system for storage and analysis of machine learning operations
… have low time and space overhead for large training dataset sizes and the overhead is independent of the dataset size.
-
Apklausų dalyvių aktyvumo analizė, pritaikant įvairius binarinio klasifikavimo algoritmus /
… classifier, with oversampling applied to the training dataset. The best model achieved a sensitivity (recall) metric of 82 % and a specificity of 87 %. The worst results were obtained using logistic regression, with a sensitivity metric of 78 % and a specificity of 75 %.
-
Application of Machine Learning Techniques for Real-time Classification of Sensor Array Data
… six machine learning methods to classify a dataset collected using a chemical sensor array: K-Nearest Neighbor (KNN), Support Vector Machine (SVM), Classification and Regression Trees (CART), Random Forest (RF), Naïve Bayes Classifier (NB), and Principal Component Regression (PCR). A total …
-
Test2Vec: An Execution Trace Embedding for Behaviour Coverage Analysis of Test Cases
… traces including the input values as the training dataset and evaluated our approach on 93 real faulty versions of 4 open source Java projects. Our results show that Test2Vec can rank the failing test in the top 5, 10, and 20 tests, on average, 20.43%, 32.26%, and 44.09% of the times, …
-
The Effect of Dataset Size on the Performance of Classification Algorithms for Credit Scoring
… and the generation and collection of massive datasets, there has been a great deal of research benchmarking the performance of these powerful machine learning algorithms against traditional techniques used in credit scoring. This dissertation extends the research into the benchmarking of …
-
Data Driven Models for Language Evolution
… stable, regardless of the variation of the training dataset dimension. When applied to phylogenetic inference of the Indo-European language family, whose higher structure does not yet have consensus, our method has estimated phylogenies which are compatible with the benchmark tree and has …
-
Integrating Multi-Source Weather Data for Deep Learning
… The project subsequently shifts to generating training datasets along with annotations to be ingested by Mask R-CNN [6] network architecture. Finally, it passes the generated training dataset as an input for Detectron [5] software application and attempts to train network for the given 2017 and …
-
Data Augmentation with Seq2Seq Models
… sparsity is an issue that complicates the training process of question answering systems: syntactically diverse but semantically equivalent sentences can have significant disparities in predicted output probabilities. We propose a method for generating an augmented paraphrase corpus for the …
-
Towards answering unanswerable questions: data augmentation for enhanced medical domain question answering
… the effectiveness of fine tuning the model on a training dataset that includes these out-of-schema questions and their corresponding schema. Secondly, this research examines how external medical knowledge sources, related to diagnoses and medication, can be used in data augmentation (either …
-
Protein Fold Recognition Using Adaboost Learning Strategy
… tasks: (i) carry out cross validation within the training dataset, and (ii) test on unseen validation dataset, in which 90% of the proteins have less than 25% sequence identity in training samples. Our result yields 64.7% successful rate in classifying independent validation dataset into 27 types …
Page 1 of 7