Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 20 of 933 for “"training data"”.
-
Controlled training data generation with diffusion models
… a text-to-image generative model to produce training data specifically “useful” for supervised learning. Unlike previous works that employ an open-loop approach and pre-define prompts to generate new data using either a language model or human expertise, we develop an automated closed-loop …
-
Improving wordspotting performance with limited training data
Thesis (Ph. D.)--Massachusetts Institute of Technology, Dept. of Electrical Engineering and Computer Science, 1995.
-
Strategic Selection of Training Data for Domain-Specific Speech Recognition
… uses by allowing users to upload their own training data for making custom models that augment Watson's general model. This requires deciding a strategy for picking the training model. This thesis experiments with different training choices for custom language models that augment Watson's …
-
ModelPred: A Framework for Predicting Trained Model from Training Data
… helps to understand the impact of changes in training data on a trained model. This is critical for building trust in various stages of a machine learning pipeline: from cleaning poor-quality samples and tracking important ones to be collected during data preparation, to calibrating …
-
Steering Vision at Scale: From the Model Weights to Training Data
… to suppress target content without modifying training data or requiring model retraining. This approach enhances ethical alignment and enables greater user control in generative systems. We then turn to the complementary problem: incorporating new concepts. We present a few-shot motion …
-
Teach2Learn : gamifying education to gather training data for natural language processing
… assign labels to text samples from an unlabeled data set, thereby teaching superised machine learning algorithms how to interpret new samples. In return, students can learn how that algorithm works by unlocking lessons written by researchers. This aligns the incentives of researchers and learners …
-
Image series prediction via convolutional recurrent neural networks with limited training data
… used to forecast the image series under limited training data. Specifically, we study the problem of using a pine tree's existing appearance images to predict its future appearance images.</p>
-
Improving clinical risk-stratification tools : instance-transfer for selecting relevant training data
… models for medical applications is that the data are often noisy, incomplete, and suffer from high class-imbalance. This problem becomes more severe when the total amount of data relevant to the task of interest is small. We address this problem in the context of risk-stratifying patients …
-
Assessing the Quality of Synthetic Speech when using Enhanced Speech as Training Data
Both speech synthesis and speech enhancement are well researched fields, but their interaction remains under-explored. In particular, the effectiveness of using enhanced speech to train a speech synthesis model is still relatively unknown. This thesis investigates the effects of using enhanced …
-
A Deterministic Approach to Partitioning Neural Network Training Data for the Classification Problem
… be hindered because of failings in the use of training data. This problem can be exacerbated because of small data set size. In this dissertation, we identify and discuss a number of potential problems with typical random partitioning of neural network training data for the classification …
-
Does it have to be trees? : Data-driven dependency parsing with incomplete and noisy training data
We present a novel approach to training data-driven dependency parsers on incomplete annotations. Our parsers are simple modifications of two well-known dependency parsers, the transition-based Malt parser and the graph-based MST parser. While previous work on parsing with incomplete data has …
-
On the Ethics and Linguistic Impacts of Using the Bible as Training Data for Yucatec Maya-to-Spanish Machine Translation
… the Christian Bible, are commonly used as training data for low-resource machine translation (MT) systems because they constitute some of the most extensive and systematically digitized parallel corpora available for many languages. However, this practice raises both linguistic and ethical …
-
Increasing the Precision of Forest Area Estimates through Improved Sampling for Nearest Neighbor Satellite Image Classification
The impacts of training data sample size and sampling method on the accuracy of forest/nonforest classifications of three mosaicked Landsat ETM+ images with the nearest neighbor decision rule were explored. Large training data pools of single pixels were used in simulations to create samples with …
-
Toward automatic model adaptation for structured domains
… not overfit to irrelevant characteristics of the training data. The optimal model is not only a function of the task to which it is applied, but also the amount of training data available. Copious training data can justify a complex model that includes many of the ``true'' domain interaction. But …
-
Comparison of accuracy and efficiency of five digital image classification algorithms
… of land cover features and two types of image data (Landsat MSS and Thematic Mapper) were represented. Classification algorithms were selected from the General Image Processing System (GIPSY) at the Spatial Data Analysis Laboratory at Virginia Polytechnic Institute and State University, …
-
Robust and Fair Machine Learning under Distribution Shift
… learning algorithms, we usually assume that the training data and test data are independently and identically distributed (iid), indicating that the model learned from the training data can be well applied to the test data with good prediction performance. However, this assumption is quite …
-
Machine Learning through the Lens of Data
… debugging model behavior or selecting good training data—require us to relate outputs of models back to the training data. The goal of predictive data attribution, the focus of this thesis, is to precisely characterize the resulting model behavior as a function of the training data in order …
-
Designing Bayesian networks for highly expert-involved problem diagnosis domains
… has been made difficult by limited available training data. The probabilistic diagnostic methods that do not require a substantial amount of available training data usually require considerable expert involvement in design. This thesis proposes a model which balances the amount of expert …
-
TOWARDS RELIABLE AI UNDER DISTRIBUTION SHIFTS: A DATA-CENTRIC PERSPECTIVE
… often rely on spurious correlations in the training data, leading to performance degradation and unreliability when processing inputs under distribution shifts. This thesis systematically studies the robustness to distribution shifts for ML models from a data-centric perspective. First, we …
-
Secure and reliable deep learning in signal processing
… need to manually extract features from raw data that can better describe the underlying problem. Such a process requires strong domain knowledge about the given problems. On the contrary, deep learning-based signal processing algorithms can discover features and patterns that would not be …
Page 1 of 47