Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 20 of 234 for “"data generation"”.

  1. Knowledge-aware data generation

    Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-20 without embargo terms

    uiuc Repository record for Knowledge-aware data generation (opens in a new tab)

  2. BIG-THICK DATA GENERATION VIA LIFE SEQUENCES

    The large-scale personal big data is collected at high speeds by smart devices, mainly in the form of sensor data. While this extensive type of data allows for the analysis and inference about the environment where it was collected (e.g., locations) and certain aspects of human behavior (e.g., …

    trento Repository record for BIG-THICK DATA GENERATION VIA LIFE SEQUENCES (opens in a new tab)

  3. Controlled training data generation with diffusion models

    … generative model to produce training data specifically “useful” for supervised learning. Unlike previous works that employ an open-loop approach and pre-define prompts to generate new data using either a language model or human expertise, we develop an automated closed-loop system …

    texas Repository record for Controlled training data generation with diffusion models (opens in a new tab)

  4. New Approaches to Synthetic Tabular Data Generation

    Synthetic data generation, while already becoming well-known as part of Generative AI (GenAI), has been primarily focused on images, voice, and text, which mostly have homogeneous data formats. This dissertation focuses on the modeling and generation of synthetic tables, which involve a range of …

    vt Repository record for New Approaches to Synthetic Tabular Data Generation (opens in a new tab)

  5. Multiple classifier combination through ensembles and data generation

    This thesis introduces new approaches, namely the DataBoost and DataBoost-IM algorithms, to extend Boosting algorithms' predictive performance. The DataBoost algorithm is designed to assist Boosting algorithms to avoid over-emphasizing hard examples. In the DataBoost algorithm, new synthetic data

    ottawa-retro Repository record for Multiple classifier combination through ensembles and data generation (opens in a new tab)

  6. Differentially Private Synthetic Data Generation for Relational Databases

    Existing differentially private (DP) synthetic data generation mechanisms typically assume a single-source table. In practice, data is often distributed across multiple tables with relationships across tables. This study presents the first-of-its-kind algorithm that can be combined with \emph{any} …

    mit Repository record for Differentially Private Synthetic Data Generation for Relational Databases (opens in a new tab)

  7. Automated blackbox GUI specifications enhancement and test data generation

    … dissertation deals with the problem of automatic generation of relevant test data for parameterized GUI events (i.e., events associated with widgets that accept user inputs such as textboxes and textareas). Current techniques either manipulate the source code of the application under test (AUT) to …

    iastate Repository record for Automated blackbox GUI specifications enhancement and test data generation (opens in a new tab)

  8. SDV : an open source library for synthetic data generation

    … system that can accurately generate synthetic data. The goals of this thesis were to separate the different components in synthetic data generation into their own libraries. We identified these components as consisting of a way to transform the data, a way to model the data, and a way to …

    mit Repository record for SDV : an open source library for synthetic data generation (opens in a new tab)

  9. Privacy-Preserving Synthetic Medical Data Generation with Deep Learning

    … Language Processing. However, the utilization of data-driven methods in healthcare raises privacy concerns, which creates limitations for collaborative research. A remedy to this problem is to generate and employ synthetic data to address privacy concerns. Existing methods for artificial data

    vt Repository record for Privacy-Preserving Synthetic Medical Data Generation with Deep Learning (opens in a new tab)

  10. A Topology-Guided Diffusion Process for Synthetic Tabular Data Generation

    Synthesizing realistic tabular data is crucial for any analytical application, including policy evaluation related to household energy use. However, detailed household-level consumption data, necessary for such evaluation, are scare at fine geographic scales, as public surveys like the U.S. …

    mit Repository record for A Topology-Guided Diffusion Process for Synthetic Tabular Data Generation (opens in a new tab)

  11. Utilizing Recurrent Neural Networks for Temporal Data Generation and Prediction

    The Falling Creek Reservoir (FCR) in Roanoke is monitored for water quality and other key measurements to distribute clean and safe water to the community. Forecasting these measurements is critical for management of the FCR. However, current techniques are limited by inherent Gaussian linearity …

    vt Repository record for Utilizing Recurrent Neural Networks for Temporal Data Generation and Prediction (opens in a new tab)

  12. Automatic Software Test Data Generation from Z Specifications Using Evolutionary Algorithms

    Test data sets have been automatically generated for both numerical and string data types to test the functionality of simple procedures and a good sized UNIX filing system from their Z specifications. Different structured properties of software systems are covered, such as arithmetic expressions, …

    southwales Repository record for Automatic Software Test Data Generation from Z Specifications Using Evolutionary Algorithms (opens in a new tab)

  13. Conversation understanding and realistic artificial crash data generation with deep learning

    … conversation understanding and realistic crash data generation with deep learning. Conversation understanding includes conversational outcome, formality, and politeness prediction. For conversation outcome prediction, we use recorded audio calls collected from a partnering Fortune 500 firm that …

    missouri Repository record for Conversation understanding and realistic artificial crash data generation with deep learning (opens in a new tab)

  14. Design and development of an LLM-based framework for synthetic data generation

    The increasing demand for high-quality datasets in fields such as healthcare, finance, and cybersecurity is hindered by challenges such as data scarcity, privacy concerns, and regulatory restrictions. This thesis introduces a novel framework for generating synthetic data using fine-tuned Large …

    uoit Repository record for Design and development of an LLM-based framework for synthetic data generation (opens in a new tab)

  15. Design and Evaluation of an AI-Driven Pipeline for Synthetic Tabular Data Generation

    The increasing reliance on cloud environments for data-driven applications has created a critical tension between operational efficiency and regulatory compliance. Organisations require high-quality, representative data for effective software testing, but traditional Test Data Management (TDM) …

    stellenbosch Repository record for Design and Evaluation of an AI-Driven Pipeline for Synthetic Tabular Data Generation (opens in a new tab)

  16. Proton Computed Tomography: Matrix Data Generation Through General Purpose Graphics Processing Unit Reconstruction

    … that was developed to generate realistic pCT data sets. Simulated data sets were used to compare the performance of a BIP implementation against a SAP implementation on a single GPGPU with the data stored in a sparse matrix structure called the compressed sparse row (CSR) format. Both BIP and …

    csusb Repository record for Proton Computed Tomography: Matrix Data Generation Through General Purpose Graphics Processing Unit Reconstruction (opens in a new tab)

  17. Modeling exascale data generation and storage for the large hadron collider computing network

    … CMS operations produce 100 Petabytes of physics data per year, which is stored within a globally distributed grid network of 70 scientific institutions. By 2027, upgrades to the LHC and CMS detector will allow unprecedented probes of microscopic physics, but in doing so generate 2,000 Petabytes …

    mit Repository record for Modeling exascale data generation and storage for the large hadron collider computing network (opens in a new tab)

  18. WiSDM: a platform for crowd-sourced data acquisition, analytics, and synthetic data generation

    … a result, it is desirable to collect behavioral data before and during a disease outbreak. Such data can help in creating better computer models that can, in turn, be used by epidemiologists and policy makers to better plan and respond to infectious disease outbreaks. However, traditional data

    vt Repository record for WiSDM: a platform for crowd-sourced data acquisition, analytics, and synthetic data generation (opens in a new tab)

  19. Synthetic Data Generation and Sampling for Online Training of DNN in Manufacturing Supervised Learning Problems

    … of Industrial Internet offers abundant passive data from manufacturing systems and networks, which enables data-driven modeling with high-data-demand, advanced statistical models such as Deep Neural Networks (DNNs). Deep Neural Networks (DNNs) have proven to be remarkably effective in supervised …

    vt Repository record for Synthetic Data Generation and Sampling for Online Training of DNN in Manufacturing Supervised Learning Problems (opens in a new tab)

Page 1 of 12