Stellenbosch : Stellenbosch University
Design and Evaluation of an AI-Driven Pipeline for Synthetic Tabular Data Generation
Abstract
dc:description.abstractThe increasing reliance on cloud environments for data-driven applications has created a critical tension between operational efficiency and regulatory compliance. Organisations require high-quality, representative data for effective software testing, but traditional Test Data Management (TDM) methods are often resource-intensive and tend to produce low-quality, overly curated test datasets. The use of production data for testing poses significant legal and ethical risks due to the presence of Personally Identifiable Information (PII). Synthetic data – artificially generated data that statistically mimics real-world data without compromising PII – offers a compelling solution. However, generating realistic, high-quality tabular data is a non-trivial task, particularly when the source data is incomplete and messy. This thesis addresses these challenges by developing, implementing, and rigorously evaluating an end-to-end AI-driven pipeline for generating high-quality synthetic tabular data. The pipeline is modular and cloud-native, incorporating robust preprocessing techniques and Missing Value Imputation (MVI) as foundational steps. A formal evaluation framework was developed to assess the quality of synthetic data based on three core dimensions: Fidelity – measuring statistical similarities between synthetic and real data; Utility – measuring performance in downstream machine learning tasks; and Privacy – measuring empirical privacy risks. The experimental results reveal that the quality of the final synthetic data is highly dependent on the initial imputation step, with the Mice Forest algorithm significantly outperforming naive row deletion. In a comparative analysis of generative models, the Tabular Variational Autoencoder (TVAE) emerged as the leading generalist model, achieving the highest fidelity and classification utility. Gaussian Copula, however, demonstrated task-specific excellence in regression tasks where preserving explicit correlation is essential. Importantly, this research provides strong empirical evidence that a trade-off between fidelity and privacy is not inevitable; it is possible to achieve both high fidelity and low privacy risk simultaneously. This thesis establishes a robust foundation for a comprehensive framework designed to enable efficient testing and promote data collaboration through high-quality, privacypreserving synthetic data. It provides a validated mechanism for generating superior test data, mitigating the high costs and compliance risks associated with traditional TDM.
Degree
thesis:*- Grantor dc:publisher
- Stellenbosch : Stellenbosch University
- Year dc:date.issued
- 2026
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Sims-Handcock, Chad Calvin
- Advisor dc:contributor.advisor
-
- Theart, Rensu
Rights
- Language dc:language.iso
- en
Identifiers
dc:identifier.*- Repository record dc:identifier.uri
- https://scholar.sun.ac.za/handle/10019.1/135845
- OAI identifier oai:identifier
- oai:scholar.sun.ac.za:10019.1/135845