Back to results

University of Cambridge

Generating Trustworthy Synthetic Data

Abstract

dc:description.abstract

Science relies on data. Real data can be severely limiting: it may be privacy-sensitive, unfair, unbalanced, unrepresentative, or it may simply not exist for the setting of interest. These limitations continuously constrain how the machine learning (ML) community operates---from training to testing, from model development to deployment. Synthetic data is a type of data that is generated as to resemble (some part of) the real data, while overcoming some of these real data limitations and improving AI trustworthiness. Advances in deep generative modelling have made synthetic data more real than ever, and as a result there is a steep rise in research that aims to replace real data with synthetic data. The first uses of synthetic data were mostly privacy-focused, aiming to create realistic synthetic data that mimics the real data but does not disclose sensitive information. More recently, however, there has been an increasing interest in extending synthetic data to use cases where it improves upon real data, for example providing better fairness, augmenting the dataset size, and creating or simulating data for unseen domains. The latest and most widely acclaimed development is user-prompted data, epitomized by OpenAI's ChatGPT. In this thesis, I first aim to categorize different applications and challenges of synthetic data. Subsequently, I delve deeper into three of these applications, providing synthetic data methods for fairness, augmenting dataset size for small subgroups, and simulating data. Next I turn to the core critique of synthetic data---to what extent can we trust the synthetic data? I include two chapters on this topic. The first explores how we should publish and use synthetic data, such that we can understand whether the downstream results (e.g. predictions of a downstream model) are indeed reliable. The second explores how we can test whether synthetic data is indeed private. At last, we include an outlook on promising synthetic data research directions.

Degree

thesis:*
Name dc:type.qualificationname
Doctor of Philosophy (PhD)
Level dc:type.qualificationlevel
Doctoral
Grantor dc:publisher.institution
University of Cambridge
Year dc:date.issued
2024

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Van Breugel, Boris
Advisor dc:contributor.advisor
  • van der Schaar, Mihaela

Subjects

dc:subject × 6

Rights

dc:rights
Language dc:language
eng

Identifiers

dc:identifier.*
DOI dc:identifier.doi
https://doi.org/10.17863/CAM.121707
OAI identifier oai:identifier
oai:www.repository.cam.ac.uk:1810/390004

Chain of custody

source
Harvested from
Cambridge University
Base URL
api.repository.cam.ac.uk/server/oai/request
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Van Breugel, Boris. Generating Trustworthy Synthetic Data. Doctoral thesis, University of Cambridge, 2024. https://doi.org/10.17863/CAM.121707