Abstract
dc:description.abstractScience relies on data. Real data can be severely limiting: it may be privacy-sensitive, unfair, unbalanced, unrepresentative, or it may simply not exist for the setting of interest. These limitations continuously constrain how the machine learning (ML) community operates---from training to testing, from model development to deployment. Synthetic data is a type of data that is generated as to resemble (some part of) the real data, while overcoming some of these real data limitations and improving AI trustworthiness. Advances in deep generative modelling have made synthetic data more real than ever, and as a result there is a steep rise in research that aims to replace real data with synthetic data. The first uses of synthetic data were mostly privacy-focused, aiming to create realistic synthetic data that mimics the real data but does not disclose sensitive information. More recently, however, there has been an increasing interest in extending synthetic data to use cases where it improves upon real data, for example providing better fairness, augmenting the dataset size, and creating or simulating data for unseen domains. The latest and most widely acclaimed development is user-prompted data, epitomized by OpenAI's ChatGPT. In this thesis, I first aim to categorize different applications and challenges of synthetic data. Subsequently, I delve deeper into three of these applications, providing synthetic data methods for fairness, augmenting dataset size for small subgroups, and simulating data. Next I turn to the core critique of synthetic data---to what extent can we trust the synthetic data? I include two chapters on this topic. The first explores how we should publish and use synthetic data, such that we can understand whether the downstream results (e.g. predictions of a downstream model) are indeed reliable. The second explores how we can test whether synthetic data is indeed private. At last, we include an outlook on promising synthetic data research directions.
Degree
thesis:*- Name dc:type.qualificationname
- Doctor of Philosophy (PhD)
- Level dc:type.qualificationlevel
- Doctoral
- Grantor dc:publisher.institution
- University of Cambridge
- Year dc:date.issued
- 2024
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Van Breugel, Boris
- Advisor dc:contributor.advisor
-
- van der Schaar, Mihaela
Subjects
dc:subject × 6Rights
dc:rights- Licence
- Language dc:language
- eng
Identifiers
dc:identifier.*- DOI dc:identifier.doi
- https://doi.org/10.17863/CAM.121707
- OAI identifier oai:identifier
- oai:www.repository.cam.ac.uk:1810/390004