{"id":{"repo_id":"cambridge","oai_identifier":"oai:www.repository.cam.ac.uk:1810/390004"},"canonical_url":"https://search.dev.ndltd.org/etd/cambridge/oai:www.repository.cam.ac.uk:1810/390004","repository":{"repo_id":"cambridge","name":"Cambridge University","base_url":"https://api.repository.cam.ac.uk/server/oai/request"},"display":{"title":"Generating Trustworthy Synthetic Data","abstract":"Science relies on data. Real data can be severely limiting: it may be privacy-sensitive, unfair, unbalanced, unrepresentative, or it may simply not exist for the setting of interest. These limitations continuously constrain how the machine learning (ML) community operates---from training to testing, from model development to deployment. Synthetic data is a type of data that is generated as to resemble (some part of) the real data, while overcoming some of these real data limitations and improving AI trustworthiness. Advances in deep generative modelling have made synthetic data more real than ever, and as a result there is a steep rise in research that aims to replace real data with synthetic data. The first uses of synthetic data were mostly privacy-focused, aiming to create realistic synthetic data that mimics the real data but does not disclose sensitive information. More recently, however, there has been an increasing interest in extending synthetic data to use cases where it improves upon real data, for example providing better fairness, augmenting the dataset size, and creating or simulating data for unseen domains. The latest and most widely acclaimed development is user-prompted data, epitomized by OpenAI's ChatGPT. In this thesis, I first aim to categorize different applications and challenges of synthetic data. Subsequently, I delve deeper into three of these applications, providing synthetic data methods for fairness, augmenting dataset size for small subgroups, and simulating data. Next I turn to the core critique of synthetic data---to what extent can we trust the synthetic data? I include two chapters on this topic. The first explores how we should publish and use synthetic data, such that we can understand whether the downstream results (e.g. predictions of a downstream model) are indeed reliable. The second explores how we can test whether synthetic data is indeed private. At last, we include an outlook on promising synthetic data research directions.","abstract_html":"Science relies on data. Real data can be severely limiting: it may be privacy-sensitive, unfair, unbalanced, unrepresentative, or it may simply not exist for the setting of interest. These limitations continuously constrain how the machine learning (ML) community operates---from training to testing, from model development to deployment. Synthetic data is a type of data that is generated as to resemble (some part of) the real data, while overcoming some of these real data limitations and improving AI trustworthiness. Advances in deep generative modelling have made synthetic data more real than ever, and as a result there is a steep rise in research that aims to replace real data with synthetic data. The first uses of synthetic data were mostly privacy-focused, aiming to create realistic synthetic data that mimics the real data but does not disclose sensitive information. More recently, however, there has been an increasing interest in extending synthetic data to use cases where it improves upon real data, for example providing better fairness, augmenting the dataset size, and creating or simulating data for unseen domains. The latest and most widely acclaimed development is user-prompted data, epitomized by OpenAI&#x27;s ChatGPT. In this thesis, I first aim to categorize different applications and challenges of synthetic data. Subsequently, I delve deeper into three of these applications, providing synthetic data methods for fairness, augmenting dataset size for small subgroups, and simulating data. Next I turn to the core critique of synthetic data---to what extent can we trust the synthetic data? I include two chapters on this topic. The first explores how we should publish and use synthetic data, such that we can understand whether the downstream results (e.g. predictions of a downstream model) are indeed reliable. The second explores how we can test whether synthetic data is indeed private. At last, we include an outlook on promising synthetic data research directions.","abstract_has_math":false,"creators":["Van Breugel, Boris"],"institution":"University of Cambridge","degree_name":"Doctor of Philosophy (PhD)","degree_level":"Doctoral","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["van der Schaar, Mihaela"],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-09-18","date_published":"2024-09-18","updated_at":"2026-07-22T22:24:00Z","subjects":["Generative models","Synthetic data","Machine learning","Artificial intelligence","Fairness","Privacy"],"languages":["eng"],"rights":[],"rights_urls":["https://www.repository.cam.ac.uk/bitstreams/cf429ed2-ecd2-4a6a-89f0-7ae56f6c6d5a/download","https://creativecommons.org/licenses/by/4.0/"],"identifier_entries":[]},"links":{"outbound_url":"https://doi.org/10.17863/CAM.121707","outbound_label":"DOI","outbound_source":"dc:identifier.doi"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["van der Schaar, Mihaela"]},{"key":"dc:creator","label":"Author","values":["Van Breugel, Boris"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2024-09-18"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["University of Cambridge"]},{"key":"dc:relation.isreferencedby.uri","label":"Dc Relation Isreferencedby URI","values":["https://www.repository.cam.ac.uk/handle/1810/390004"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["Doctoral"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["Doctor of Philosophy (PhD)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Generative models","Synthetic data","Machine learning","Artificial intelligence","Fairness","Privacy"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["https://www.repository.cam.ac.uk/bitstreams/cf429ed2-ecd2-4a6a-89f0-7ae56f6c6d5a/download","https://creativecommons.org/licenses/by/4.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.17863/CAM.121707"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://www.repository.cam.ac.uk/bitstreams/c7ae89c1-d579-4c6c-85de-d78e1291671f/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Science relies on data. Real data can be severely limiting: it may be privacy-sensitive, unfair, unbalanced, unrepresentative, or it may simply not exist for the setting of interest. These limitations continuously constrain how the machine learning (ML) community operates---from training to testing, from model development to deployment. Synthetic data is a type of data that is generated as to resemble (some part of) the real data, while overcoming some of these real data limitations and improving AI trustworthiness. Advances in deep generative modelling have made synthetic data more real than ever, and as a result there is a steep rise in research that aims to replace real data with synthetic data. The first uses of synthetic data were mostly privacy-focused, aiming to create realistic synthetic data that mimics the real data but does not disclose sensitive information. More recently, however, there has been an increasing interest in extending synthetic data to use cases where it improves upon real data, for example providing better fairness, augmenting the dataset size, and creating or simulating data for unseen domains. The latest and most widely acclaimed development is user-prompted data, epitomized by OpenAI's ChatGPT. In this thesis, I first aim to categorize different applications and challenges of synthetic data. Subsequently, I delve deeper into three of these applications, providing synthetic data methods for fairness, augmenting dataset size for small subgroups, and simulating data. Next I turn to the core critique of synthetic data---to what extent can we trust the synthetic data? I include two chapters on this topic. The first explores how we should publish and use synthetic data, such that we can understand whether the downstream results (e.g. predictions of a downstream model) are indeed reliable. The second explores how we can test whether synthetic data is indeed private. At last, we include an outlook on promising synthetic data research directions."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["71dbb26741850297508f4d75c776d5f9","87eda9de84448d1f82354d60eee3eb5f"]},{"key":"dc:title","label":"Title","values":["Generating Trustworthy Synthetic Data"]}]}],"canonical_facts":{"dc:contributor.advisor":["van der Schaar, Mihaela"],"dc:creator":["Van Breugel, Boris"],"dc:date.issued":["2024-09-18"],"dc:description.abstract":["Science relies on data. Real data can be severely limiting: it may be privacy-sensitive, unfair, unbalanced, unrepresentative, or it may simply not exist for the setting of interest. These limitations continuously constrain how the machine learning (ML) community operates---from training to testing, from model development to deployment. Synthetic data is a type of data that is generated as to resemble (some part of) the real data, while overcoming some of these real data limitations and improving AI trustworthiness. Advances in deep generative modelling have made synthetic data more real than ever, and as a result there is a steep rise in research that aims to replace real data with synthetic data. The first uses of synthetic data were mostly privacy-focused, aiming to create realistic synthetic data that mimics the real data but does not disclose sensitive information. More recently, however, there has been an increasing interest in extending synthetic data to use cases where it improves upon real data, for example providing better fairness, augmenting the dataset size, and creating or simulating data for unseen domains. The latest and most widely acclaimed development is user-prompted data, epitomized by OpenAI's ChatGPT. In this thesis, I first aim to categorize different applications and challenges of synthetic data. Subsequently, I delve deeper into three of these applications, providing synthetic data methods for fairness, augmenting dataset size for small subgroups, and simulating data. Next I turn to the core critique of synthetic data---to what extent can we trust the synthetic data? I include two chapters on this topic. The first explores how we should publish and use synthetic data, such that we can understand whether the downstream results (e.g. predictions of a downstream model) are indeed reliable. The second explores how we can test whether synthetic data is indeed private. At last, we include an outlook on promising synthetic data research directions."],"dc:format.checksum.md5":["71dbb26741850297508f4d75c776d5f9","87eda9de84448d1f82354d60eee3eb5f"],"dc:identifier.doi":["https://doi.org/10.17863/CAM.121707"],"dc:identifier.uri":["https://www.repository.cam.ac.uk/bitstreams/c7ae89c1-d579-4c6c-85de-d78e1291671f/download"],"dc:language":["eng"],"dc:publisher.institution":["University of Cambridge"],"dc:relation.isreferencedby.uri":["https://www.repository.cam.ac.uk/handle/1810/390004"],"dc:rights":["https://www.repository.cam.ac.uk/bitstreams/cf429ed2-ecd2-4a6a-89f0-7ae56f6c6d5a/download","https://creativecommons.org/licenses/by/4.0/"],"dc:subject":["Generative models","Synthetic data","Machine learning","Artificial intelligence","Fairness","Privacy"],"dc:title":["Generating Trustworthy Synthetic Data"],"dc:type":["Thesis"],"dc:type.qualificationlevel":["Doctoral"],"dc:type.qualificationname":["Doctor of Philosophy (PhD)"]},"updated_at":"2026-07-22T22:24:00Z"}