{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/101661"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/101661","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Privacy-preserving seedbased data synthesis","abstract":"How can we share sensitive datasets in such a way as to maximize utility while simultaneously safeguarding privacy? This thesis proposes an answer to this question by re-framing the problem of sharing sensitive datasets as a data synthesis task. Specifically, we propose a framework to synthesize full data records in a privacy-preserving way so that they can be shared instead of the original sensitive data. The core the framework is a technique called seedbased data synthesis. Seedbased data synthesis produces data records by conditioning the output of a generative model on some input data record called the seed. This technique produces synthetic records that are similar to their seeds, which results in high quality outputs. But it simultaneously introduces statistical dependence between synthetic records and their seeds, which may compromise privacy. As a countermeasure, we introduce a new class of techniques that can achieve strong privacy notions in this setting: privacy tests. Privacy tests are algorithms that probabilistically reject candidate synthetics records which are determined to leak sensitive information. Synthetic records that fail the test are simply discarded, whereas those that pass the test are deemed safe and included in the synthetic dataset to be shared. We design two privacy tests that provably yield differential privacy. We analyze the quality of synthetic datasets based on a cryptography-inspired definition of distinguishability: if synthetic data records are indistinguishable from real records, then they are (by definition) as useful as real data. On the theory front, we characterize the utility-privacy trade-off of seedbased data synthesis. On the experimental front, we design an efficient procedure to experimentally quantify distinguishability. We experimentally validate the seedbased data synthesis framework using five probabilistic generative models. Specifically, using real-world datasets as input, we produce synthetic data records for four different application scenarios and data types: location trajectories, census microdata, medical data, and facial images. We evaluate the quality of the produced synthetic records using both application-dependent utility metrics and distinguishability, and show that the framework is capable of producing highly realistic synthetic data records while providing differential privacy for conservative parameters.","abstract_html":"How can we share sensitive datasets in such a way as to maximize utility while simultaneously safeguarding privacy? This thesis proposes an answer to this question by re-framing the problem of sharing sensitive datasets as a data synthesis task. Specifically, we propose a framework to synthesize full data records in a privacy-preserving way so that they can be shared instead of the original sensitive data. The core the framework is a technique called seedbased data synthesis. Seedbased data synthesis produces data records by conditioning the output of a generative model on some input data record called the seed. This technique produces synthetic records that are similar to their seeds, which results in high quality outputs. But it simultaneously introduces statistical dependence between synthetic records and their seeds, which may compromise privacy. As a countermeasure, we introduce a new class of techniques that can achieve strong privacy notions in this setting: privacy tests. Privacy tests are algorithms that probabilistically reject candidate synthetics records which are determined to leak sensitive information. Synthetic records that fail the test are simply discarded, whereas those that pass the test are deemed safe and included in the synthetic dataset to be shared. We design two privacy tests that provably yield differential privacy. We analyze the quality of synthetic datasets based on a cryptography-inspired definition of distinguishability: if synthetic data records are indistinguishable from real records, then they are (by definition) as useful as real data. On the theory front, we characterize the utility-privacy trade-off of seedbased data synthesis. On the experimental front, we design an efficient procedure to experimentally quantify distinguishability. We experimentally validate the seedbased data synthesis framework using five probabilistic generative models. Specifically, using real-world datasets as input, we produce synthetic data records for four different application scenarios and data types: location trajectories, census microdata, medical data, and facial images. We evaluate the quality of the produced synthetic records using both application-dependent utility metrics and distinguishability, and show that the framework is capable of producing highly realistic synthetic data records while providing differential privacy for conservative parameters.","abstract_has_math":false,"creators":["Bindschadler, Vincent"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Gunter, Carl A.","Zhai, ChengXiang","Borisov, Nikita","Smith, Adam D"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2018,"date_issued":"2018-09-27T16:30:15Z","date_published":"2018-09-27T16:30:15Z","updated_at":"2026-07-22T22:24:40Z","subjects":["Synthetic Data","Privacy","Data Privacy"],"languages":["en"],"rights":["Copyright 2018 Vincent Bindschadler"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/101661","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Gunter, Carl A.","Zhai, ChengXiang","Borisov, Nikita","Smith, Adam D"]},{"key":"dc:creator","label":"Author","values":["Bindschadler, Vincent"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2018-09-27T16:30:15Z","2020-09-28T09:15:24Z","2018-07-02","2018-08"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Synthetic Data","Privacy","Data Privacy"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2018 Vincent Bindschadler"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/101661"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["How can we share sensitive datasets in such a way as to maximize utility while simultaneously safeguarding privacy? This thesis proposes an answer to this question by re-framing the problem of sharing sensitive datasets as a data synthesis task. Specifically, we propose a framework to synthesize full data records in a privacy-preserving way so that they can be shared instead of the original sensitive data. The core the framework is a technique called seedbased data synthesis. Seedbased data synthesis produces data records by conditioning the output of a generative model on some input data record called the seed. This technique produces synthetic records that are similar to their seeds, which results in high quality outputs. But it simultaneously introduces statistical dependence between synthetic records and their seeds, which may compromise privacy. As a countermeasure, we introduce a new class of techniques that can achieve strong privacy notions in this setting: privacy tests. Privacy tests are algorithms that probabilistically reject candidate synthetics records which are determined to leak sensitive information. Synthetic records that fail the test are simply discarded, whereas those that pass the test are deemed safe and included in the synthetic dataset to be shared. We design two privacy tests that provably yield differential privacy. We analyze the quality of synthetic datasets based on a cryptography-inspired definition of distinguishability: if synthetic data records are indistinguishable from real records, then they are (by definition) as useful as real data. On the theory front, we characterize the utility-privacy trade-off of seedbased data synthesis. On the experimental front, we design an efficient procedure to experimentally quantify distinguishability. We experimentally validate the seedbased data synthesis framework using five probabilistic generative models. Specifically, using real-world datasets as input, we produce synthetic data records for four different application scenarios and data types: location trajectories, census microdata, medical data, and facial images. We evaluate the quality of the produced synthetic records using both application-dependent utility metrics and distinguishability, and show that the framework is capable of producing highly realistic synthetic data records while providing differential privacy for conservative parameters.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2020-08-01","The student, Vincent Bindschadler, accepted the attached license on 2018-07-01 at 14:15.","The student, Vincent Bindschadler, submitted this Dissertation for approval on 2018-07-01 at 14:20.","This Dissertation was approved for publication on 2018-07-02 at 14:38.","DSpace SAF Submission Ingestion Package generated from Vireo submission #12694 on 2018-09-27 at 11:16:14","Made available in DSpace on 2018-09-27T16:30:15Z (GMT). No. of bitstreams: 2 BINDSCHADLER-DISSERTATION-2018.pdf: 2582831 bytes, checksum: f398eec4c74ecd58a5b249e531dfee31 (MD5) LICENSE.txt: 4217 bytes, checksum: 276505650d0694753fef6ac2ce2835b6 (MD5) Previous issue date: 2018-07-02","Embargo set by: Seth Robbins for item 107759 Lift date: 2020-09-27T16:30:34Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 107759 Lift date: 2020-09-27T16:31:43Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 107759 Lift date: 2020-09-27T16:34:29Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 107759 on 2020-09-28T09:15:24Z."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Privacy-preserving seedbased data synthesis"]}]}],"canonical_facts":{"dc:contributor":["Gunter, Carl A.","Zhai, ChengXiang","Borisov, Nikita","Smith, Adam D"],"dc:creator":["Bindschadler, Vincent"],"dc:date":["2018-09-27T16:30:15Z","2020-09-28T09:15:24Z","2018-07-02","2018-08"],"dc:description":["How can we share sensitive datasets in such a way as to maximize utility while simultaneously safeguarding privacy? This thesis proposes an answer to this question by re-framing the problem of sharing sensitive datasets as a data synthesis task. Specifically, we propose a framework to synthesize full data records in a privacy-preserving way so that they can be shared instead of the original sensitive data. The core the framework is a technique called seedbased data synthesis. Seedbased data synthesis produces data records by conditioning the output of a generative model on some input data record called the seed. This technique produces synthetic records that are similar to their seeds, which results in high quality outputs. But it simultaneously introduces statistical dependence between synthetic records and their seeds, which may compromise privacy. As a countermeasure, we introduce a new class of techniques that can achieve strong privacy notions in this setting: privacy tests. Privacy tests are algorithms that probabilistically reject candidate synthetics records which are determined to leak sensitive information. Synthetic records that fail the test are simply discarded, whereas those that pass the test are deemed safe and included in the synthetic dataset to be shared. We design two privacy tests that provably yield differential privacy. We analyze the quality of synthetic datasets based on a cryptography-inspired definition of distinguishability: if synthetic data records are indistinguishable from real records, then they are (by definition) as useful as real data. On the theory front, we characterize the utility-privacy trade-off of seedbased data synthesis. On the experimental front, we design an efficient procedure to experimentally quantify distinguishability. We experimentally validate the seedbased data synthesis framework using five probabilistic generative models. Specifically, using real-world datasets as input, we produce synthetic data records for four different application scenarios and data types: location trajectories, census microdata, medical data, and facial images. We evaluate the quality of the produced synthetic records using both application-dependent utility metrics and distinguishability, and show that the framework is capable of producing highly realistic synthetic data records while providing differential privacy for conservative parameters.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2020-08-01","The student, Vincent Bindschadler, accepted the attached license on 2018-07-01 at 14:15.","The student, Vincent Bindschadler, submitted this Dissertation for approval on 2018-07-01 at 14:20.","This Dissertation was approved for publication on 2018-07-02 at 14:38.","DSpace SAF Submission Ingestion Package generated from Vireo submission #12694 on 2018-09-27 at 11:16:14","Made available in DSpace on 2018-09-27T16:30:15Z (GMT). No. of bitstreams: 2 BINDSCHADLER-DISSERTATION-2018.pdf: 2582831 bytes, checksum: f398eec4c74ecd58a5b249e531dfee31 (MD5) LICENSE.txt: 4217 bytes, checksum: 276505650d0694753fef6ac2ce2835b6 (MD5) Previous issue date: 2018-07-02","Embargo set by: Seth Robbins for item 107759 Lift date: 2020-09-27T16:30:34Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 107759 Lift date: 2020-09-27T16:31:43Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 107759 Lift date: 2020-09-27T16:34:29Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 107759 on 2020-09-28T09:15:24Z."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/101661"],"dc:language":["en"],"dc:rights":["Copyright 2018 Vincent Bindschadler"],"dc:subject":["Synthetic Data","Privacy","Data Privacy"],"dc:title":["Privacy-preserving seedbased data synthesis"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:40Z"}