{"id":{"repo_id":"mit","oai_identifier":"oai:dspace.mit.edu:1721.1/151313"},"canonical_url":"https://search.dev.ndltd.org/etd/mit/oai:dspace.mit.edu:1721.1/151313","repository":{"repo_id":"mit","name":"MIT","base_url":"https://dspace.mit.edu/oai/request"},"display":{"title":"Privacy-Preserving Natural Language Dataset Generation","abstract":"As we depend on data more heavily to power the insights made by machine learning systems, it becomes imperative that we design guarantees for protecting the privacy of such data. Recent research has shown the ease with which attacks such as membership inference or model inversion can extract potentially sensitive training data given the model alone. To prevent curious or malevolent users from gleaning training data through these attacks, we propose the generation of private synthetic datasets to replace the original datasets in training and testing the model. These synthetic datasets will have the same semantic and statistical distribution as the original dataset, but will be differentially private, thus preventing individuals in the dataset from being identified. This would guarantee that no sensitive information from the original dataset can be extracted from the generated synthetic dataset. Compared to related works that dealt with either structured data or unstructured data separately, our work developed a pipeline for generating synthetic datasets given a complex dataset consisting of structured and unstructured text, as well as numerical data. We used a number of metrics to evaluate the generation pipeline according to its statistical similarity to the original dataset, its utility, and its privacy. Our experiments focused on varying the degree of privacy across the sub-modules of the pipeline. We found that we can generate differentially private synthetic datasets whose structured and unstructured components each achieve good performance in similarity, utility, and privacy.","abstract_html":"As we depend on data more heavily to power the insights made by machine learning systems, it becomes imperative that we design guarantees for protecting the privacy of such data. Recent research has shown the ease with which attacks such as membership inference or model inversion can extract potentially sensitive training data given the model alone. To prevent curious or malevolent users from gleaning training data through these attacks, we propose the generation of private synthetic datasets to replace the original datasets in training and testing the model. These synthetic datasets will have the same semantic and statistical distribution as the original dataset, but will be differentially private, thus preventing individuals in the dataset from being identified. This would guarantee that no sensitive information from the original dataset can be extracted from the generated synthetic dataset. Compared to related works that dealt with either structured data or unstructured data separately, our work developed a pipeline for generating synthetic datasets given a complex dataset consisting of structured and unstructured text, as well as numerical data. We used a number of metrics to evaluate the generation pipeline according to its statistical similarity to the original dataset, its utility, and its privacy. Our experiments focused on varying the degree of privacy across the sub-modules of the pipeline. We found that we can generate differentially private synthetic datasets whose structured and unstructured components each achieve good performance in similarity, utility, and privacy.","abstract_has_math":false,"creators":["Chen, Ashley"],"institution":"Massachusetts Institute of Technology","degree_name":"Master","degree_level":null,"degree_discipline":null,"degree_department":"Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science","school":null,"contributors":[],"advisors":["Kagal, Lalana"],"committee_chairs":[],"committee_members":[],"year":2023,"date_issued":"2023-06","date_published":"2023-06","updated_at":"2026-07-22T22:22:11Z","subjects":[],"languages":[],"rights":["In Copyright - Educational Use Permitted","Copyright retained by author(s)"],"rights_urls":["https://rightsstatements.org/page/InC-EDU/1.0/"],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/1721.1/151313","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Kagal, Lalana"]},{"key":"dc:contributor.department","label":"Department","values":["Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science"]},{"key":"dc:creator","label":"Author","values":["Chen, Ashley"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2023-07-31T19:30:37Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2023-07-31T19:30:37Z"]},{"key":"dc:date.issued","label":"Date","values":["2023-06"]},{"key":"dc:publisher","label":"Institution","values":["Massachusetts Institute of Technology"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master","Master of Engineering in Electrical Engineering and Computer Science"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["In Copyright - Educational Use Permitted","Copyright retained by author(s)"]},{"key":"dc:rights.uri","label":"Rights URI","values":["https://rightsstatements.org/page/InC-EDU/1.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/1721.1/151313"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["As we depend on data more heavily to power the insights made by machine learning systems, it becomes imperative that we design guarantees for protecting the privacy of such data. Recent research has shown the ease with which attacks such as membership inference or model inversion can extract potentially sensitive training data given the model alone. To prevent curious or malevolent users from gleaning training data through these attacks, we propose the generation of private synthetic datasets to replace the original datasets in training and testing the model. These synthetic datasets will have the same semantic and statistical distribution as the original dataset, but will be differentially private, thus preventing individuals in the dataset from being identified. This would guarantee that no sensitive information from the original dataset can be extracted from the generated synthetic dataset. Compared to related works that dealt with either structured data or unstructured data separately, our work developed a pipeline for generating synthetic datasets given a complex dataset consisting of structured and unstructured text, as well as numerical data. We used a number of metrics to evaluate the generation pipeline according to its statistical similarity to the original dataset, its utility, and its privacy. Our experiments focused on varying the degree of privacy across the sub-modules of the pipeline. We found that we can generate differentially private synthetic datasets whose structured and unstructured components each achieve good performance in similarity, utility, and privacy."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["M.Eng."]},{"key":"dc:title","label":"Title","values":["Privacy-Preserving Natural Language Dataset Generation"]}]}],"canonical_facts":{"dc:contributor.advisor":["Kagal, Lalana"],"dc:contributor.department":["Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science"],"dc:creator":["Chen, Ashley"],"dc:date.accessioned":["2023-07-31T19:30:37Z"],"dc:date.available":["2023-07-31T19:30:37Z"],"dc:date.issued":["2023-06"],"dc:description.abstract":["As we depend on data more heavily to power the insights made by machine learning systems, it becomes imperative that we design guarantees for protecting the privacy of such data. Recent research has shown the ease with which attacks such as membership inference or model inversion can extract potentially sensitive training data given the model alone. To prevent curious or malevolent users from gleaning training data through these attacks, we propose the generation of private synthetic datasets to replace the original datasets in training and testing the model. These synthetic datasets will have the same semantic and statistical distribution as the original dataset, but will be differentially private, thus preventing individuals in the dataset from being identified. This would guarantee that no sensitive information from the original dataset can be extracted from the generated synthetic dataset. Compared to related works that dealt with either structured data or unstructured data separately, our work developed a pipeline for generating synthetic datasets given a complex dataset consisting of structured and unstructured text, as well as numerical data. We used a number of metrics to evaluate the generation pipeline according to its statistical similarity to the original dataset, its utility, and its privacy. Our experiments focused on varying the degree of privacy across the sub-modules of the pipeline. We found that we can generate differentially private synthetic datasets whose structured and unstructured components each achieve good performance in similarity, utility, and privacy."],"dc:description.degree":["M.Eng."],"dc:identifier.uri":["https://hdl.handle.net/1721.1/151313"],"dc:publisher":["Massachusetts Institute of Technology"],"dc:rights":["In Copyright - Educational Use Permitted","Copyright retained by author(s)"],"dc:rights.uri":["https://rightsstatements.org/page/InC-EDU/1.0/"],"dc:title":["Privacy-Preserving Natural Language Dataset Generation"],"dc:type":["Thesis"],"thesis:degree_name":["Master","Master of Engineering in Electrical Engineering and Computer Science"]},"updated_at":"2026-07-22T22:22:11Z"}