{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/42341"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/42341","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Privacy-preserving data publishing and analytics using data cubes","abstract":"Data cubes play an essential role in data analysis and decision support. In a data cube, data from a fact table is aggregated on subsets of the table's dimensions, forming a collection of smaller tables called cuboids. When the fact table includes sensitive data such as salary or diagnosis, publishing even a subset of its cuboids may compromise individuals' privacy. In this thesis, we address several problems about privacy-preserving publishing of data cubes using differential privacy or its extensions, which provide privacy guarantees for individuals by adding noise to query answers. The first problem is about how to improve the data quality in privacy-preserving data cubes. Our noise-control frameworks choose noise source in a data cube, i.e., an initial subset of cuboids to compute directly from the fact table with certain amount of noise to be injected to each of them, and then compute the remaining cuboids from them. We show that it is NP-hard to choose proper noise source for certain noise-control objetives, but provide efficient approximation algorithms. The second problem is about how to enforce consistency in the published cuboids. We proposed several approaches with provable guarantee on the noise bound and one of them can even improve the utility of differentially private cuboids (reducing error). The third problem is about how to calibrate noise in data cubes subject to certain exact background knowledge while we are trying to improve the data quality. The notation of generic differential privacy is applied, and we generalize its properties to plug it into our noise-control frameworks for handling background knowledge. Techniques proposed in this thesis provide advanced principles and major parts of a complete solution towards privacy-preserving publishing of data cubes.","abstract_html":"Data cubes play an essential role in data analysis and decision support. In a data cube, data from a fact table is aggregated on subsets of the table&#x27;s dimensions, forming a collection of smaller tables called cuboids. When the fact table includes sensitive data such as salary or diagnosis, publishing even a subset of its cuboids may compromise individuals&#x27; privacy. In this thesis, we address several problems about privacy-preserving publishing of data cubes using differential privacy or its extensions, which provide privacy guarantees for individuals by adding noise to query answers. The first problem is about how to improve the data quality in privacy-preserving data cubes. Our noise-control frameworks choose noise source in a data cube, i.e., an initial subset of cuboids to compute directly from the fact table with certain amount of noise to be injected to each of them, and then compute the remaining cuboids from them. We show that it is NP-hard to choose proper noise source for certain noise-control objetives, but provide efficient approximation algorithms. The second problem is about how to enforce consistency in the published cuboids. We proposed several approaches with provable guarantee on the noise bound and one of them can even improve the utility of differentially private cuboids (reducing error). The third problem is about how to calibrate noise in data cubes subject to certain exact background knowledge while we are trying to improve the data quality. The notation of generic differential privacy is applied, and we generalize its properties to plug it into our noise-control frameworks for handling background knowledge. Techniques proposed in this thesis provide advanced principles and major parts of a complete solution towards privacy-preserving publishing of data cubes.","abstract_has_math":false,"creators":["Ding, Bolin"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Han, Jiawei","Winslett, Marianne","Zhai, ChengXiang","Machanavajjhala, Ashwin"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2013,"date_issued":"2013-02-03T19:35:52Z","date_published":"2013-02-03T19:35:52Z","updated_at":"2026-07-22T22:25:33Z","subjects":["online analytical processing (OLAP)","data cube","differential privacy","private data analysis"],"languages":["en"],"rights":["Copyright 2012 Bolin Ding"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/42341","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Han, Jiawei","Winslett, Marianne","Zhai, ChengXiang","Machanavajjhala, Ashwin"]},{"key":"dc:creator","label":"Author","values":["Ding, Bolin"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2013-02-03T19:35:52Z","2012-12"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["online analytical processing (OLAP)","data cube","differential privacy","private data analysis"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2012 Bolin Ding"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/42341"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Data cubes play an essential role in data analysis and decision support. In a data cube, data from a fact table is aggregated on subsets of the table's dimensions, forming a collection of smaller tables called cuboids. When the fact table includes sensitive data such as salary or diagnosis, publishing even a subset of its cuboids may compromise individuals' privacy. In this thesis, we address several problems about privacy-preserving publishing of data cubes using differential privacy or its extensions, which provide privacy guarantees for individuals by adding noise to query answers. The first problem is about how to improve the data quality in privacy-preserving data cubes. Our noise-control frameworks choose noise source in a data cube, i.e., an initial subset of cuboids to compute directly from the fact table with certain amount of noise to be injected to each of them, and then compute the remaining cuboids from them. We show that it is NP-hard to choose proper noise source for certain noise-control objetives, but provide efficient approximation algorithms. The second problem is about how to enforce consistency in the published cuboids. We proposed several approaches with provable guarantee on the noise bound and one of them can even improve the utility of differentially private cuboids (reducing error). The third problem is about how to calibrate noise in data cubes subject to certain exact background knowledge while we are trying to improve the data quality. The notation of generic differential privacy is applied, and we generalize its properties to plug it into our noise-control frameworks for handling background knowledge. Techniques proposed in this thesis provide advanced principles and major parts of a complete solution towards privacy-preserving publishing of data cubes.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2012-08-17T15:23:36Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 2 Ding_Bolin.zip: 4388695 bytes, checksum: 36e770866fce974b9c02c2c9062d7aa1 (MD5) Ding_Bolin.pdf: 1196147 bytes, checksum: 04c4b2123bade572f10ac9d10fb6a012 (MD5)","Made available in DSpace on 2013-02-03T19:35:52Z (GMT). No. of bitstreams: 3 Bolin_Ding.pdf: 1196147 bytes, checksum: 04c4b2123bade572f10ac9d10fb6a012 (MD5) Ding_Bolin.zip: 4388695 bytes, checksum: 36e770866fce974b9c02c2c9062d7aa1 (MD5) license.txt: 4058 bytes, checksum: 421999b23eaaa358512dbbc1a8aef242 (MD5)"]},{"key":"dc:title","label":"Title","values":["Privacy-preserving data publishing and analytics using data cubes"]}]}],"canonical_facts":{"dc:contributor":["Han, Jiawei","Winslett, Marianne","Zhai, ChengXiang","Machanavajjhala, Ashwin"],"dc:creator":["Ding, Bolin"],"dc:date":["2013-02-03T19:35:52Z","2012-12"],"dc:description":["Data cubes play an essential role in data analysis and decision support. In a data cube, data from a fact table is aggregated on subsets of the table's dimensions, forming a collection of smaller tables called cuboids. When the fact table includes sensitive data such as salary or diagnosis, publishing even a subset of its cuboids may compromise individuals' privacy. In this thesis, we address several problems about privacy-preserving publishing of data cubes using differential privacy or its extensions, which provide privacy guarantees for individuals by adding noise to query answers. The first problem is about how to improve the data quality in privacy-preserving data cubes. Our noise-control frameworks choose noise source in a data cube, i.e., an initial subset of cuboids to compute directly from the fact table with certain amount of noise to be injected to each of them, and then compute the remaining cuboids from them. We show that it is NP-hard to choose proper noise source for certain noise-control objetives, but provide efficient approximation algorithms. The second problem is about how to enforce consistency in the published cuboids. We proposed several approaches with provable guarantee on the noise bound and one of them can even improve the utility of differentially private cuboids (reducing error). The third problem is about how to calibrate noise in data cubes subject to certain exact background knowledge while we are trying to improve the data quality. The notation of generic differential privacy is applied, and we generalize its properties to plug it into our noise-control frameworks for handling background knowledge. Techniques proposed in this thesis provide advanced principles and major parts of a complete solution towards privacy-preserving publishing of data cubes.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2012-08-17T15:23:36Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 2 Ding_Bolin.zip: 4388695 bytes, checksum: 36e770866fce974b9c02c2c9062d7aa1 (MD5) Ding_Bolin.pdf: 1196147 bytes, checksum: 04c4b2123bade572f10ac9d10fb6a012 (MD5)","Made available in DSpace on 2013-02-03T19:35:52Z (GMT). No. of bitstreams: 3 Bolin_Ding.pdf: 1196147 bytes, checksum: 04c4b2123bade572f10ac9d10fb6a012 (MD5) Ding_Bolin.zip: 4388695 bytes, checksum: 36e770866fce974b9c02c2c9062d7aa1 (MD5) license.txt: 4058 bytes, checksum: 421999b23eaaa358512dbbc1a8aef242 (MD5)"],"dc:identifier":["http://hdl.handle.net/2142/42341"],"dc:language":["en"],"dc:rights":["Copyright 2012 Bolin Ding"],"dc:subject":["online analytical processing (OLAP)","data cube","differential privacy","private data analysis"],"dc:title":["Privacy-preserving data publishing and analytics using data cubes"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:33Z"}