{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/113122"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/113122","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Attacks on medical data and centralized social platforms: PU learning and graphical inference","abstract":"Embargo set by: Seth Robbins for item 121048 Lift date: 2024-01-12T22:35:30Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","abstract_html":"Embargo set by: Seth Robbins for item 121048 Lift date: 2024-01-12T22:35:30Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","abstract_has_math":false,"creators":["Su, Du"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Lu, Yi","Hajek, Bruce","Srikant, Rayadurgam","Viswanath, Pramod"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-01-12T22:34:44Z","date_published":"2022-01-12T22:34:44Z","updated_at":"2026-07-22T22:24:53Z","subjects":["privacy","social network","data-mining","density evolution","neural networks"],"languages":["en"],"rights":["Copyright 2021 Du Su"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/113122","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Lu, Yi","Hajek, Bruce","Srikant, Rayadurgam","Viswanath, Pramod"]},{"key":"dc:creator","label":"Author","values":["Su, Du"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022-01-12T22:34:44Z","2024-01-12T22:35:30Z","2021-06-15","2021-08"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["privacy","social network","data-mining","density evolution","neural networks"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2021 Du Su"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/113122"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Embargo set by: Seth Robbins for item 121048 Lift date: 2024-01-12T22:35:30Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only","Sensitive information in user data, such as health status, financial history, and personal preference is facing a high risk of being exposed to adversarial entities, as well as being abused by the service provider. Prohibiting individual-level operations on user records is one natural proposal to protect the sensitive information of user data. Instead, only aggregate statistics of sensitive data are allowed to be stored or released. It is widely believed that aggregation is an effective way of preserving the typical characteristics of users for data mining tasks while preventing the infringement of individual privacy, as the aggregated statistics carry little information of each individual user. However, these privacy-preserving techniques may fail to eliminate the risks of sensitive data leakage. Aggregation is not able to fully hide features of the dataset, and adversarial entities are able to exploit the features of aggregate data to recover the sensitive information in user records. Moreover, social platforms with a high degree of centralization, such as Tiktok, are able to design special aggregation algorithm, in a way that the individual information can be recovered completely from aggregated statistics, allowing them to exploit user preferences. In this thesis, we study the two aforementioned scenarios where aggregation fails to protect individual user's sensitive data. First, we show the hazard of re-identification of sensitive class labels caused by revealing a noisy sample mean of one class. With a novel formulation of the re-identification attack as a generalized positive-unlabeled learning problem, we prove that the risk function of the re-identification problem is closely related to that of learning with complete data. We demonstrate that with a one-sided noisy sample mean, an effective re-identification attack can be devised with existing PU learning algorithms. We then propose a novel algorithm, growPU, that exploits the unique property of sample mean and consistently outperforms existing PU learning algorithms on the re-identification task. GrowPU achieves re-identification accuracy of $93.6\\%$ on the MNIST dataset and $88.1\\%$ on an online behavioral dataset with noiseless sample mean. With noise that guarantees $0.01$-differential privacy, growPU achieves $91.9\\%$ on the MNIST dataset and $84.6\\%$ on the online behavioral dataset. In the second part of the thesis, we present a randomized article-push algorithm and a message-passing reconstruction algorithm, such that social media platforms are able to infer user preferences from only the publicly available aggregate data of article-reads, without storing any individual users' actions. Its $O(n)$ complexity allows the reconstruction algorithm to scale to a large population, as is typical of social media platforms. Moreover, the feasibility of the privacy attack depends on the algorithm using as few articles as possible. We determine the minimum number of articles needed for high probability inference. Given the proportion of users, $0<\\eps<1$, who prefer a given topic, the push algorithm and reconstruction algorithm can achieve an article-to-user ratio $\\beta=\\sqrt{\\eps(1-\\eps)}$, at which phase transition occurs. By formulating the inference problem as a compressed sensing problem, we show that our phase transition threshold $\\sqrt{\\eps(1-\\eps)}$ is extremely close to that of compressed sensing, even when the latter algorithm is of a worst-case $O(n^3)$ complexity and uses a dense Gaussian measurement matrix.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2023-08-01","The student, Du Su, accepted the attached license on 2021-06-11 at 21:06.","The student, Du Su, submitted this Dissertation for approval on 2021-06-11 at 21:14.","This Dissertation was approved for publication on 2021-06-15 at 16:44.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16683 on 2022-01-12 at 12:52:15","Made available in DSpace on 2022-01-12T22:34:44Z (GMT). No. of bitstreams: 2 SU-DISSERTATION-2021.pdf: 1691587 bytes, checksum: d8fed788a05c0d211eaeab4eb3120d46 (MD5) LICENSE.txt: 4202 bytes, checksum: b9e91baf5ff0bb33179e273d405ac038 (MD5) Previous issue date: 2021-06-15"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Attacks on medical data and centralized social platforms: PU learning and graphical inference"]}]}],"canonical_facts":{"dc:contributor":["Lu, Yi","Hajek, Bruce","Srikant, Rayadurgam","Viswanath, Pramod"],"dc:creator":["Su, Du"],"dc:date":["2022-01-12T22:34:44Z","2024-01-12T22:35:30Z","2021-06-15","2021-08"],"dc:description":["Embargo set by: Seth Robbins for item 121048 Lift date: 2024-01-12T22:35:30Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only","Sensitive information in user data, such as health status, financial history, and personal preference is facing a high risk of being exposed to adversarial entities, as well as being abused by the service provider. Prohibiting individual-level operations on user records is one natural proposal to protect the sensitive information of user data. Instead, only aggregate statistics of sensitive data are allowed to be stored or released. It is widely believed that aggregation is an effective way of preserving the typical characteristics of users for data mining tasks while preventing the infringement of individual privacy, as the aggregated statistics carry little information of each individual user. However, these privacy-preserving techniques may fail to eliminate the risks of sensitive data leakage. Aggregation is not able to fully hide features of the dataset, and adversarial entities are able to exploit the features of aggregate data to recover the sensitive information in user records. Moreover, social platforms with a high degree of centralization, such as Tiktok, are able to design special aggregation algorithm, in a way that the individual information can be recovered completely from aggregated statistics, allowing them to exploit user preferences. In this thesis, we study the two aforementioned scenarios where aggregation fails to protect individual user's sensitive data. First, we show the hazard of re-identification of sensitive class labels caused by revealing a noisy sample mean of one class. With a novel formulation of the re-identification attack as a generalized positive-unlabeled learning problem, we prove that the risk function of the re-identification problem is closely related to that of learning with complete data. We demonstrate that with a one-sided noisy sample mean, an effective re-identification attack can be devised with existing PU learning algorithms. We then propose a novel algorithm, growPU, that exploits the unique property of sample mean and consistently outperforms existing PU learning algorithms on the re-identification task. GrowPU achieves re-identification accuracy of $93.6\\%$ on the MNIST dataset and $88.1\\%$ on an online behavioral dataset with noiseless sample mean. With noise that guarantees $0.01$-differential privacy, growPU achieves $91.9\\%$ on the MNIST dataset and $84.6\\%$ on the online behavioral dataset. In the second part of the thesis, we present a randomized article-push algorithm and a message-passing reconstruction algorithm, such that social media platforms are able to infer user preferences from only the publicly available aggregate data of article-reads, without storing any individual users' actions. Its $O(n)$ complexity allows the reconstruction algorithm to scale to a large population, as is typical of social media platforms. Moreover, the feasibility of the privacy attack depends on the algorithm using as few articles as possible. We determine the minimum number of articles needed for high probability inference. Given the proportion of users, $0<\\eps<1$, who prefer a given topic, the push algorithm and reconstruction algorithm can achieve an article-to-user ratio $\\beta=\\sqrt{\\eps(1-\\eps)}$, at which phase transition occurs. By formulating the inference problem as a compressed sensing problem, we show that our phase transition threshold $\\sqrt{\\eps(1-\\eps)}$ is extremely close to that of compressed sensing, even when the latter algorithm is of a worst-case $O(n^3)$ complexity and uses a dense Gaussian measurement matrix.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2023-08-01","The student, Du Su, accepted the attached license on 2021-06-11 at 21:06.","The student, Du Su, submitted this Dissertation for approval on 2021-06-11 at 21:14.","This Dissertation was approved for publication on 2021-06-15 at 16:44.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16683 on 2022-01-12 at 12:52:15","Made available in DSpace on 2022-01-12T22:34:44Z (GMT). No. of bitstreams: 2 SU-DISSERTATION-2021.pdf: 1691587 bytes, checksum: d8fed788a05c0d211eaeab4eb3120d46 (MD5) LICENSE.txt: 4202 bytes, checksum: b9e91baf5ff0bb33179e273d405ac038 (MD5) Previous issue date: 2021-06-15"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/113122"],"dc:language":["en"],"dc:rights":["Copyright 2021 Du Su"],"dc:subject":["privacy","social network","data-mining","density evolution","neural networks"],"dc:title":["Attacks on medical data and centralized social platforms: PU learning and graphical inference"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:53Z"}