{"id":{"repo_id":"bu","oai_identifier":"oai:open.bu.edu:2144/53099"},"canonical_url":"https://search.dev.ndltd.org/etd/bu/oai:open.bu.edu:2144/53099","repository":{"repo_id":"bu","name":"Boston University","base_url":"https://open.bu.edu/oai/request"},"display":{"title":"Statistical learning with differential privacy","abstract":"In an era characterized by the unprecedented growth of data, the preservation of individual privacy has become a prominent challenge in data-driven decision-making across diverse domains. The concept of differential privacy, a robust mathematical framework, has emerged as the gold standard for providing rigorous data privacy protections. Since its introduction, efforts have been made to develop methodologies for modern statistics and machine learning, integrating the robust guarantees offered by differential privacy. Differential privacy tools mostly operate by introducing carefully calibrated randomness into the training process, thereby ensuring privacy protections. However, this approach introduces a complex privacy-utility trade-off. While random noise enhances privacy, it concurrently poses a threat to utility by potentially undermining the accuracy of the results. Thus, achieving a delicate balance between privacy and utility is crucial for effective problem-solving in data science. This dissertation addresses the challenge by proposing methodologies for statistical learning tasks while upholding privacy principles. The tasks include (1) constructing private confidence intervals for population proportions in survey sampling, (2) addressing private linear regression with uncertainties from prior data linkage, and (3) investigating private stochastic gradient descent algorithms. Differentially private algorithms for the first two tasks are designed using the differential privacy toolbox.The utility of the proposed algorithms, covering confidence intervals and estimator accuracy, is thoroughly examined through theoretical analyses and numerical experiments. For the third task, a novel analytical method based on diffusion processes is developed to accurately capture the utility dynamics of the algorithm in high-dimensional settings for a least-squares problem The analyses within this dissertation leverage a range of mathematical tools, including convex optimization, random matrix theory, and stochastic differential equations. By integrating these tools into the examination of differentially private methodologies, this research aims to contribute both theoretical insights and practical solutions to the intricate interplay between privacy and utility in the realm of statistical learning.","abstract_html":"In an era characterized by the unprecedented growth of data, the preservation of individual privacy has become a prominent challenge in data-driven decision-making across diverse domains. The concept of differential privacy, a robust mathematical framework, has emerged as the gold standard for providing rigorous data privacy protections. Since its introduction, efforts have been made to develop methodologies for modern statistics and machine learning, integrating the robust guarantees offered by differential privacy. Differential privacy tools mostly operate by introducing carefully calibrated randomness into the training process, thereby ensuring privacy protections. However, this approach introduces a complex privacy-utility trade-off. While random noise enhances privacy, it concurrently poses a threat to utility by potentially undermining the accuracy of the results. Thus, achieving a delicate balance between privacy and utility is crucial for effective problem-solving in data science. This dissertation addresses the challenge by proposing methodologies for statistical learning tasks while upholding privacy principles. The tasks include (1) constructing private confidence intervals for population proportions in survey sampling, (2) addressing private linear regression with uncertainties from prior data linkage, and (3) investigating private stochastic gradient descent algorithms. Differentially private algorithms for the first two tasks are designed using the differential privacy toolbox.The utility of the proposed algorithms, covering confidence intervals and estimator accuracy, is thoroughly examined through theoretical analyses and numerical experiments. For the third task, a novel analytical method based on diffusion processes is developed to accurately capture the utility dynamics of the algorithm in high-dimensional settings for a least-squares problem The analyses within this dissertation leverage a range of mathematical tools, including convex optimization, random matrix theory, and stochastic differential equations. By integrating these tools into the examination of differentially private methodologies, this research aims to contribute both theoretical insights and practical solutions to the intricate interplay between privacy and utility in the realm of statistical learning.","abstract_has_math":false,"creators":["Lin, Shurong"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Kolaczyk, Eric D."],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024","date_published":"2024","updated_at":"2026-07-24T01:26:20Z","subjects":["Statistics","Computer science","Data privacy","Differential privacy","Optimization","Regression","Statistical learning","Uncertainty quantification"],"languages":["en_US"],"rights":["Attribution 4.0 International"],"rights_urls":["http://creativecommons.org/licenses/by/4.0/"],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2144/53099","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Kolaczyk, Eric D."]},{"key":"dc:creator","label":"Author","values":["Lin, Shurong"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2026-06-11T13:32:26Z"]},{"key":"dc:date.issued","label":"Date","values":["2024"]},{"key":"dc:type","label":"Dc Type","values":["Thesis/Dissertation"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Statistics","Computer science","Data privacy","Differential privacy","Optimization","Regression","Statistical learning","Uncertainty quantification"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en_US"]},{"key":"dc:rights","label":"Dc Rights","values":["Attribution 4.0 International"]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://creativecommons.org/licenses/by/4.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/2144/53099"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["2024"]},{"key":"dc:description.abstract","label":"Abstract","values":["In an era characterized by the unprecedented growth of data, the preservation of individual privacy has become a prominent challenge in data-driven decision-making across diverse domains. The concept of differential privacy, a robust mathematical framework, has emerged as the gold standard for providing rigorous data privacy protections. Since its introduction, efforts have been made to develop methodologies for modern statistics and machine learning, integrating the robust guarantees offered by differential privacy. Differential privacy tools mostly operate by introducing carefully calibrated randomness into the training process, thereby ensuring privacy protections. However, this approach introduces a complex privacy-utility trade-off. While random noise enhances privacy, it concurrently poses a threat to utility by potentially undermining the accuracy of the results. Thus, achieving a delicate balance between privacy and utility is crucial for effective problem-solving in data science. This dissertation addresses the challenge by proposing methodologies for statistical learning tasks while upholding privacy principles. The tasks include (1) constructing private confidence intervals for population proportions in survey sampling, (2) addressing private linear regression with uncertainties from prior data linkage, and (3) investigating private stochastic gradient descent algorithms. Differentially private algorithms for the first two tasks are designed using the differential privacy toolbox.The utility of the proposed algorithms, covering confidence intervals and estimator accuracy, is thoroughly examined through theoretical analyses and numerical experiments. For the third task, a novel analytical method based on diffusion processes is developed to accurately capture the utility dynamics of the algorithm in high-dimensional settings for a least-squares problem The analyses within this dissertation leverage a range of mathematical tools, including convex optimization, random matrix theory, and stochastic differential equations. By integrating these tools into the examination of differentially private methodologies, this research aims to contribute both theoretical insights and practical solutions to the intricate interplay between privacy and utility in the realm of statistical learning."]},{"key":"dc:title","label":"Title","values":["Statistical learning with differential privacy"]}]}],"canonical_facts":{"dc:contributor.advisor":["Kolaczyk, Eric D."],"dc:creator":["Lin, Shurong"],"dc:date.accessioned":["2026-06-11T13:32:26Z"],"dc:date.issued":["2024"],"dc:description":["2024"],"dc:description.abstract":["In an era characterized by the unprecedented growth of data, the preservation of individual privacy has become a prominent challenge in data-driven decision-making across diverse domains. The concept of differential privacy, a robust mathematical framework, has emerged as the gold standard for providing rigorous data privacy protections. Since its introduction, efforts have been made to develop methodologies for modern statistics and machine learning, integrating the robust guarantees offered by differential privacy. Differential privacy tools mostly operate by introducing carefully calibrated randomness into the training process, thereby ensuring privacy protections. However, this approach introduces a complex privacy-utility trade-off. While random noise enhances privacy, it concurrently poses a threat to utility by potentially undermining the accuracy of the results. Thus, achieving a delicate balance between privacy and utility is crucial for effective problem-solving in data science. This dissertation addresses the challenge by proposing methodologies for statistical learning tasks while upholding privacy principles. The tasks include (1) constructing private confidence intervals for population proportions in survey sampling, (2) addressing private linear regression with uncertainties from prior data linkage, and (3) investigating private stochastic gradient descent algorithms. Differentially private algorithms for the first two tasks are designed using the differential privacy toolbox.The utility of the proposed algorithms, covering confidence intervals and estimator accuracy, is thoroughly examined through theoretical analyses and numerical experiments. For the third task, a novel analytical method based on diffusion processes is developed to accurately capture the utility dynamics of the algorithm in high-dimensional settings for a least-squares problem The analyses within this dissertation leverage a range of mathematical tools, including convex optimization, random matrix theory, and stochastic differential equations. By integrating these tools into the examination of differentially private methodologies, this research aims to contribute both theoretical insights and practical solutions to the intricate interplay between privacy and utility in the realm of statistical learning."],"dc:identifier.uri":["https://hdl.handle.net/2144/53099"],"dc:language.iso":["en_US"],"dc:rights":["Attribution 4.0 International"],"dc:rights.uri":["http://creativecommons.org/licenses/by/4.0/"],"dc:subject":["Statistics","Computer science","Data privacy","Differential privacy","Optimization","Regression","Statistical learning","Uncertainty quantification"],"dc:title":["Statistical learning with differential privacy"],"dc:type":["Thesis/Dissertation"]},"updated_at":"2026-07-24T01:26:20Z"}