{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/108673"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/108673","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Statistical inference for high-dimensional data","abstract":"Statistical inference is a procedure of using collected observations to deduce properties of the underlying data generating process. In this thesis, we investigate three important problems in high-dimensional statistics and develop some new methods and theory, which show the limitation of some existing approaches and motivate the use of our proposed methods. In the first chapter, we study distance covariance, Hilbert-Schmidt covariance (aka Hilbert-Schmidt independence criterion [Gretton et al. (2008)] and related independence tests under the high dimensional scenario. We show that the sample distance/Hilbert-Schmidt covariance between two random vectors can be approximated by the sum of squared componentwise sample cross-covariances up to an asymptotically constant factor, which indicates that the distance/Hilbert-Schmidt covariance based test can only capture linear dependence in high dimension. Under the assumption that the components within each high-dimensional vector are weakly dependent, the distance correlation based t test developed by Szekely and Rizzo (2013) for independence is shown to have trivial limiting power when the two random vectors are nonlinearly dependent but component-wisely uncorrelated. This new and surprising phenomenon, which seems to be discovered for the first time, is further confirmed in our simulation study. As a remedy, we propose tests based on an aggregation of marginal sample distance/Hilbert-Schmidt covariances and show their superior power behavior against their joint counterparts in simulations. We further extend the distance correlation based $t$ test to those based on Hilbert-Schmidt covariance and marginal distance/Hilbert-Schmidt covariance. A novel unified approach is developed to analyze the studentized sample distance/Hilbert-Schmidt covariance as well as the studentized sample marginal distance covariance under both null and alternative hypothesis. Our theoretical and simulation results shed light on the limitation of distance/Hilbert-Schmidt covariance when used jointly in the high dimensional setting and suggest the aggregation of marginal distance/Hilbert-Schmidt covariance as a useful alternative. In the second chapter, we study a class of two sample test statistics based on inter-point distances in the high dimensional and low/medium sample size setting. Our test statistics include the well-known energy distance and maximum mean discrepancy with Gaussian and Laplacian kernels, and the critical values are obtained via permutations. We show that all these tests are inconsistent when the two high dimensional distributions correspond to the same marginal distributions but differ in other aspects of the distributions. The tests based on energy distance and maximum mean discrepancy mainly target the differences between marginal means and variances, whereas the test based on L1-distance can capture the difference in marginal distributions. Our theory sheds new light on the limitation of inter-point distance based tests, the impact of different distance metrics, and the behavior of permutation tests in high dimension. Some simulation results and a real data illustration are also presented to corroborate our theoretical findings. In the third chapter, we propose a new methodology for change point detection of a high-dimensional time series. We extend the U-statistic based approach of Wang et al. (2019) by applying the trimming technique and utilizing the self-normalization principle. Under the fixed-b asymptotics, where we fix the proportion of trimming parameter over the sample size, we derive the limiting distributions of our test statistic under both the null and local alternatives of a single mean change. Furthermore, we combine our test statistic with the wild binary segmentation procedure to perform the change-point estimation. Empirical simulations demonstrate that the trimming technique is effective and necessary for both testing and estimation when there is strong temporal dependence. As an important theoretical contribution, we derive the weak convergence of the U-statistic based processes for high-dimensional linear process and show the applicability of BN decomposition in high dimension.","abstract_html":"Statistical inference is a procedure of using collected observations to deduce properties of the underlying data generating process. In this thesis, we investigate three important problems in high-dimensional statistics and develop some new methods and theory, which show the limitation of some existing approaches and motivate the use of our proposed methods. In the first chapter, we study distance covariance, Hilbert-Schmidt covariance (aka Hilbert-Schmidt independence criterion [Gretton et al. (2008)] and related independence tests under the high dimensional scenario. We show that the sample distance/Hilbert-Schmidt covariance between two random vectors can be approximated by the sum of squared componentwise sample cross-covariances up to an asymptotically constant factor, which indicates that the distance/Hilbert-Schmidt covariance based test can only capture linear dependence in high dimension. Under the assumption that the components within each high-dimensional vector are weakly dependent, the distance correlation based t test developed by Szekely and Rizzo (2013) for independence is shown to have trivial limiting power when the two random vectors are nonlinearly dependent but component-wisely uncorrelated. This new and surprising phenomenon, which seems to be discovered for the first time, is further confirmed in our simulation study. As a remedy, we propose tests based on an aggregation of marginal sample distance/Hilbert-Schmidt covariances and show their superior power behavior against their joint counterparts in simulations. We further extend the distance correlation based $t$ test to those based on Hilbert-Schmidt covariance and marginal distance/Hilbert-Schmidt covariance. A novel unified approach is developed to analyze the studentized sample distance/Hilbert-Schmidt covariance as well as the studentized sample marginal distance covariance under both null and alternative hypothesis. Our theoretical and simulation results shed light on the limitation of distance/Hilbert-Schmidt covariance when used jointly in the high dimensional setting and suggest the aggregation of marginal distance/Hilbert-Schmidt covariance as a useful alternative. In the second chapter, we study a class of two sample test statistics based on inter-point distances in the high dimensional and low/medium sample size setting. Our test statistics include the well-known energy distance and maximum mean discrepancy with Gaussian and Laplacian kernels, and the critical values are obtained via permutations. We show that all these tests are inconsistent when the two high dimensional distributions correspond to the same marginal distributions but differ in other aspects of the distributions. The tests based on energy distance and maximum mean discrepancy mainly target the differences between marginal means and variances, whereas the test based on L1-distance can capture the difference in marginal distributions. Our theory sheds new light on the limitation of inter-point distance based tests, the impact of different distance metrics, and the behavior of permutation tests in high dimension. Some simulation results and a real data illustration are also presented to corroborate our theoretical findings. In the third chapter, we propose a new methodology for change point detection of a high-dimensional time series. We extend the U-statistic based approach of Wang et al. (2019) by applying the trimming technique and utilizing the self-normalization principle. Under the fixed-b asymptotics, where we fix the proportion of trimming parameter over the sample size, we derive the limiting distributions of our test statistic under both the null and local alternatives of a single mean change. Furthermore, we combine our test statistic with the wild binary segmentation procedure to perform the change-point estimation. Empirical simulations demonstrate that the trimming technique is effective and necessary for both testing and estimation when there is strong temporal dependence. As an important theoretical contribution, we derive the weak convergence of the U-statistic based processes for high-dimensional linear process and show the applicability of BN decomposition in high dimension.","abstract_has_math":true,"creators":["Zhu, Changbo"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Statistics","degree_department":null,"school":null,"contributors":["Shao, Xiaofeng","Chen, Xiaohui","Fellouris, Geogious","Marden, John I."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2020,"date_issued":"2020-10-07T22:49:47Z","date_published":"2020-10-07T22:49:47Z","updated_at":"2026-07-22T22:24:48Z","subjects":["High dimensionality","U-statistics","Independence test","Two Sample Test","Change Point Detection","Power Analysis"],"languages":["en"],"rights":["Copyright 2020 Changbo Zhu"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/108673","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Shao, Xiaofeng","Chen, Xiaohui","Fellouris, Geogious","Marden, John I."]},{"key":"dc:creator","label":"Author","values":["Zhu, Changbo"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2020-10-07T22:49:47Z","2022-10-07T22:50:13Z","2020-07-15","2020-08"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Statistics"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["High dimensionality","U-statistics","Independence test","Two Sample Test","Change Point Detection","Power Analysis"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2020 Changbo Zhu"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/108673"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Statistical inference is a procedure of using collected observations to deduce properties of the underlying data generating process. In this thesis, we investigate three important problems in high-dimensional statistics and develop some new methods and theory, which show the limitation of some existing approaches and motivate the use of our proposed methods. In the first chapter, we study distance covariance, Hilbert-Schmidt covariance (aka Hilbert-Schmidt independence criterion [Gretton et al. (2008)] and related independence tests under the high dimensional scenario. We show that the sample distance/Hilbert-Schmidt covariance between two random vectors can be approximated by the sum of squared componentwise sample cross-covariances up to an asymptotically constant factor, which indicates that the distance/Hilbert-Schmidt covariance based test can only capture linear dependence in high dimension. Under the assumption that the components within each high-dimensional vector are weakly dependent, the distance correlation based t test developed by Szekely and Rizzo (2013) for independence is shown to have trivial limiting power when the two random vectors are nonlinearly dependent but component-wisely uncorrelated. This new and surprising phenomenon, which seems to be discovered for the first time, is further confirmed in our simulation study. As a remedy, we propose tests based on an aggregation of marginal sample distance/Hilbert-Schmidt covariances and show their superior power behavior against their joint counterparts in simulations. We further extend the distance correlation based $t$ test to those based on Hilbert-Schmidt covariance and marginal distance/Hilbert-Schmidt covariance. A novel unified approach is developed to analyze the studentized sample distance/Hilbert-Schmidt covariance as well as the studentized sample marginal distance covariance under both null and alternative hypothesis. Our theoretical and simulation results shed light on the limitation of distance/Hilbert-Schmidt covariance when used jointly in the high dimensional setting and suggest the aggregation of marginal distance/Hilbert-Schmidt covariance as a useful alternative. In the second chapter, we study a class of two sample test statistics based on inter-point distances in the high dimensional and low/medium sample size setting. Our test statistics include the well-known energy distance and maximum mean discrepancy with Gaussian and Laplacian kernels, and the critical values are obtained via permutations. We show that all these tests are inconsistent when the two high dimensional distributions correspond to the same marginal distributions but differ in other aspects of the distributions. The tests based on energy distance and maximum mean discrepancy mainly target the differences between marginal means and variances, whereas the test based on L1-distance can capture the difference in marginal distributions. Our theory sheds new light on the limitation of inter-point distance based tests, the impact of different distance metrics, and the behavior of permutation tests in high dimension. Some simulation results and a real data illustration are also presented to corroborate our theoretical findings. In the third chapter, we propose a new methodology for change point detection of a high-dimensional time series. We extend the U-statistic based approach of Wang et al. (2019) by applying the trimming technique and utilizing the self-normalization principle. Under the fixed-b asymptotics, where we fix the proportion of trimming parameter over the sample size, we derive the limiting distributions of our test statistic under both the null and local alternatives of a single mean change. Furthermore, we combine our test statistic with the wild binary segmentation procedure to perform the change-point estimation. Empirical simulations demonstrate that the trimming technique is effective and necessary for both testing and estimation when there is strong temporal dependence. As an important theoretical contribution, we derive the weak convergence of the U-statistic based processes for high-dimensional linear process and show the applicability of BN decomposition in high dimension.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2022-08-01","The student, Changbo Zhu, accepted the attached license on 2020-07-05 at 10:52.","The student, Changbo Zhu, submitted this Dissertation for approval on 2020-07-05 at 11:13.","This Dissertation was approved for publication on 2020-07-15 at 07:49.","DSpace SAF Submission Ingestion Package generated from Vireo submission #15495 on 2020-10-02 at 15:49:52","Made available in DSpace on 2020-10-07T22:49:47Z (GMT). No. of bitstreams: 3 ZHU-DISSERTATION-2020.pdf: 1197049 bytes, checksum: aa22e0fa2636166b6c3345c1c96010d6 (MD5) LICENSE.txt: 4208 bytes, checksum: 114946439a859953796e22cb47873c49 (MD5) PROQUEST_LICENSE.txt: 4554 bytes, checksum: 03a995fe9d285b3e2ff432c24f2b43e0 (MD5) Previous issue date: 2020-07-15","Embargo set by: Seth Robbins for item 116302 Lift date: 2022-10-07T22:50:13Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Statistical inference for high-dimensional data"]}]}],"canonical_facts":{"dc:contributor":["Shao, Xiaofeng","Chen, Xiaohui","Fellouris, Geogious","Marden, John I."],"dc:creator":["Zhu, Changbo"],"dc:date":["2020-10-07T22:49:47Z","2022-10-07T22:50:13Z","2020-07-15","2020-08"],"dc:description":["Statistical inference is a procedure of using collected observations to deduce properties of the underlying data generating process. In this thesis, we investigate three important problems in high-dimensional statistics and develop some new methods and theory, which show the limitation of some existing approaches and motivate the use of our proposed methods. In the first chapter, we study distance covariance, Hilbert-Schmidt covariance (aka Hilbert-Schmidt independence criterion [Gretton et al. (2008)] and related independence tests under the high dimensional scenario. We show that the sample distance/Hilbert-Schmidt covariance between two random vectors can be approximated by the sum of squared componentwise sample cross-covariances up to an asymptotically constant factor, which indicates that the distance/Hilbert-Schmidt covariance based test can only capture linear dependence in high dimension. Under the assumption that the components within each high-dimensional vector are weakly dependent, the distance correlation based t test developed by Szekely and Rizzo (2013) for independence is shown to have trivial limiting power when the two random vectors are nonlinearly dependent but component-wisely uncorrelated. This new and surprising phenomenon, which seems to be discovered for the first time, is further confirmed in our simulation study. As a remedy, we propose tests based on an aggregation of marginal sample distance/Hilbert-Schmidt covariances and show their superior power behavior against their joint counterparts in simulations. We further extend the distance correlation based $t$ test to those based on Hilbert-Schmidt covariance and marginal distance/Hilbert-Schmidt covariance. A novel unified approach is developed to analyze the studentized sample distance/Hilbert-Schmidt covariance as well as the studentized sample marginal distance covariance under both null and alternative hypothesis. Our theoretical and simulation results shed light on the limitation of distance/Hilbert-Schmidt covariance when used jointly in the high dimensional setting and suggest the aggregation of marginal distance/Hilbert-Schmidt covariance as a useful alternative. In the second chapter, we study a class of two sample test statistics based on inter-point distances in the high dimensional and low/medium sample size setting. Our test statistics include the well-known energy distance and maximum mean discrepancy with Gaussian and Laplacian kernels, and the critical values are obtained via permutations. We show that all these tests are inconsistent when the two high dimensional distributions correspond to the same marginal distributions but differ in other aspects of the distributions. The tests based on energy distance and maximum mean discrepancy mainly target the differences between marginal means and variances, whereas the test based on L1-distance can capture the difference in marginal distributions. Our theory sheds new light on the limitation of inter-point distance based tests, the impact of different distance metrics, and the behavior of permutation tests in high dimension. Some simulation results and a real data illustration are also presented to corroborate our theoretical findings. In the third chapter, we propose a new methodology for change point detection of a high-dimensional time series. We extend the U-statistic based approach of Wang et al. (2019) by applying the trimming technique and utilizing the self-normalization principle. Under the fixed-b asymptotics, where we fix the proportion of trimming parameter over the sample size, we derive the limiting distributions of our test statistic under both the null and local alternatives of a single mean change. Furthermore, we combine our test statistic with the wild binary segmentation procedure to perform the change-point estimation. Empirical simulations demonstrate that the trimming technique is effective and necessary for both testing and estimation when there is strong temporal dependence. As an important theoretical contribution, we derive the weak convergence of the U-statistic based processes for high-dimensional linear process and show the applicability of BN decomposition in high dimension.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2022-08-01","The student, Changbo Zhu, accepted the attached license on 2020-07-05 at 10:52.","The student, Changbo Zhu, submitted this Dissertation for approval on 2020-07-05 at 11:13.","This Dissertation was approved for publication on 2020-07-15 at 07:49.","DSpace SAF Submission Ingestion Package generated from Vireo submission #15495 on 2020-10-02 at 15:49:52","Made available in DSpace on 2020-10-07T22:49:47Z (GMT). No. of bitstreams: 3 ZHU-DISSERTATION-2020.pdf: 1197049 bytes, checksum: aa22e0fa2636166b6c3345c1c96010d6 (MD5) LICENSE.txt: 4208 bytes, checksum: 114946439a859953796e22cb47873c49 (MD5) PROQUEST_LICENSE.txt: 4554 bytes, checksum: 03a995fe9d285b3e2ff432c24f2b43e0 (MD5) Previous issue date: 2020-07-15","Embargo set by: Seth Robbins for item 116302 Lift date: 2022-10-07T22:50:13Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/108673"],"dc:language":["en"],"dc:rights":["Copyright 2020 Changbo Zhu"],"dc:subject":["High dimensionality","U-statistics","Independence test","Two Sample Test","Change Point Detection","Power Analysis"],"dc:title":["Statistical inference for high-dimensional data"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Statistics"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:48Z"}