{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/124538"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/124538","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Statistical uncertainty quantification for machine learning models and training acceleration for graph neural networks","abstract":"Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2026-05-01","abstract_html":"Submission published under a 24 month embargo labeled &#x27;U of I Access&#x27;, the embargo will last until 2026-05-01","abstract_has_math":false,"creators":["Xu, Tianning"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Statistics","degree_department":null,"school":null,"contributors":["Zhu, Ruoqing","Shao, Xiaofeng","Yang, Yun","Zhao, Sihai Dave"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-05","date_published":"2024-05","updated_at":"2026-07-22T22:25:02Z","subjects":["U-statistics","Random Forest","Variance Estimation","Conformal Prediction"],"languages":["en","eng"],"rights":["Copyright 2024 Tianning Xu"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/124538","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Zhu, Ruoqing","Shao, Xiaofeng","Yang, Yun","Zhao, Sihai Dave"]},{"key":"dc:creator","label":"Author","values":["Xu, Tianning"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2024-05","2024-04-23"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Statistics"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["U-statistics","Random Forest","Variance Estimation","Conformal Prediction"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2024 Tianning Xu"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/124538"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2026-05-01","The student, Tianning Xu, accepted the attached license on 2024-04-17 at 22:57.","The student, Tianning Xu, submitted this Dissertation for approval on 2024-04-17 at 23:13.","This Dissertation was approved for publication on 2024-04-23 at 13:01.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20473 on 2024-09-16 at 00:43:46","Uncertainty quantification in statistical models is crucial across various applications that require credible conclusions, particularly in domains such as estimating treatment effects. A significant emphasis in current research lies in the construction of confidence sets and prediction sets to complement point estimations effectively, forming the core focus of the first and second chapters of this thesis, respectively. In addition, with the exponential growth in data sizes, the acceleration of deep neural network training has emerged as a pivotal technique, which is discussed in the third chapter within the context of graph neural networks. In the first chapter, we dive into the variance estimation for subbaging ensemble learning models, such as random forest by infinite-order U-statistics (IOUS). While normality results of IOUS have been studied extensively, its variance estimation and theoretical properties remain mostly unexplored. Existing approaches mainly utilize the leading term dominance property in the Hoeffding decomposition. However, such a view usually leads to biased estimation when the kernel size is large relative to the sample size. On the other hand, while several unbiased estimators exist in the literature, their relationships and theoretical properties (e.g., ratio consistency) have never been studied. These limitations lead to unguaranteed asymptotic coverage of constructed confidence intervals. To bridge these gaps in the literature, we propose a new view of the Hoeffding decomposition for variance estimation that leads to an unbiased estimator. Instead of leading term dominance, our view utilizes the dominance of the peak region. Moreover, we establish the connection and equivalence of our estimator with several existing unbiased variance estimators. Theoretically, we are the first to establish the ratio consistency of such a variance estimator, which justifies the coverage rate of confidence intervals constructed from random forests. Numerically, we further propose a local smoothing procedure to improve the estimator's finite sample performance. Extensive simulation studies show that our estimators enjoy lower bias and achieve targeted coverage rates. In the second chapter, we explore the domain of conformal prediction (CP) to quantify uncertainty in black-box models, which offers guaranteed marginal coverage without distributional assumptions. Localized conformal prediction (LCP) improves the conditional coverage of CP's prediction sets by prioritizing local samples based on similarity. However, existing LCP uses the Gaussian kernel, which are non-adaptive kernel weights, resulting in inadequate conditional coverage rates for prediction sets. In this paper, we introduce the integration of random forest kernels into LCP. Additionally, we propose a novel approach termed distributionally sensitive random forest (DS-forest) to learn the kernel weights, which incorporates a normalized Kolmogorov-Smirnov statistic into the split rule. Furthermore, we demonstrate the applicability of our method in a two-stage random forest to construct prediction intervals in regression tasks and propose an inexact algorithm that eliminates the need for calibration sets for forest models. Numerical experiments show the superiority of LCP with DS-forest over existing LCP techniques and quantile forests, guaranteeing marginal coverage while exhibiting notable improvements in conditional coverage. In the third chapter, we study the approximating and accelerating node embedding aggregation in graph convolutional networks (GCNs) training. Among sampling techniques, a layer-wise approach recursively performs importance sampling to select neighbors jointly for existing nodes in each layer. We revisit the approach from a matrix approximation perspective and identify two issues in the existing layer-wise sampling methods: suboptimal sampling probabilities and estimation biases induced by sampling without replacement. To address these issues, we accordingly propose two remedies: a new principle for constructing sampling probabilities and an efficient debiasing algorithm. Improvements are demonstrated by an extensive analysis of the estimation variance and experiments on common benchmarks."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Statistical uncertainty quantification for machine learning models and training acceleration for graph neural networks"]}]}],"canonical_facts":{"dc:contributor":["Zhu, Ruoqing","Shao, Xiaofeng","Yang, Yun","Zhao, Sihai Dave"],"dc:creator":["Xu, Tianning"],"dc:date":["2024-05","2024-04-23"],"dc:description":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2026-05-01","The student, Tianning Xu, accepted the attached license on 2024-04-17 at 22:57.","The student, Tianning Xu, submitted this Dissertation for approval on 2024-04-17 at 23:13.","This Dissertation was approved for publication on 2024-04-23 at 13:01.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20473 on 2024-09-16 at 00:43:46","Uncertainty quantification in statistical models is crucial across various applications that require credible conclusions, particularly in domains such as estimating treatment effects. A significant emphasis in current research lies in the construction of confidence sets and prediction sets to complement point estimations effectively, forming the core focus of the first and second chapters of this thesis, respectively. In addition, with the exponential growth in data sizes, the acceleration of deep neural network training has emerged as a pivotal technique, which is discussed in the third chapter within the context of graph neural networks. In the first chapter, we dive into the variance estimation for subbaging ensemble learning models, such as random forest by infinite-order U-statistics (IOUS). While normality results of IOUS have been studied extensively, its variance estimation and theoretical properties remain mostly unexplored. Existing approaches mainly utilize the leading term dominance property in the Hoeffding decomposition. However, such a view usually leads to biased estimation when the kernel size is large relative to the sample size. On the other hand, while several unbiased estimators exist in the literature, their relationships and theoretical properties (e.g., ratio consistency) have never been studied. These limitations lead to unguaranteed asymptotic coverage of constructed confidence intervals. To bridge these gaps in the literature, we propose a new view of the Hoeffding decomposition for variance estimation that leads to an unbiased estimator. Instead of leading term dominance, our view utilizes the dominance of the peak region. Moreover, we establish the connection and equivalence of our estimator with several existing unbiased variance estimators. Theoretically, we are the first to establish the ratio consistency of such a variance estimator, which justifies the coverage rate of confidence intervals constructed from random forests. Numerically, we further propose a local smoothing procedure to improve the estimator's finite sample performance. Extensive simulation studies show that our estimators enjoy lower bias and achieve targeted coverage rates. In the second chapter, we explore the domain of conformal prediction (CP) to quantify uncertainty in black-box models, which offers guaranteed marginal coverage without distributional assumptions. Localized conformal prediction (LCP) improves the conditional coverage of CP's prediction sets by prioritizing local samples based on similarity. However, existing LCP uses the Gaussian kernel, which are non-adaptive kernel weights, resulting in inadequate conditional coverage rates for prediction sets. In this paper, we introduce the integration of random forest kernels into LCP. Additionally, we propose a novel approach termed distributionally sensitive random forest (DS-forest) to learn the kernel weights, which incorporates a normalized Kolmogorov-Smirnov statistic into the split rule. Furthermore, we demonstrate the applicability of our method in a two-stage random forest to construct prediction intervals in regression tasks and propose an inexact algorithm that eliminates the need for calibration sets for forest models. Numerical experiments show the superiority of LCP with DS-forest over existing LCP techniques and quantile forests, guaranteeing marginal coverage while exhibiting notable improvements in conditional coverage. In the third chapter, we study the approximating and accelerating node embedding aggregation in graph convolutional networks (GCNs) training. Among sampling techniques, a layer-wise approach recursively performs importance sampling to select neighbors jointly for existing nodes in each layer. We revisit the approach from a matrix approximation perspective and identify two issues in the existing layer-wise sampling methods: suboptimal sampling probabilities and estimation biases induced by sampling without replacement. To address these issues, we accordingly propose two remedies: a new principle for constructing sampling probabilities and an efficient debiasing algorithm. Improvements are demonstrated by an extensive analysis of the estimation variance and experiments on common benchmarks."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/124538"],"dc:language":["en","eng"],"dc:rights":["Copyright 2024 Tianning Xu"],"dc:subject":["U-statistics","Random Forest","Variance Estimation","Conformal Prediction"],"dc:title":["Statistical uncertainty quantification for machine learning models and training acceleration for graph neural networks"],"dc:type":["text"],"thesis:degree_discipline":["Statistics"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:02Z"}