{"id":{"repo_id":"uthsc","oai_identifier":"oai:digitalcommons.library.tmc.edu:utgsbs_dissertations-2348"},"canonical_url":"https://search.dev.ndltd.org/etd/uthsc/oai:digitalcommons.library.tmc.edu:utgsbs_dissertations-2348","repository":{"repo_id":"uthsc","name":"University of Texas Health Science Center at Houston","base_url":"https://digitalcommons.library.tmc.edu/do/oai/"},"display":{"title":"Toward Nonparametric Propensity Score Estimation With Guaranteed Covariate Balance","abstract":"<p><em> Establishing clear causality between exposures and outcomes is often complicated </em><em>by confounders in observational studies, which leads to imbalance in covariate distributions </em><em>between treatments and biased treatment effect inference. The propensity score </em><em>(PS) has been widely used to adjust this covariate imbalance in observational data. However, </em><em>the propensity score analysis methods rely on a correctly specified parametric PS </em><em>model. When the model is misspecified, the covariate imbalance may occur, which leads </em><em>to biased estimation of the treatment effect. Therefore, it is necessary to study how to </em><em>improve the model misspecification in propensity score analysis. </em></p> <p><em></em><em> My Ph.D. dissertation consists of three aims. In Aim 1, we examined whether the </em><em>optimization of global balance — the mean balance of covariates or their transformations </em><em>in the overall study population, can circumvent the need for correct propensity score </em><em>model specification, and whether the use of a propensity score model further improves </em><em>the estimation performance compared to methods without modeling propensity score. </em><em>In Aim 2, we developed a propensity score analysis framework, the propensity score </em><em>with local balance (PSLB), which incorporates nonparametric propensity score models </em><em>and improves the balancing property of the estimated propensity score compared to </em><em>existing methods that only optimize the global balance. In Aim 3, we developed a </em><em>subgroup analysis method that is robust to propensity score model misspecification. </em><em>Specifically, we proposed a new algorithm, the guaranteed subgroup balancing propensity </em><em>score (G-SBPS), to ensure exact subgroup balance — the mean balance of covariates in </em><em>each subgroup. In addition, we implemented kernel methods in G-SBPS to improve </em><em>the propensity score model fitting. For each of the aim, we provided theoretical and </em><em>simulation-based justification for the research question or proposed methodologies, and </em><em>applied the proposed methods to the right heart catheterization (RHC) data to estimate </em><em>the length of hospital stay or the diabetes self-management training (DSMT) data to </em><em>evaluate the hospitalization rate within three years.</em></p>","abstract_html":"&lt;p&gt;&lt;em&gt; Establishing clear causality between exposures and outcomes is often complicated &lt;/em&gt;&lt;em&gt;by confounders in observational studies, which leads to imbalance in covariate distributions &lt;/em&gt;&lt;em&gt;between treatments and biased treatment effect inference. The propensity score &lt;/em&gt;&lt;em&gt;(PS) has been widely used to adjust this covariate imbalance in observational data. However, &lt;/em&gt;&lt;em&gt;the propensity score analysis methods rely on a correctly specified parametric PS &lt;/em&gt;&lt;em&gt;model. When the model is misspecified, the covariate imbalance may occur, which leads &lt;/em&gt;&lt;em&gt;to biased estimation of the treatment effect. Therefore, it is necessary to study how to &lt;/em&gt;&lt;em&gt;improve the model misspecification in propensity score analysis. &lt;/em&gt;&lt;/p&gt; &lt;p&gt;&lt;em&gt;&lt;/em&gt;&lt;em&gt; My Ph.D. dissertation consists of three aims. In Aim 1, we examined whether the &lt;/em&gt;&lt;em&gt;optimization of global balance — the mean balance of covariates or their transformations &lt;/em&gt;&lt;em&gt;in the overall study population, can circumvent the need for correct propensity score &lt;/em&gt;&lt;em&gt;model specification, and whether the use of a propensity score model further improves &lt;/em&gt;&lt;em&gt;the estimation performance compared to methods without modeling propensity score. &lt;/em&gt;&lt;em&gt;In Aim 2, we developed a propensity score analysis framework, the propensity score &lt;/em&gt;&lt;em&gt;with local balance (PSLB), which incorporates nonparametric propensity score models &lt;/em&gt;&lt;em&gt;and improves the balancing property of the estimated propensity score compared to &lt;/em&gt;&lt;em&gt;existing methods that only optimize the global balance. In Aim 3, we developed a &lt;/em&gt;&lt;em&gt;subgroup analysis method that is robust to propensity score model misspecification. &lt;/em&gt;&lt;em&gt;Specifically, we proposed a new algorithm, the guaranteed subgroup balancing propensity &lt;/em&gt;&lt;em&gt;score (G-SBPS), to ensure exact subgroup balance — the mean balance of covariates in &lt;/em&gt;&lt;em&gt;each subgroup. In addition, we implemented kernel methods in G-SBPS to improve &lt;/em&gt;&lt;em&gt;the propensity score model fitting. For each of the aim, we provided theoretical and &lt;/em&gt;&lt;em&gt;simulation-based justification for the research question or proposed methodologies, and &lt;/em&gt;&lt;em&gt;applied the proposed methods to the right heart catheterization (RHC) data to estimate &lt;/em&gt;&lt;em&gt;the length of hospital stay or the diabetes self-management training (DSMT) data to &lt;/em&gt;&lt;em&gt;evaluate the hospitalization rate within three years.&lt;/em&gt;&lt;/p&gt;","abstract_has_math":false,"creators":["Li, Yan","<p>https://orcid.org/0000-0001-9952-9737</p>"],"institution":null,"degree_name":"Doctor of Philosophy (PhD)","degree_level":"Thesis (MS)","degree_discipline":null,"degree_department":null,"school":null,"contributors":["Liang Li","Yu Shen","Ying Yuan"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2023,"date_issued":"2023-08-01T07:00:00Z","date_published":"2023-08-01T07:00:00Z","updated_at":"2026-07-24T05:48:59Z","subjects":["Observational study","Propensity score","Nonparametric modeling","Inverse probability weighting","Covariate balancing","Subgroup analysis","Biomedical Informatics"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://digitalcommons.library.tmc.edu/utgsbs_dissertations/1291","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Liang Li","Yu Shen","Ying Yuan"]},{"key":"dc:creator","label":"Author","values":["Li, Yan","<p>https://orcid.org/0000-0001-9952-9737</p>"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2024-07-25T07:00:00Z"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis (MS)"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctor of Philosophy (PhD)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Observational study","Propensity score","Nonparametric modeling","Inverse probability weighting","Covariate balancing","Subgroup analysis","Biomedical Informatics"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://digitalcommons.library.tmc.edu/utgsbs_dissertations/1291"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p><em> Establishing clear causality between exposures and outcomes is often complicated </em><em>by confounders in observational studies, which leads to imbalance in covariate distributions </em><em>between treatments and biased treatment effect inference. The propensity score </em><em>(PS) has been widely used to adjust this covariate imbalance in observational data. However, </em><em>the propensity score analysis methods rely on a correctly specified parametric PS </em><em>model. When the model is misspecified, the covariate imbalance may occur, which leads </em><em>to biased estimation of the treatment effect. Therefore, it is necessary to study how to </em><em>improve the model misspecification in propensity score analysis. </em></p> <p><em></em><em> My Ph.D. dissertation consists of three aims. In Aim 1, we examined whether the </em><em>optimization of global balance — the mean balance of covariates or their transformations </em><em>in the overall study population, can circumvent the need for correct propensity score </em><em>model specification, and whether the use of a propensity score model further improves </em><em>the estimation performance compared to methods without modeling propensity score. </em><em>In Aim 2, we developed a propensity score analysis framework, the propensity score </em><em>with local balance (PSLB), which incorporates nonparametric propensity score models </em><em>and improves the balancing property of the estimated propensity score compared to </em><em>existing methods that only optimize the global balance. In Aim 3, we developed a </em><em>subgroup analysis method that is robust to propensity score model misspecification. </em><em>Specifically, we proposed a new algorithm, the guaranteed subgroup balancing propensity </em><em>score (G-SBPS), to ensure exact subgroup balance — the mean balance of covariates in </em><em>each subgroup. In addition, we implemented kernel methods in G-SBPS to improve </em><em>the propensity score model fitting. For each of the aim, we provided theoretical and </em><em>simulation-based justification for the research question or proposed methodologies, and </em><em>applied the proposed methods to the right heart catheterization (RHC) data to estimate </em><em>the length of hospital stay or the diabetes self-management training (DSMT) data to </em><em>evaluate the hospitalization rate within three years.</em></p>"]},{"key":"dc:title","label":"Title","values":["Toward Nonparametric Propensity Score Estimation With Guaranteed Covariate Balance"]}]}],"canonical_facts":{"dc:contributor":["Liang Li","Yu Shen","Ying Yuan"],"dc:creator":["Li, Yan","<p>https://orcid.org/0000-0001-9952-9737</p>"],"dc:date.available":["2024-07-25T07:00:00Z"],"dc:description.abstract":["<p><em> Establishing clear causality between exposures and outcomes is often complicated </em><em>by confounders in observational studies, which leads to imbalance in covariate distributions </em><em>between treatments and biased treatment effect inference. The propensity score </em><em>(PS) has been widely used to adjust this covariate imbalance in observational data. However, </em><em>the propensity score analysis methods rely on a correctly specified parametric PS </em><em>model. When the model is misspecified, the covariate imbalance may occur, which leads </em><em>to biased estimation of the treatment effect. Therefore, it is necessary to study how to </em><em>improve the model misspecification in propensity score analysis. </em></p> <p><em></em><em> My Ph.D. dissertation consists of three aims. In Aim 1, we examined whether the </em><em>optimization of global balance — the mean balance of covariates or their transformations </em><em>in the overall study population, can circumvent the need for correct propensity score </em><em>model specification, and whether the use of a propensity score model further improves </em><em>the estimation performance compared to methods without modeling propensity score. </em><em>In Aim 2, we developed a propensity score analysis framework, the propensity score </em><em>with local balance (PSLB), which incorporates nonparametric propensity score models </em><em>and improves the balancing property of the estimated propensity score compared to </em><em>existing methods that only optimize the global balance. In Aim 3, we developed a </em><em>subgroup analysis method that is robust to propensity score model misspecification. </em><em>Specifically, we proposed a new algorithm, the guaranteed subgroup balancing propensity </em><em>score (G-SBPS), to ensure exact subgroup balance — the mean balance of covariates in </em><em>each subgroup. In addition, we implemented kernel methods in G-SBPS to improve </em><em>the propensity score model fitting. For each of the aim, we provided theoretical and </em><em>simulation-based justification for the research question or proposed methodologies, and </em><em>applied the proposed methods to the right heart catheterization (RHC) data to estimate </em><em>the length of hospital stay or the diabetes self-management training (DSMT) data to </em><em>evaluate the hospitalization rate within three years.</em></p>"],"dc:identifier":["https://digitalcommons.library.tmc.edu/utgsbs_dissertations/1291"],"dc:subject":["Observational study","Propensity score","Nonparametric modeling","Inverse probability weighting","Covariate balancing","Subgroup analysis","Biomedical Informatics"],"dc:title":["Toward Nonparametric Propensity Score Estimation With Guaranteed Covariate Balance"],"thesis:degree_level":["Thesis (MS)"],"thesis:degree_name":["Doctor of Philosophy (PhD)"]},"updated_at":"2026-07-24T05:48:59Z"}