{"id":{"repo_id":"temple","oai_identifier":"oai:scholarshare.temple.edu:20.500.12613/11133"},"canonical_url":"https://search.dev.ndltd.org/etd/temple/oai:scholarshare.temple.edu:20.500.12613/11133","repository":{"repo_id":"temple","name":"Temple University","base_url":"https://scholarshare.temple.edu/server/oai/request"},"display":{"title":"P-VALUE BASED VARIABLE SELECTION FOR GENERALIZED LINEAR MODELS","abstract":"This thesis presents two new p-value based methods for variable selection in generalized linear models. Generalized linear models are widely used, but their non-analytic solutions and intricate dependencies create challenges for many existing methods. Addressing these issues, our proposed contributions can select important variables for generalized linear models and control the false discovery rate under arbitrary covariance structure. Both approaches ultimately make selections via the two step multiple testing procedure of Sarkar and Tang (2022), but they differ in how they construct the necessary p-values. Our first proposed method, the generalized two step, creates appropriate p-values by solving a non-trivial linear quadratic equation to construct a data-adaptive linear transformation. This transformation is then applied to an initial fitted model to create paired estimates of the unknown true coefficient vector. Our second proposed method, the weighted two step, adjusts the original response vector and design matrix to reframe generalized linear models as homoskedastic linear regressions. After reweighting, the methods of Sarkar and Tang (2022) can be more directly applied to the generalized setting. We develop mathematical theory to prove that our contributions control the false discovery rate, and empirical evaluations demonstrate their promising performance across diverse simulation settings and ten datasets.","abstract_html":"This thesis presents two new p-value based methods for variable selection in generalized linear models. Generalized linear models are widely used, but their non-analytic solutions and intricate dependencies create challenges for many existing methods. Addressing these issues, our proposed contributions can select important variables for generalized linear models and control the false discovery rate under arbitrary covariance structure. Both approaches ultimately make selections via the two step multiple testing procedure of Sarkar and Tang (2022), but they differ in how they construct the necessary p-values. Our first proposed method, the generalized two step, creates appropriate p-values by solving a non-trivial linear quadratic equation to construct a data-adaptive linear transformation. This transformation is then applied to an initial fitted model to create paired estimates of the unknown true coefficient vector. Our second proposed method, the weighted two step, adjusts the original response vector and design matrix to reframe generalized linear models as homoskedastic linear regressions. After reweighting, the methods of Sarkar and Tang (2022) can be more directly applied to the generalized setting. We develop mathematical theory to prove that our contributions control the false discovery rate, and empirical evaluations demonstrate their promising performance across diverse simulation settings and ten datasets.","abstract_has_math":false,"creators":["Rilling, Joseph"],"institution":"Temple University. Libraries","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Tang, Cheng Yong"],"committee_chairs":[],"committee_members":["McAlinn, Kenichiro","Lee, Kuang-Yao","Yang, Yang"],"year":2025,"date_issued":"2025-05","date_published":"2025-05","updated_at":"2026-07-27T21:21:02Z","subjects":["Statistics","Hochberg, Benjamini","Generalized linear models","Multiple testing","P-value"],"languages":["eng"],"rights":["IN COPYRIGHT- This Rights Statement can be used for an Item that is in copyright. Using this statement implies that the organization making this Item available has determined that the Item is in copyright and either is the rights-holder, has obtained permission from the rights-holder(s) to make their Work(s) available, or makes the Item available under an exception or limitation to copyright (including Fair Use) that entitles it to make the Item available."],"rights_urls":["http://rightsstatements.org/vocab/InC/1.0/"],"identifier_entries":[]},"links":{"outbound_url":"https://scholarshare.temple.edu/handle/20.500.12613/11133","outbound_label":"Repository record","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Tang, Cheng Yong"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["McAlinn, Kenichiro","Lee, Kuang-Yao","Yang, Yang"]},{"key":"dc:creator","label":"Author","values":["Rilling, Joseph"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2025-07-21T19:16:38Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2025-07-21T19:16:38Z"]},{"key":"dc:date.issued","label":"Date","values":["2025-05"]},{"key":"dc:publisher","label":"Institution","values":["Temple University. Libraries"]},{"key":"dc:type","label":"Dc Type","values":["Text"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Statistics","Hochberg, Benjamini","Generalized linear models","Multiple testing","P-value"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["IN COPYRIGHT- This Rights Statement can be used for an Item that is in copyright. Using this statement implies that the organization making this Item available has determined that the Item is in copyright and either is the rights-holder, has obtained permission from the rights-holder(s) to make their Work(s) available, or makes the Item available under an exception or limitation to copyright (including Fair Use) that entitles it to make the Item available."]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://rightsstatements.org/vocab/InC/1.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://scholarshare.temple.edu/handle/20.500.12613/11133"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["This thesis presents two new p-value based methods for variable selection in generalized linear models. Generalized linear models are widely used, but their non-analytic solutions and intricate dependencies create challenges for many existing methods. Addressing these issues, our proposed contributions can select important variables for generalized linear models and control the false discovery rate under arbitrary covariance structure. Both approaches ultimately make selections via the two step multiple testing procedure of Sarkar and Tang (2022), but they differ in how they construct the necessary p-values. Our first proposed method, the generalized two step, creates appropriate p-values by solving a non-trivial linear quadratic equation to construct a data-adaptive linear transformation. This transformation is then applied to an initial fitted model to create paired estimates of the unknown true coefficient vector. Our second proposed method, the weighted two step, adjusts the original response vector and design matrix to reframe generalized linear models as homoskedastic linear regressions. After reweighting, the methods of Sarkar and Tang (2022) can be more directly applied to the generalized setting. We develop mathematical theory to prove that our contributions control the false discovery rate, and empirical evaluations demonstrate their promising performance across diverse simulation settings and ten datasets."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["Ph.D."]},{"key":"dc:title","label":"Title","values":["P-VALUE BASED VARIABLE SELECTION FOR GENERALIZED LINEAR MODELS"]}]}],"canonical_facts":{"dc:contributor.advisor":["Tang, Cheng Yong"],"dc:contributor.committeemember":["McAlinn, Kenichiro","Lee, Kuang-Yao","Yang, Yang"],"dc:creator":["Rilling, Joseph"],"dc:date.accessioned":["2025-07-21T19:16:38Z"],"dc:date.available":["2025-07-21T19:16:38Z"],"dc:date.issued":["2025-05"],"dc:description.abstract":["This thesis presents two new p-value based methods for variable selection in generalized linear models. Generalized linear models are widely used, but their non-analytic solutions and intricate dependencies create challenges for many existing methods. Addressing these issues, our proposed contributions can select important variables for generalized linear models and control the false discovery rate under arbitrary covariance structure. Both approaches ultimately make selections via the two step multiple testing procedure of Sarkar and Tang (2022), but they differ in how they construct the necessary p-values. Our first proposed method, the generalized two step, creates appropriate p-values by solving a non-trivial linear quadratic equation to construct a data-adaptive linear transformation. This transformation is then applied to an initial fitted model to create paired estimates of the unknown true coefficient vector. Our second proposed method, the weighted two step, adjusts the original response vector and design matrix to reframe generalized linear models as homoskedastic linear regressions. After reweighting, the methods of Sarkar and Tang (2022) can be more directly applied to the generalized setting. We develop mathematical theory to prove that our contributions control the false discovery rate, and empirical evaluations demonstrate their promising performance across diverse simulation settings and ten datasets."],"dc:description.degree":["Ph.D."],"dc:identifier.uri":["https://scholarshare.temple.edu/handle/20.500.12613/11133"],"dc:language.iso":["eng"],"dc:publisher":["Temple University. Libraries"],"dc:rights":["IN COPYRIGHT- This Rights Statement can be used for an Item that is in copyright. Using this statement implies that the organization making this Item available has determined that the Item is in copyright and either is the rights-holder, has obtained permission from the rights-holder(s) to make their Work(s) available, or makes the Item available under an exception or limitation to copyright (including Fair Use) that entitles it to make the Item available."],"dc:rights.uri":["http://rightsstatements.org/vocab/InC/1.0/"],"dc:subject":["Statistics","Hochberg, Benjamini","Generalized linear models","Multiple testing","P-value"],"dc:title":["P-VALUE BASED VARIABLE SELECTION FOR GENERALIZED LINEAR MODELS"],"dc:type":["Text"]},"updated_at":"2026-07-27T21:21:02Z"}