Back to search

University of Illinois - Chicago

Variable Selection and Hypothesis Testing for High-Dimensional Models in Mental Health Research

Abstract

dc:description

Mental health research increasingly relies on complex, high-dimensional data arising from heterogeneous populations, longitudinal designs, hierarchical structures, and outcomes exhibiting excess zeros. These characteristics pose significant challenges for variable selection, model estimation, hypothesis testing, and sample size determination, often undermining statistical stability, interpretability, and computational efficiency. This dissertation develops and evaluates novel statistical methodologies aimed at improving inference in high-dimensional and structurally complex mental health data. The first component introduces a high-dimensional multinomial logistic regression framework incorporating combined L0 and L2 regularization to achieve exact variable selection with enhanced numerical stability. The proposed method integrates iteratively reweighted least squares with gradient scaling, gradient hard-thresholding to enforce sparsity, and systematic pairwise swapping to iteratively refine selected predictors. Simulation studies and a real-world genomics application demonstrate improved convergence, reduced false discoveries, and superior variable selection performance compared to conventional regularization approaches. The second component addresses estimation, hypothesis testing, and sample size determination for zero-inflated mixed models commonly encountered in mental health services research. A unified framework is developed using Laplace approximation and Newton–Raphson optimization for efficient estimation of fixed effects, random effects, and variance components. Hypothesis testing is conducted using Wald-based and parametric bootstrap methods, while power and sample size procedures are extended to accommodate zero inflation, hierarchical clustering, longitudinal correlation, and attrition. Simulation results show improved Type I error control, power estimation, and robustness across realistic study designs. Collectively, this dissertation advances statistical methodology for high-dimensional and zero-inflated mental health data, providing scalable and interpretable tools that enhance hypothesis testing, inference, and study design in modern public health research.

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Avisek Datta (13531380)

Subjects

dc:subject × 2

Rights

dc:rights
Statement dc:rights
  • In Copyright

Identifiers

dc:identifier.*
OAI identifier oai:identifier
oai:figshare.com:article/32991929

Chain of custody

source
Harvested from
University of Illinois - Chicago
Base URL
api.figshare.com/v2/oai
Last updated
2026-07-27
Source record
OAI-PMH GetRecord
citation

Avisek Datta (13531380). Variable Selection and Hypothesis Testing for High-Dimensional Models in Mental Health Research. 2026. https://doi.org/10.25417/uic.32991929.v1