University of Illinois - Chicago
Optimizing Machine Learning Performance on Tabular Clinical Data: A Pipeline Approach
Abstract
dc:descriptionThis MSc thesis addresses the pivotal challenge of developing robust and effective machine learning solutions for mixed-type tabular clinical data. Such datasets, prevalent in healthcare, are characterized by a complex interplay of numerical and categorical features, often leading to intricate, non-linear dependencies that hinder traditional analysis and model efficacy. Confronting inherent issues like data scarcity, imbalance, noise, and sensitivity, this work introduces a comprehensive, adaptable machine learning pipeline engineered for optimal predictive performance. A distinctive feature of this pipeline is its integral use of state-of-the-art generative models, specifically Tabsyn, to strategically augment and synthesize data. This innovation tackles limitations posed by insufficient real-world samples, thereby enhancing model robustness and generalizability. Beyond data generation, the pipeline systematically incorporates advanced preprocessing (imputation, encoding, scaling), rigorous feature engineering, judicious model selection across a diverse suite of algorithms (including MLPs, SVMs, Decision Trees, Logistic Regression, Random Forests, and XGBoost), and sophisticated Bayesian hyperparameter optimization. The pipeline's dual effectiveness, both in achieving high predictive accuracy and exposing fundamental challenges, is rigorously demonstrated across two distinct clinical tabular datasets. While proving highly successful in optimizing models for well-structured data, the research critically illustrates a "ceiling effect" when applied to datasets with inherently less informative features. This empirical finding underscores that even the most advanced models within an exhaustively optimized framework cannot overcome the limitations imposed by poor feature quality, emphasizing that domain expertise and meticulous data curation are paramount to real-world clinical utility. In conclusion, this thesis advocates for a holistic machine learning strategy in healthcare, where algorithmic sophistication, empowered by synthetic data generation, is inextricably linked with the integrity of data preparation and the relevance of mixed-type features. The validated pipeline and insights provided offer a foundational framework for advancing precision and reliability in clinical decision-making.
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Davide Salonico (24399683)
Subjects
dc:subject × 1Rights
dc:rights- Statement dc:rights
-
- In Copyright
Identifiers
dc:identifier.*- DOI dc:identifier
- https://doi.org/10.25417/uic.32994122.v1
- OAI identifier oai:identifier
- oai:figshare.com:article/32994122