Back to results

University of Illinois - Chicago

Optimizing Machine Learning Performance on Tabular Clinical Data: A Pipeline Approach

Abstract

dc:description

This MSc thesis addresses the pivotal challenge of developing robust and effective machine learning solutions for mixed-type tabular clinical data. Such datasets, prevalent in healthcare, are characterized by a complex interplay of numerical and categorical features, often leading to intricate, non-linear dependencies that hinder traditional analysis and model efficacy. Confronting inherent issues like data scarcity, imbalance, noise, and sensitivity, this work introduces a comprehensive, adaptable machine learning pipeline engineered for optimal predictive performance. A distinctive feature of this pipeline is its integral use of state-of-the-art generative models, specifically Tabsyn, to strategically augment and synthesize data. This innovation tackles limitations posed by insufficient real-world samples, thereby enhancing model robustness and generalizability. Beyond data generation, the pipeline systematically incorporates advanced preprocessing (imputation, encoding, scaling), rigorous feature engineering, judicious model selection across a diverse suite of algorithms (including MLPs, SVMs, Decision Trees, Logistic Regression, Random Forests, and XGBoost), and sophisticated Bayesian hyperparameter optimization. The pipeline's dual effectiveness, both in achieving high predictive accuracy and exposing fundamental challenges, is rigorously demonstrated across two distinct clinical tabular datasets. While proving highly successful in optimizing models for well-structured data, the research critically illustrates a "ceiling effect" when applied to datasets with inherently less informative features. This empirical finding underscores that even the most advanced models within an exhaustively optimized framework cannot overcome the limitations imposed by poor feature quality, emphasizing that domain expertise and meticulous data curation are paramount to real-world clinical utility. In conclusion, this thesis advocates for a holistic machine learning strategy in healthcare, where algorithmic sophistication, empowered by synthetic data generation, is inextricably linked with the integrity of data preparation and the relevance of mixed-type features. The validated pipeline and insights provided offer a foundational framework for advancing precision and reliability in clinical decision-making.

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Davide Salonico (24399683)

Subjects

dc:subject × 1

Rights

dc:rights
Statement dc:rights
  • In Copyright

Identifiers

dc:identifier.*
OAI identifier oai:identifier
oai:figshare.com:article/32994122

Chain of custody

source
Harvested from
University of Illinois - Chicago
Base URL
api.figshare.com/v2/oai
Last updated
2026-07-27
Source record
OAI-PMH GetRecord
related terms
citation

Davide Salonico (24399683). Optimizing Machine Learning Performance on Tabular Clinical Data: A Pipeline Approach. 2026. https://doi.org/10.25417/uic.32994122.v1