Helsingin yliopisto
Evaluating the Role of Balanced Causal and Non-Causal Features in Predictive Modeling
Abstract
dc:description.abstractDeveloping Machine learning models that provides a stable performance across domains unseen during training, remains a persistent challenge. This generalizability of models is particularly challenging when the relationship between input features and target labels varies across domains. This study is motivated by recent work suggesting that causal features generalize better across domains as opposed to non-causal features. However, the evaluations comprised of varying sizes of input information across the models being compared. This study addresses this limitation by balancing the number of input features across all models, thereby improving the fairness of the comparison. Experiments were conducted on four real-world tabular datasets with different application domains - ANES(Voting), ASSISTments, BRFSS (Diabetes), and NHANES (Blood Lead). The evaluations utilized tree-based classifiers, MLPs and domain generalization approaches, including GroupDRO and REx. Each model was trained on data from one domain and evaluated on another unseen test domain to assess robustness under a strictly binary domain shift. Results show that models trained on causal features achieved more stable generalization, particularly in datasets with strong causal relationships to the target label. Tree-based models consistently outperformed MLPs. Despite the strictly binary experimental setup, domain generalization methods that are designed for group settings were implemented for the purpose of exploration. As expected, Rex and GroupDRO performed poorly in this setting. Interpretability analyses using SHAP and LIME were done on all datasets, in an XGBoost model, which offered insights on the level of influence the features have to the target, which supported the results. Overall, this work investigates the role of causal features in enhancing robustness under domain shift, by implementing a fair, interpretable experiment pipeline across real-world datasets.
Degree
thesis:*- Grantor dc:publisher
- Helsingin yliopisto
- Year dc:date.issued
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Subramanian, Janani
Subjects
dc:subject × 4Rights
dc:rights- Statement dc:rights
-
- In Copyright 1.0
- Language dc:language.iso
- eng
Identifiers
dc:identifier.*- Handle dc:identifier.uri
- http://hdl.handle.net/10138/601034
- OAI identifier oai:identifier
- oai:helda.helsinki.fi:10138/601034