University of Illinois - Chicago
Imitation Learning with Superhuman Policy Gradient Optimization for Sequential Cancer Treatment Decisions
Abstract
dc:descriptionWe propose a simulator-driven imitation learning framework for sequential deci- sion making in head and neck cancer (HNC) treatment. Our method, Superhu- man Policy Gradient Optimization (SPGO), integrates inverse reinforcement learning principles with policy gradient updates to derive three-stage treat- ment policies directly from recorded physician decisions. A pre-trained clinical simulator—combining a variational autoencoder (VAE) and gradient boosting (XGBoost) models—generates complete, temporally consistent patient trajec- tories, enabling safe and reproducible training. Unlike conventional behavior cloning, SPGO optimizes a subdominance loss that explicitly rewards surpassing the expert across multiple clinical outcomes, including relapse at year three and patient-reported toxicities at multiple follow- up times. We systematically compare six subdominance configurations (absolute vs. relative, sum vs. max aggregation, per-feature vs. max-only α updates) to assess how loss design affects convergence and treatment quality. Our best configuration—relative differences with sum aggregation and per- feature α updates—achieves over 70% superhuman dominance across clinically relevant features on held-out patients. The learned policies reproduce expert decisions on acute measures while significantly reducing predicted late toxicities and relapse risk, demonstrating generalization beyond the training distribution.
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Filippo Corna (23292115)
Subjects
dc:subject × 1Rights
dc:rights- Statement dc:rights
-
- In Copyright
- Open Access after 2028-01-01
Identifiers
dc:identifier.*- DOI dc:identifier
- https://doi.org/10.25417/uic.31451836.v1
- OAI identifier oai:identifier
- oai:figshare.com:article/31451836