{"id":{"repo_id":"uic","oai_identifier":"oai:figshare.com:article/31451836"},"canonical_url":"https://search.dev.ndltd.org/etd/uic/oai:figshare.com:article/31451836","repository":{"repo_id":"uic","name":"University of Illinois - Chicago","base_url":"https://api.figshare.com/v2/oai"},"display":{"title":"Imitation Learning with Superhuman Policy Gradient Optimization for Sequential Cancer Treatment Decisions","abstract":"We propose a simulator-driven imitation learning framework for sequential deci- sion making in head and neck cancer (HNC) treatment. Our method, Superhu- man Policy Gradient Optimization (SPGO), integrates inverse reinforcement learning principles with policy gradient updates to derive three-stage treat- ment policies directly from recorded physician decisions. A pre-trained clinical simulator—combining a variational autoencoder (VAE) and gradient boosting (XGBoost) models—generates complete, temporally consistent patient trajec- tories, enabling safe and reproducible training. Unlike conventional behavior cloning, SPGO optimizes a subdominance loss that explicitly rewards surpassing the expert across multiple clinical outcomes, including relapse at year three and patient-reported toxicities at multiple follow- up times. We systematically compare six subdominance configurations (absolute vs. relative, sum vs. max aggregation, per-feature vs. max-only α updates) to assess how loss design affects convergence and treatment quality. Our best configuration—relative differences with sum aggregation and per- feature α updates—achieves over 70% superhuman dominance across clinically relevant features on held-out patients. The learned policies reproduce expert decisions on acute measures while significantly reducing predicted late toxicities and relapse risk, demonstrating generalization beyond the training distribution.","abstract_html":"We propose a simulator-driven imitation learning framework for sequential deci- sion making in head and neck cancer (HNC) treatment. Our method, Superhu- man Policy Gradient Optimization (SPGO), integrates inverse reinforcement learning principles with policy gradient updates to derive three-stage treat- ment policies directly from recorded physician decisions. A pre-trained clinical simulator—combining a variational autoencoder (VAE) and gradient boosting (XGBoost) models—generates complete, temporally consistent patient trajec- tories, enabling safe and reproducible training. Unlike conventional behavior cloning, SPGO optimizes a subdominance loss that explicitly rewards surpassing the expert across multiple clinical outcomes, including relapse at year three and patient-reported toxicities at multiple follow- up times. We systematically compare six subdominance configurations (absolute vs. relative, sum vs. max aggregation, per-feature vs. max-only α updates) to assess how loss design affects convergence and treatment quality. Our best configuration—relative differences with sum aggregation and per- feature α updates—achieves over 70% superhuman dominance across clinically relevant features on held-out patients. The learned policies reproduce expert decisions on acute measures while significantly reducing predicted late toxicities and relapse risk, demonstrating generalization beyond the training distribution.","abstract_has_math":false,"creators":["Filippo Corna (23292115)"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-12-01T00:00:00Z","date_published":"2025-12-01T00:00:00Z","updated_at":"2026-07-27T21:34:33Z","subjects":["Artificial Intelligence"],"languages":[],"rights":["In Copyright","Open Access after 2028-01-01"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://doi.org/10.25417/uic.31451836.v1","outbound_label":"DOI","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["Filippo Corna (23292115)"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-12-01T00:00:00Z"]},{"key":"dc:relation","label":"Dc Relation","values":["https://figshare.com/articles/thesis/Imitation_Learning_with_Superhuman_Policy_Gradient_Optimization_for_Sequential_Cancer_Treatment_Decisions/31451836"]},{"key":"dc:type","label":"Dc Type","values":["Text","Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Artificial Intelligence"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["In Copyright","Open Access after 2028-01-01"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["10.25417/uic.31451836.v1"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["We propose a simulator-driven imitation learning framework for sequential deci- sion making in head and neck cancer (HNC) treatment. Our method, Superhu- man Policy Gradient Optimization (SPGO), integrates inverse reinforcement learning principles with policy gradient updates to derive three-stage treat- ment policies directly from recorded physician decisions. A pre-trained clinical simulator—combining a variational autoencoder (VAE) and gradient boosting (XGBoost) models—generates complete, temporally consistent patient trajec- tories, enabling safe and reproducible training. Unlike conventional behavior cloning, SPGO optimizes a subdominance loss that explicitly rewards surpassing the expert across multiple clinical outcomes, including relapse at year three and patient-reported toxicities at multiple follow- up times. We systematically compare six subdominance configurations (absolute vs. relative, sum vs. max aggregation, per-feature vs. max-only α updates) to assess how loss design affects convergence and treatment quality. Our best configuration—relative differences with sum aggregation and per- feature α updates—achieves over 70% superhuman dominance across clinically relevant features on held-out patients. The learned policies reproduce expert decisions on acute measures while significantly reducing predicted late toxicities and relapse risk, demonstrating generalization beyond the training distribution."]},{"key":"dc:title","label":"Title","values":["Imitation Learning with Superhuman Policy Gradient Optimization for Sequential Cancer Treatment Decisions"]}]}],"canonical_facts":{"dc:creator":["Filippo Corna (23292115)"],"dc:date":["2025-12-01T00:00:00Z"],"dc:description":["We propose a simulator-driven imitation learning framework for sequential deci- sion making in head and neck cancer (HNC) treatment. Our method, Superhu- man Policy Gradient Optimization (SPGO), integrates inverse reinforcement learning principles with policy gradient updates to derive three-stage treat- ment policies directly from recorded physician decisions. A pre-trained clinical simulator—combining a variational autoencoder (VAE) and gradient boosting (XGBoost) models—generates complete, temporally consistent patient trajec- tories, enabling safe and reproducible training. Unlike conventional behavior cloning, SPGO optimizes a subdominance loss that explicitly rewards surpassing the expert across multiple clinical outcomes, including relapse at year three and patient-reported toxicities at multiple follow- up times. We systematically compare six subdominance configurations (absolute vs. relative, sum vs. max aggregation, per-feature vs. max-only α updates) to assess how loss design affects convergence and treatment quality. Our best configuration—relative differences with sum aggregation and per- feature α updates—achieves over 70% superhuman dominance across clinically relevant features on held-out patients. The learned policies reproduce expert decisions on acute measures while significantly reducing predicted late toxicities and relapse risk, demonstrating generalization beyond the training distribution."],"dc:identifier":["10.25417/uic.31451836.v1"],"dc:relation":["https://figshare.com/articles/thesis/Imitation_Learning_with_Superhuman_Policy_Gradient_Optimization_for_Sequential_Cancer_Treatment_Decisions/31451836"],"dc:rights":["In Copyright","Open Access after 2028-01-01"],"dc:subject":["Artificial Intelligence"],"dc:title":["Imitation Learning with Superhuman Policy Gradient Optimization for Sequential Cancer Treatment Decisions"],"dc:type":["Text","Thesis"]},"updated_at":"2026-07-27T21:34:33Z"}