{"id":{"repo_id":"umkc","oai_identifier":"oai:mospace.umsystem.edu:10355/110342"},"canonical_url":"https://search.dev.ndltd.org/etd/umkc/oai:mospace.umsystem.edu:10355/110342","repository":{"repo_id":"umkc","name":"University of Missouri - Kansas City","base_url":"https://mospace.umsystem.edu/oai/request"},"display":{"title":"Methods for imbalanced data in sports analytics: improving injury prediction models","abstract":"With the unprecedented growth of sports streaming and the increasing use of machine learning insports analytics, a more in-depth understanding of real-world datasets has become essential, along with methods that can handle noisy data that degrade predictive accuracy. Sports injuries are a central concern because an injury can permanently derail an athlete’s career. Yet even with extensive daily training records, injuries remain rare events, and many existing injury prediction models struggle to identify injury cases accurately. This research evaluates both real-world sports injury datasets and artificial datasets to develop and illustrate a practical framework for modeling and evaluating prediction under extreme class imbalance, rather than to identify a single best-performing classifier. Experiments use XGBoost as a standard baseline and examine how methodological choices affect model behavior, including regularization with early stopping to control overfitting, alternative outcome-generation schemes for synthetic data, decision-rule restructuring through constructed risk-score ensembles, and incremental increases in injury prevalence to study how discrimination changes as the base rate shifts. Results indicate that regularized XGBoost with early stopping reduces overfitting, but injury-case prediction remains limited under extreme class imbalance and weak signal. Combining multiple prediction model forms and systematically increasing injury prevalence improves AUC, although injury prediction accuracy still requires further improvement. Overall, the findings aim to guide future work on methodology choices for imbalanced sports datasets and support more effective injury-risk screening and interpretation for athletes.","abstract_html":"With the unprecedented growth of sports streaming and the increasing use of machine learning insports analytics, a more in-depth understanding of real-world datasets has become essential, along with methods that can handle noisy data that degrade predictive accuracy. Sports injuries are a central concern because an injury can permanently derail an athlete’s career. Yet even with extensive daily training records, injuries remain rare events, and many existing injury prediction models struggle to identify injury cases accurately. This research evaluates both real-world sports injury datasets and artificial datasets to develop and illustrate a practical framework for modeling and evaluating prediction under extreme class imbalance, rather than to identify a single best-performing classifier. Experiments use XGBoost as a standard baseline and examine how methodological choices affect model behavior, including regularization with early stopping to control overfitting, alternative outcome-generation schemes for synthetic data, decision-rule restructuring through constructed risk-score ensembles, and incremental increases in injury prevalence to study how discrimination changes as the base rate shifts. Results indicate that regularized XGBoost with early stopping reduces overfitting, but injury-case prediction remains limited under extreme class imbalance and weak signal. Combining multiple prediction model forms and systematically increasing injury prevalence improves AUC, although injury prediction accuracy still requires further improvement. Overall, the findings aim to guide future work on methodology choices for imbalanced sports datasets and support more effective injury-risk screening and interpretation for athletes.","abstract_has_math":false,"creators":["Hu, Di (Graduate student at University of Missouri--Kansas City)"],"institution":"University of Missouri--Kansas City","degree_name":"M.S. (Master of Science)","degree_level":"Masters","degree_discipline":"Mathematics (UMKC)","degree_department":null,"school":null,"contributors":[],"advisors":["Cao, Shuhao"],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025","date_published":"2025","updated_at":"2026-07-24T05:19:15Z","subjects":[],"languages":["en_US"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/10355/110342","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Cao, Shuhao"]},{"key":"dc:creator","label":"Author","values":["Hu, Di (Graduate student at University of Missouri--Kansas City)"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2026-01-26T20:03:15Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2026-01-26T20:03:15Z"]},{"key":"dc:date.issued","label":"Date","values":["2025"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Mathematics (UMKC)","Statistics (UMKC)"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Masters"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S. (Master of Science)"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Missouri--Kansas City"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en_US"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10355/110342"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Title from PDF of title page viewed January 27, 2026","Thesis advisor: Shuhao Cao","Vita","Includes bibliographical references (pages 38-39)","Thesis (M.S.)--Department of Mathematics and Statistics. University of Missouri--Kansas City, 2025"]},{"key":"dc:description.abstract","label":"Abstract","values":["With the unprecedented growth of sports streaming and the increasing use of machine learning insports analytics, a more in-depth understanding of real-world datasets has become essential, along with methods that can handle noisy data that degrade predictive accuracy. Sports injuries are a central concern because an injury can permanently derail an athlete’s career. Yet even with extensive daily training records, injuries remain rare events, and many existing injury prediction models struggle to identify injury cases accurately. This research evaluates both real-world sports injury datasets and artificial datasets to develop and illustrate a practical framework for modeling and evaluating prediction under extreme class imbalance, rather than to identify a single best-performing classifier. Experiments use XGBoost as a standard baseline and examine how methodological choices affect model behavior, including regularization with early stopping to control overfitting, alternative outcome-generation schemes for synthetic data, decision-rule restructuring through constructed risk-score ensembles, and incremental increases in injury prevalence to study how discrimination changes as the base rate shifts. Results indicate that regularized XGBoost with early stopping reduces overfitting, but injury-case prediction remains limited under extreme class imbalance and weak signal. Combining multiple prediction model forms and systematically increasing injury prevalence improves AUC, although injury prediction accuracy still requires further improvement. Overall, the findings aim to guide future work on methodology choices for imbalanced sports datasets and support more effective injury-risk screening and interpretation for athletes."]},{"key":"dc:title","label":"Title","values":["Methods for imbalanced data in sports analytics: improving injury prediction models"]}]}],"canonical_facts":{"dc:contributor.advisor":["Cao, Shuhao"],"dc:creator":["Hu, Di (Graduate student at University of Missouri--Kansas City)"],"dc:date.accessioned":["2026-01-26T20:03:15Z"],"dc:date.available":["2026-01-26T20:03:15Z"],"dc:date.issued":["2025"],"dc:description":["Title from PDF of title page viewed January 27, 2026","Thesis advisor: Shuhao Cao","Vita","Includes bibliographical references (pages 38-39)","Thesis (M.S.)--Department of Mathematics and Statistics. University of Missouri--Kansas City, 2025"],"dc:description.abstract":["With the unprecedented growth of sports streaming and the increasing use of machine learning insports analytics, a more in-depth understanding of real-world datasets has become essential, along with methods that can handle noisy data that degrade predictive accuracy. Sports injuries are a central concern because an injury can permanently derail an athlete’s career. Yet even with extensive daily training records, injuries remain rare events, and many existing injury prediction models struggle to identify injury cases accurately. This research evaluates both real-world sports injury datasets and artificial datasets to develop and illustrate a practical framework for modeling and evaluating prediction under extreme class imbalance, rather than to identify a single best-performing classifier. Experiments use XGBoost as a standard baseline and examine how methodological choices affect model behavior, including regularization with early stopping to control overfitting, alternative outcome-generation schemes for synthetic data, decision-rule restructuring through constructed risk-score ensembles, and incremental increases in injury prevalence to study how discrimination changes as the base rate shifts. Results indicate that regularized XGBoost with early stopping reduces overfitting, but injury-case prediction remains limited under extreme class imbalance and weak signal. Combining multiple prediction model forms and systematically increasing injury prevalence improves AUC, although injury prediction accuracy still requires further improvement. Overall, the findings aim to guide future work on methodology choices for imbalanced sports datasets and support more effective injury-risk screening and interpretation for athletes."],"dc:identifier.uri":["https://hdl.handle.net/10355/110342"],"dc:language.iso":["en_US"],"dc:title":["Methods for imbalanced data in sports analytics: improving injury prediction models"],"dc:type":["Thesis"],"thesis:degree_discipline":["Mathematics (UMKC)","Statistics (UMKC)"],"thesis:degree_level":["Masters"],"thesis:degree_name":["M.S. (Master of Science)"],"thesis:institution_name":["University of Missouri--Kansas City"]},"updated_at":"2026-07-24T05:19:15Z"}