{"id":{"repo_id":"venda","oai_identifier":"oai:univendspace.univen.ac.za:11602/3028"},"canonical_url":"https://search.dev.ndltd.org/etd/venda/oai:univendspace.univen.ac.za:11602/3028","repository":{"repo_id":"venda","name":"University of Venda","base_url":"https://univendspace.univen.ac.za/server/oai/request"},"display":{"title":"Comparative Analysis of Discrimination and Calibration Accuracy of Discrete Survival, Random Forests, and Neural Networks in Health-Related Survival Prediction Models","abstract":"Prediction models for survival analysis are commonly used in biomedical sciences to understand the onset of certain diseases. Traditional statistical models have been employed for the previous years, however, their limitations and inability to handle big data sets has made a way for the introduction of machine learning methods which gained recognition due to their ability to learn complex algorithms. However, existing literature indicates that the predictive accuracy of machine learning and statistical models for survival analysis varies significantly across different data sets. This variability underscores the need for further research utilizing data sets with diverse characteristics. Such research is essential to develop generalizable insights into the conditions under which each method performs best. In this research project, we compared the predictive performance of traditional statistical method and machine learning algorithms in discrete survival analysis. The machine learning methods include discrete-time survival trees, discrete-time random survival forests, and discrete-time neural networks. The study uses calibration (measured by the prediction error curves) to assess model fit and discrimination (measured by the Concordance index and area under curve) to evaluate predictive accuracy. These methods were applied to data sets: Breast cancer, age at first alcohol intake and CRASH-2. The discrete-time neural network had the best prediction performance as compared to the rest of the models for survival of breast cancer. The discrete-time random forest with hellinger distance had the overall prediction performance on the age at first alcohol intake. The discrete-time survival model outperformed the rest of the models in predicting survival of bleeding trauma patients from the CRASH-2 data .","abstract_html":"Prediction models for survival analysis are commonly used in biomedical sciences to understand the onset of certain diseases. Traditional statistical models have been employed for the previous years, however, their limitations and inability to handle big data sets has made a way for the introduction of machine learning methods which gained recognition due to their ability to learn complex algorithms. However, existing literature indicates that the predictive accuracy of machine learning and statistical models for survival analysis varies significantly across different data sets. This variability underscores the need for further research utilizing data sets with diverse characteristics. Such research is essential to develop generalizable insights into the conditions under which each method performs best. In this research project, we compared the predictive performance of traditional statistical method and machine learning algorithms in discrete survival analysis. The machine learning methods include discrete-time survival trees, discrete-time random survival forests, and discrete-time neural networks. The study uses calibration (measured by the prediction error curves) to assess model fit and discrimination (measured by the Concordance index and area under curve) to evaluate predictive accuracy. These methods were applied to data sets: Breast cancer, age at first alcohol intake and CRASH-2. The discrete-time neural network had the best prediction performance as compared to the rest of the models for survival of breast cancer. The discrete-time random forest with hellinger distance had the overall prediction performance on the age at first alcohol intake. The discrete-time survival model outperformed the rest of the models in predicting survival of bleeding trauma patients from the CRASH-2 data .","abstract_has_math":false,"creators":["Ramachela, Audrey Tshepho"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Bere, Alphonce","Mulaudzi, Tshilidzi","Motsuku, Lactatia"],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-09-05","date_published":"2025-09-05","updated_at":"2026-07-27T21:57:28Z","subjects":["Discrete-time survival analysis","Statistical methods","Machine learning","Calibration","Discrimination","UCTD"],"languages":["en"],"rights":["University of Venda"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://univendspace.univen.ac.za/handle/11602/3028","outbound_label":"Repository record","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Bere, Alphonce","Mulaudzi, Tshilidzi","Motsuku, Lactatia"]},{"key":"dc:creator","label":"Author","values":["Ramachela, Audrey Tshepho"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025"]},{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2025-11-07T05:34:50Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2025-11-07T05:34:50Z"]},{"key":"dc:date.issued","label":"Date","values":["2025-09-05"]},{"key":"dc:relation","label":"Dc Relation","values":["PDF"]},{"key":"dc:type","label":"Dc Type","values":["Dissertation"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Discrete-time survival analysis","Statistical methods","Machine learning","Calibration","Discrimination","UCTD"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["University of Venda"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://univendspace.univen.ac.za/handle/11602/3028"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["MSc in Statistics","Department of Mathematical and Computational Sciences"]},{"key":"dc:description.abstract","label":"Abstract","values":["Prediction models for survival analysis are commonly used in biomedical sciences to understand the onset of certain diseases. Traditional statistical models have been employed for the previous years, however, their limitations and inability to handle big data sets has made a way for the introduction of machine learning methods which gained recognition due to their ability to learn complex algorithms. However, existing literature indicates that the predictive accuracy of machine learning and statistical models for survival analysis varies significantly across different data sets. This variability underscores the need for further research utilizing data sets with diverse characteristics. Such research is essential to develop generalizable insights into the conditions under which each method performs best. In this research project, we compared the predictive performance of traditional statistical method and machine learning algorithms in discrete survival analysis. The machine learning methods include discrete-time survival trees, discrete-time random survival forests, and discrete-time neural networks. The study uses calibration (measured by the prediction error curves) to assess model fit and discrimination (measured by the Concordance index and area under curve) to evaluate predictive accuracy. These methods were applied to data sets: Breast cancer, age at first alcohol intake and CRASH-2. The discrete-time neural network had the best prediction performance as compared to the rest of the models for survival of breast cancer. The discrete-time random forest with hellinger distance had the overall prediction performance on the age at first alcohol intake. The discrete-time survival model outperformed the rest of the models in predicting survival of bleeding trauma patients from the CRASH-2 data ."]},{"key":"dc:title","label":"Title","values":["Comparative Analysis of Discrimination and Calibration Accuracy of Discrete Survival, Random Forests, and Neural Networks in Health-Related Survival Prediction Models"]}]}],"canonical_facts":{"dc:contributor.advisor":["Bere, Alphonce","Mulaudzi, Tshilidzi","Motsuku, Lactatia"],"dc:creator":["Ramachela, Audrey Tshepho"],"dc:date":["2025"],"dc:date.accessioned":["2025-11-07T05:34:50Z"],"dc:date.available":["2025-11-07T05:34:50Z"],"dc:date.issued":["2025-09-05"],"dc:description":["MSc in Statistics","Department of Mathematical and Computational Sciences"],"dc:description.abstract":["Prediction models for survival analysis are commonly used in biomedical sciences to understand the onset of certain diseases. Traditional statistical models have been employed for the previous years, however, their limitations and inability to handle big data sets has made a way for the introduction of machine learning methods which gained recognition due to their ability to learn complex algorithms. However, existing literature indicates that the predictive accuracy of machine learning and statistical models for survival analysis varies significantly across different data sets. This variability underscores the need for further research utilizing data sets with diverse characteristics. Such research is essential to develop generalizable insights into the conditions under which each method performs best. In this research project, we compared the predictive performance of traditional statistical method and machine learning algorithms in discrete survival analysis. The machine learning methods include discrete-time survival trees, discrete-time random survival forests, and discrete-time neural networks. The study uses calibration (measured by the prediction error curves) to assess model fit and discrimination (measured by the Concordance index and area under curve) to evaluate predictive accuracy. These methods were applied to data sets: Breast cancer, age at first alcohol intake and CRASH-2. The discrete-time neural network had the best prediction performance as compared to the rest of the models for survival of breast cancer. The discrete-time random forest with hellinger distance had the overall prediction performance on the age at first alcohol intake. The discrete-time survival model outperformed the rest of the models in predicting survival of bleeding trauma patients from the CRASH-2 data ."],"dc:identifier.uri":["https://univendspace.univen.ac.za/handle/11602/3028"],"dc:language.iso":["en"],"dc:relation":["PDF"],"dc:rights":["University of Venda"],"dc:subject":["Discrete-time survival analysis","Statistical methods","Machine learning","Calibration","Discrimination","UCTD"],"dc:title":["Comparative Analysis of Discrimination and Calibration Accuracy of Discrete Survival, Random Forests, and Neural Networks in Health-Related Survival Prediction Models"],"dc:type":["Dissertation"]},"updated_at":"2026-07-27T21:57:28Z"}