{"id":{"repo_id":"cape-town","oai_identifier":"oai:open.uct.ac.za:11427/41623"},"canonical_url":"https://search.dev.ndltd.org/etd/cape-town/oai:open.uct.ac.za:11427/41623","repository":{"repo_id":"cape-town","name":"University of Cape Town","base_url":"https://open.uct.ac.za/oai/request"},"display":{"title":"Credit scorecards in retail banking: enhancing interpretability through shapley values and evaluating the effectiveness of alternative data for improved accuracy","abstract":"This research addresses the dual challenges of improving credit scorecard accuracy and maintaining interpretability. While machine learning algorithms like random forest and eXtreme gradient boosting outperform traditional logistic regression in accuracy, their complex predictor variable representation hinders interpretability. To reconcile this, the study discretizes numerical variables, applies one-hot encoding, and employs Shapley values to derive interpretable credit scores for random forest, eXtreme gradient boosting, light gradient boosting machine, and categorical boosting models. This approach produces credit scorecards that align with industry standards. Additionally, the investigation into the role of alternative data in credit scoring reveals its impact on model accuracy. By analysing unique predictor variables such as an applicant's social circle default status, regional ratings, and local population size, the significance of alternative data is demonstrated. Leveraging the model-X knockoffs framework for predictor variable selection contributes to superior model performance, achieving the highest area under the curve on the Kaggle home credit data.","abstract_html":"This research addresses the dual challenges of improving credit scorecard accuracy and maintaining interpretability. While machine learning algorithms like random forest and eXtreme gradient boosting outperform traditional logistic regression in accuracy, their complex predictor variable representation hinders interpretability. To reconcile this, the study discretizes numerical variables, applies one-hot encoding, and employs Shapley values to derive interpretable credit scores for random forest, eXtreme gradient boosting, light gradient boosting machine, and categorical boosting models. This approach produces credit scorecards that align with industry standards. Additionally, the investigation into the role of alternative data in credit scoring reveals its impact on model accuracy. By analysing unique predictor variables such as an applicant&#x27;s social circle default status, regional ratings, and local population size, the significance of alternative data is demonstrated. Leveraging the model-X knockoffs framework for predictor variable selection contributes to superior model performance, achieving the highest area under the curve on the Kaggle home credit data.","abstract_has_math":false,"creators":["Hlongwane, Rivalani"],"institution":"Graduate School of Business (GSB)","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Ramaboa, Kutlwano"],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025","date_published":"2025","updated_at":"2026-07-22T22:23:31Z","subjects":["credit scorecard"],"languages":["en"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/11427/41623","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Ramaboa, Kutlwano"]},{"key":"dc:creator","label":"Author","values":["Hlongwane, Rivalani"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2025-08-26T09:00:37Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2025-08-26T09:00:37Z"]},{"key":"dc:date.issued","label":"Date","values":["2025"]},{"key":"dc:publisher.department","label":"Dc Publisher Department","values":["Graduate School of Business (GSB)"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["University of Cape Town"]},{"key":"dc:type","label":"Dc Type","values":["Thesis / Dissertation"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["Doctoral","PhD"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["credit scorecard"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["http://hdl.handle.net/11427/41623"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["This research addresses the dual challenges of improving credit scorecard accuracy and maintaining interpretability. While machine learning algorithms like random forest and eXtreme gradient boosting outperform traditional logistic regression in accuracy, their complex predictor variable representation hinders interpretability. To reconcile this, the study discretizes numerical variables, applies one-hot encoding, and employs Shapley values to derive interpretable credit scores for random forest, eXtreme gradient boosting, light gradient boosting machine, and categorical boosting models. This approach produces credit scorecards that align with industry standards. Additionally, the investigation into the role of alternative data in credit scoring reveals its impact on model accuracy. By analysing unique predictor variables such as an applicant's social circle default status, regional ratings, and local population size, the significance of alternative data is demonstrated. Leveraging the model-X knockoffs framework for predictor variable selection contributes to superior model performance, achieving the highest area under the curve on the Kaggle home credit data."]},{"key":"dc:title","label":"Title","values":["Credit scorecards in retail banking: enhancing interpretability through shapley values and evaluating the effectiveness of alternative data for improved accuracy"]}]}],"canonical_facts":{"dc:contributor.advisor":["Ramaboa, Kutlwano"],"dc:creator":["Hlongwane, Rivalani"],"dc:date.accessioned":["2025-08-26T09:00:37Z"],"dc:date.available":["2025-08-26T09:00:37Z"],"dc:date.issued":["2025"],"dc:description.abstract":["This research addresses the dual challenges of improving credit scorecard accuracy and maintaining interpretability. While machine learning algorithms like random forest and eXtreme gradient boosting outperform traditional logistic regression in accuracy, their complex predictor variable representation hinders interpretability. To reconcile this, the study discretizes numerical variables, applies one-hot encoding, and employs Shapley values to derive interpretable credit scores for random forest, eXtreme gradient boosting, light gradient boosting machine, and categorical boosting models. This approach produces credit scorecards that align with industry standards. Additionally, the investigation into the role of alternative data in credit scoring reveals its impact on model accuracy. By analysing unique predictor variables such as an applicant's social circle default status, regional ratings, and local population size, the significance of alternative data is demonstrated. Leveraging the model-X knockoffs framework for predictor variable selection contributes to superior model performance, achieving the highest area under the curve on the Kaggle home credit data."],"dc:identifier.uri":["http://hdl.handle.net/11427/41623"],"dc:language.iso":["en"],"dc:publisher.department":["Graduate School of Business (GSB)"],"dc:publisher.institution":["University of Cape Town"],"dc:subject":["credit scorecard"],"dc:title":["Credit scorecards in retail banking: enhancing interpretability through shapley values and evaluating the effectiveness of alternative data for improved accuracy"],"dc:type":["Thesis / Dissertation"],"dc:type.qualificationlevel":["Doctoral","PhD"]},"updated_at":"2026-07-22T22:23:31Z"}