Stellenbosch : Stellenbosch University
Money Laundering Alert Prioritisation Using Machine Learning and Benford’s Law
Abstract
dc:description.abstractTymeBank’s traditional rule-based system for Anti-Money Laundering generates a high volume of false-positive alerts. From an industrial engineering perspective, this creates a significant process inefficiency, resulting in a costly and unscalable investigative bottleneck that consumes analyst resources and diverts focus from genuine financial crime. This research addresses the problem by developing a machine learning risk-scoring system. The system pursues multiple objectives, aiming to enhance fraud detection by identifying genuine mule accounts while simultaneously enabling risk-based prioritisation to reduce wasted investigative effort. Using TymeBank’s anonymised data, a controlled experiment compared baseline models (Model A), trained on the full dataset, against specialised models (Model B). These specialised configurations were trained only on high-transaction profiles and incorporated additional features derived from Benford’s Law. A range of algorithms, including Logistic Regression, Na¨ıve Bayes, Decision Trees, K-Nearest Neighbours, and Random Forests, was systematically evaluated. To assess models against the multiple objectives, a bespoke evaluation approach with a businessoriented focus was developed. Standard statistical metrics, such as accuracy, are often misleading in the context of imbalanced data, while precision and recall present inherent trade-offs. To address this, a Total Cost function was introduced to determine the balance between effective fraud detection and the high cost of unnecessary investigations, with the latter being directly quantified through a Resource Efficiency metric. The results suggest that a Random Forest, Model A provides the most cost-effective solution, achieving the lowest Total Cost and the highest Resource Efficiency among all tested configurations. In contrast, the specialised Model B exhibited a pronounced trade-off. Although it achieved an increase in the Fraud Capture Rate, this was accompanied by a decrease in Resource Efficiency, leading to a higher overall Total Cost. These results indicate that the predictive value of a larger, general dataset, as utilised by Model A, was more influential on overall performance than the niche statistical features employed by Model B. Finally, SHAP analysis confirmed that the model’s logic is sound, basing its high-risk predictions on intuitive patterns, such as rapid balance depletion. This research contributes an interpretable model and a specialised, cost-centric framework for evaluating AML systems. It provides TymeBank with a validated, data-driven method to enhance operational efficiency, reduce resource waste, and reinforce the fundamental trust that underpins the financial system.
Degree
thesis:*- Grantor dc:publisher
- Stellenbosch : Stellenbosch University
- Year dc:date.issued
- 2026
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Spies, Luka
- Advisor dc:contributor.advisor
-
- Bekker, James
Rights
- Language dc:language.iso
- en
Identifiers
dc:identifier.*- Repository record dc:identifier.uri
- https://scholar.sun.ac.za/handle/10019.1/135818
- OAI identifier oai:identifier
- oai:scholar.sun.ac.za:10019.1/135818