Kennesaw State University
Classifying Imbalanced Financial Fraud Data Utilizing Enhanced Random Forest Algorithm
Abstract
dc:description.abstract<p>Imbalanced datasets have been a unique challenge for machine learning, requiring specialized approaches to correctly classify the minority class. Financial fraud detection involves using highly imbalanced datasets with a class imbalance of up to .01% frauds to 99.99% regular transactions. It is essential to identify all frauds in financial fraud detection, even if some classifications' precision is low. I developed a random forest assembly that separates fraudulent transactions into tiers of precision. With this approach, 96% of fraudulent transactions are identified, showing an 8% increase in recall when compared to standard approaches. 59% of fraud classifications' precision increases by 10% up to 98% by optimizing several random forests on different fitness functions. These models are then combined to act as a sieve with increasing tolerance for low precision classifications. The effectiveness of random forest for financial fraud detection is also improved through feature extraction techniques. Random forest is weak at detecting patterns between interdepended features. This problem is address through unsupervised feature extraction. I will demonstrate a new random forest architecture PCA-embedded random forest, which increased random forest performance.</p>
Degree
thesis:*- Name thesis:degree_name
- Master of Science in Computer Science (MSCS)
- Level thesis:degree_level
- Thesis
- Discipline thesis:degree_discipline
- Computer Science
- Year dc:date.available
- 2020
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Gardner, Charles
- Contributors dc:contributor
-
- Dr. Dan lo
- Dr. Yong Shi
Subjects
dc:subject × 6Identifiers
dc:identifier.*- Repository record dc:identifier
- https://digitalcommons.kennesaw.edu/cs_etd/40
- OAI identifier oai:identifier
- oai:digitalcommons.kennesaw.edu:cs_etd-1042