Back to results

NJIT

Decision tree rule-based feature selection for imbalanced data

Abstract

dc:description.abstract

A class imbalance problem appears in many real world applications, e.g., fault diagnosis, text categorization and fraud detection. When dealing with an imbalanced dataset, feature selection becomes an important issue. To address it, this work proposes a feature selection method that is based on a decision tree rule and weighted Gini index. The effectiveness of the proposed methods is verified by classifying a dataset from Santander Bank and two datasets from UCI machine learning repository. The results show that our methods can achieve higher Area Under the Curve (AUC) and F-measure. We also compare them with filter-based feature selection approaches, i.e., Chi-Square and F-statistic. The results show that they outperform them but need slightly more computational efforts.

Degree

thesis:*
Name thesis:degree_name
Master of Science in Computer Engineering - (M.S.)
Discipline thesis:degree_discipline
Electrical and Computer Engineering
Year
2017

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Liu, Haoyue
Contributors dc:contributor
  • MengChu Zhou
  • Osvaldo Simeone
  • Yun Q. Shi

Subjects

dc:subject × 3

Identifiers

dc:identifier.*
Repository record dc:identifier
https://digitalcommons.njit.edu/theses/25
OAI identifier oai:identifier
oai:digitalcommons.njit.edu:theses-1024

Chain of custody

source
Harvested from
NJIT
Base URL
digitalcommons.njit.edu/do/oai/
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Liu, Haoyue. Decision tree rule-based feature selection for imbalanced data. 2017. https://digitalcommons.njit.edu/theses/25