University of Ottawa (Canada)
Multiple classifier combination through ensembles and data generation
Abstract
dc:descriptionThis thesis introduces new approaches, namely the DataBoost and DataBoost-IM algorithms, to extend Boosting algorithms' predictive performance. The DataBoost algorithm is designed to assist Boosting algorithms to avoid over-emphasizing hard examples. In the DataBoost algorithm, new synthetic data with bias information towards hard examples are added to the original training set when training the component classifiers. The DataBoost approach was evaluated against ten data sets, using both decision trees and neural networks as base classifiers. The experiments show promising results, in terms of overall accuracy when compared to a standard benchmarking Boosting algorithm. The DataBoost-IM algorithm is developed to learn from two-class imbalanced data sets. In the DataBoost-IM approach, the class frequencies and the total weights against different classes within the ensemble's training set are rebalanced by adding new synthetic data. The DataBoost-IM method was evaluated, in terms of the F-measures, G-mean and overall accuracy, against seventeen highly and moderately imbalanced data sets using decision trees as base classifiers. (Abstract shortened by UMI.)
Degree
thesis:*- Grantor dc:publisher
- University of Ottawa (Canada)
- Year dc:date
- 2013
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Guo, Hong Yu
Subjects
dc:subject × 1Rights
- Language dc:language
- en
Identifiers
dc:identifier.*- Identifier
-
Source: Masters Abstracts International, Volume: 43-06, page: 2406.
http://dx.doi.org/10.20381/ruor-9729 - OAI identifier oai:identifier
- oai:ruor.uottawa.ca:10393/26648