University of Cincinnati
A Review and Comparative Study on Univariate Feature Selection Techniques
Abstract
dc:description<p>Huge volume of dataset with extremely high dimension are emerging in a notable variety of fields, going from bioinformatics to text mining, which gives rise to a crucial technique, termed as feature selection, as the preprocessing strategy in extracting information and knowledge from dataset. As of 1997, feature selection has been extensively and intensively explored in both areas of statistics and machine learning, and it is still a critical issue in today’s data mining area. There are different basic taxonomies of feature selection wildly accepted by researchers, the two well accepted ones being: (1) filter, wrapper and embedded, and (2) multivariate and univariate. This thesis is devoted to study of univariate feature selection technique, which assesses the discriminative power of each feature individually by pairwise evaluating the dependency between response and each feature based on certain metric.</p><p>Univariate feature selection techniques are primarily employed in tackling the extremely high-dimensional dataset due to its speed and cost-effectiveness. Up to now, quite a few of univariate feature selection techniques have been invented by researchers in machine learning area, or borrowed directly from statistics, and these techniques have been evaluated on some specific datasets and normally proven to be both effective and efficient in terms of model accuracy. In the literature, however, I found generally no justification as to why one technique is suitable for one certain dataset is provided, and thorough comparative studies of univariate feature selection techniques are rarely carried out.</p><p>In this study, I firstly collect the most important existing univariate feature selection techniques then classify them into the following 5 categories: (1) <i>Hypothesis test</i>, (2) <i>Uncertainty reduction</i>, (3) <i>Distance metric</i>, (4) <i>Discriminative power</i> and (5) <i>Correlation coefficient</i>. A part of typical techniques are further applied and assessed on the artificial datasets then a comparative study is carried out. The artificial datasets are virtually a pair of input (feature) and response characterized by data type of input and response, degree of dependency, form of dependency and statistical distribution of input value. Hopefully the artificial datasets with various combinations of above-mentioned factors can represent the real dataset as far as possible, so that the advantages and shortcomings of these techniques can be highlighted. Finally, a guideline is established to provide researchers the advices of choosing most appropriate univariate feature selection techniques to one specific dataset based on the given prior information, and further evaluated by a case study.</p>
Degree
thesis:*- Name thesis:degree_name
- MS
- Level thesis:degree_level
- masters
- Discipline thesis:degree_discipline
- Engineering and Applied Science: Mechanical Engineering
- Grantor dc:publisher
- University of Cincinnati
- Year dc:date
- 2012
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Ni, Weizeng
- Contributors dc:contributor
-
- Huang, Hongdao
Subjects
dc:subject × 4Rights
dc:rights- Statement dc:rights
-
- unrestricted
- This thesis or dissertation is protected by copyright: all rights reserved. It may not be copied or redistributed beyond the terms of applicable copyright laws.
- Language dc:language
- English
Identifiers
dc:identifier.*- Repository record dc:identifier
- http://rave.ohiolink.edu/etdc/view?acc_num=ucin1353156184
- OAI identifier oai:identifier
- oai:etd.ohiolink.edu:ucin1353156184