{"id":{"repo_id":"colostate","oai_identifier":"oai:mountainscholar.org:10217/244101"},"canonical_url":"https://search.dev.ndltd.org/etd/colostate/oai:mountainscholar.org:10217/244101","repository":{"repo_id":"colostate","name":"Colorado State University","base_url":"https://api.mountainscholar.org/server/oai/request"},"display":{"title":"Feature selection from huge feature sets in the context of computer vision","abstract":"Most learning systems use hand-picked sets of features as input data for their learning algorithms. This is particularly true of computer vision systems, where the number of features that can be computed over an image is, for practical purposes, limitless. Unfortunately, most of these features are irrelevant or redundant to a given task, and no feature selection algorithm to date can handle such large feature sets. Moreover, many standard feature selection algorithms perform poorly when faced with many irrelevant and redundant features. This work addresses the feature selection problem by proposing a three-step algorithm. The first step uses an algorithm based on the well known algorithm called Relief [54] to remove irrelevance: the second step clusters features using K-means to remove redundancy: and the third step is a standard feature selection algorithm. This three-step algorithm is shown to be more effective than standard feature selection algorithms for data with lots of irrelevance and redundancy. In other experiment a data set with 4096 features was reduced to 5% of its original size with very little information loss. In addition, we modify Relief to remove its bias against non-monotonic features and use correlation as the distance measure for K-means.","abstract_html":"Most learning systems use hand-picked sets of features as input data for their learning algorithms. This is particularly true of computer vision systems, where the number of features that can be computed over an image is, for practical purposes, limitless. Unfortunately, most of these features are irrelevant or redundant to a given task, and no feature selection algorithm to date can handle such large feature sets. Moreover, many standard feature selection algorithms perform poorly when faced with many irrelevant and redundant features. This work addresses the feature selection problem by proposing a three-step algorithm. The first step uses an algorithm based on the well known algorithm called Relief [54] to remove irrelevance: the second step clusters features using K-means to remove redundancy: and the third step is a standard feature selection algorithm. This three-step algorithm is shown to be more effective than standard feature selection algorithms for data with lots of irrelevance and redundancy. In other experiment a data set with 4096 features was reduced to 5% of its original size with very little information loss. In addition, we modify Relief to remove its bias against non-monotonic features and use correlation as the distance measure for K-means.","abstract_has_math":false,"creators":["Bins Filho, José Carlos, author","Draper, Bruce A., advisor","Kirby, Michael, committee member","Beveridge, J. Ross, committee member","Anderson, Charles W., committee member"],"institution":"Colorado State University. Libraries","degree_name":"Doctor of Philosophy (Ph.D.)","degree_level":"Doctoral","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2000,"date_issued":"2000","date_published":"2000","updated_at":"2026-07-27T19:13:11Z","subjects":["computer science"],"languages":["eng","English"],"rights":["Copyright and other restrictions may apply. User is responsible for compliance with all applicable laws. For information about copyright law, please see https://libguides.colostate.edu/copyright."],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://doi.org/10.25675/3.026725"],"render_values":[{"text":"https://doi.org/10.25675/3.026725","href":"https://doi.org/10.25675/3.026725","code":true}]}]},"links":{"outbound_url":"https://hdl.handle.net/10217/244101","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["Bins Filho, José Carlos, author","Draper, Bruce A., advisor","Kirby, Michael, committee member","Beveridge, J. Ross, committee member","Anderson, Charles W., committee member"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2026-04-22T18:19:07Z"]},{"key":"dc:date.issued","label":"Date","values":["2000"]},{"key":"dc:publisher","label":"Institution","values":["Colorado State University. Libraries"]},{"key":"dc:type","label":"Dc Type","values":["Text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Doctoral"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctor of Philosophy (Ph.D.)"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Colorado State University"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["computer science"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["English"]},{"key":"dc:language.iso","label":"Language (ISO)","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright and other restrictions may apply. User is responsible for compliance with all applicable laws. For information about copyright law, please see https://libguides.colostate.edu/copyright."]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10217/244101","https://doi.org/10.25675/3.026725"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Most learning systems use hand-picked sets of features as input data for their learning algorithms. This is particularly true of computer vision systems, where the number of features that can be computed over an image is, for practical purposes, limitless. Unfortunately, most of these features are irrelevant or redundant to a given task, and no feature selection algorithm to date can handle such large feature sets. Moreover, many standard feature selection algorithms perform poorly when faced with many irrelevant and redundant features. This work addresses the feature selection problem by proposing a three-step algorithm. The first step uses an algorithm based on the well known algorithm called Relief [54] to remove irrelevance: the second step clusters features using K-means to remove redundancy: and the third step is a standard feature selection algorithm. This three-step algorithm is shown to be more effective than standard feature selection algorithms for data with lots of irrelevance and redundancy. In other experiment a data set with 4096 features was reduced to 5% of its original size with very little information loss. In addition, we modify Relief to remove its bias against non-monotonic features and use correlation as the distance measure for K-means."]},{"key":"dc:format.medium","label":"Dc Format Medium","values":["doctoral dissertations"]},{"key":"dc:title","label":"Title","values":["Feature selection from huge feature sets in the context of computer vision"]}]}],"canonical_facts":{"dc:creator":["Bins Filho, José Carlos, author","Draper, Bruce A., advisor","Kirby, Michael, committee member","Beveridge, J. Ross, committee member","Anderson, Charles W., committee member"],"dc:date.accessioned":["2026-04-22T18:19:07Z"],"dc:date.issued":["2000"],"dc:description.abstract":["Most learning systems use hand-picked sets of features as input data for their learning algorithms. This is particularly true of computer vision systems, where the number of features that can be computed over an image is, for practical purposes, limitless. Unfortunately, most of these features are irrelevant or redundant to a given task, and no feature selection algorithm to date can handle such large feature sets. Moreover, many standard feature selection algorithms perform poorly when faced with many irrelevant and redundant features. This work addresses the feature selection problem by proposing a three-step algorithm. The first step uses an algorithm based on the well known algorithm called Relief [54] to remove irrelevance: the second step clusters features using K-means to remove redundancy: and the third step is a standard feature selection algorithm. This three-step algorithm is shown to be more effective than standard feature selection algorithms for data with lots of irrelevance and redundancy. In other experiment a data set with 4096 features was reduced to 5% of its original size with very little information loss. In addition, we modify Relief to remove its bias against non-monotonic features and use correlation as the distance measure for K-means."],"dc:format.medium":["doctoral dissertations"],"dc:identifier.uri":["https://hdl.handle.net/10217/244101","https://doi.org/10.25675/3.026725"],"dc:language":["English"],"dc:language.iso":["eng"],"dc:publisher":["Colorado State University. Libraries"],"dc:rights":["Copyright and other restrictions may apply. User is responsible for compliance with all applicable laws. For information about copyright law, please see https://libguides.colostate.edu/copyright."],"dc:subject":["computer science"],"dc:title":["Feature selection from huge feature sets in the context of computer vision"],"dc:type":["Text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Doctoral"],"thesis:degree_name":["Doctor of Philosophy (Ph.D.)"],"thesis:institution_name":["Colorado State University"]},"updated_at":"2026-07-27T19:13:11Z"}