{"id":{"repo_id":"purdue-thes","oai_identifier":"oai:docs.lib.purdue.edu:open_access_dissertations-1159"},"canonical_url":"https://search.dev.ndltd.org/etd/purdue-thes/oai:docs.lib.purdue.edu:open_access_dissertations-1159","repository":{"repo_id":"purdue-thes","name":"Purdue University","base_url":"https://docs.lib.purdue.edu/do/oai/"},"display":{"title":"New Covariance-Based Feature Extraction Methods for Classification and Prediction of High-Dimensional Data","abstract":"<p>When analyzing high dimensional data sets, it is often necessary to implement feature extraction methods in order to capture relevant discriminating information useful for the purposes of classification and prediction. The relevant information can typically be represented in lower-dimensional feature spaces, and a widely used approach for this is the principal component analysis (PCA) method. PCA efficiently compresses information into lower dimensions; however, studies indicate that it is not optimal for feature extraction especially when dealing with classification problems. Furthermore, for high-dimensional data having limited observations, as is typically the case with remote sensing data and nonstationary data such as financial data, covariance matrix estimation becomes unreliable, and this adversely affects the representation of data in the PCA domain. In this thesis, we first introduce a new feature extraction method called summed component analysis (SCA), which makes use of the structure of eigenvectors of the common covariance matrix to generate new features as sums of certain original features. Secondly, we present a variation of SCA, known as class summed component analysis (CSCA). CSCA takes advantage of the relative ease of computing the class covariance matrices and uses them to determine data transformations. Since the new features consist of simple sums of the original features, we are able to gain a conceptual meaning of the new representation of the data which is appealing for man-machine interface. We evaluate these methods on data sets with varying sample sizes and on financial time series, and are able to show improved classification and prediction accuracies.</p>","abstract_html":"&lt;p&gt;When analyzing high dimensional data sets, it is often necessary to implement feature extraction methods in order to capture relevant discriminating information useful for the purposes of classification and prediction. The relevant information can typically be represented in lower-dimensional feature spaces, and a widely used approach for this is the principal component analysis (PCA) method. PCA efficiently compresses information into lower dimensions; however, studies indicate that it is not optimal for feature extraction especially when dealing with classification problems. Furthermore, for high-dimensional data having limited observations, as is typically the case with remote sensing data and nonstationary data such as financial data, covariance matrix estimation becomes unreliable, and this adversely affects the representation of data in the PCA domain. In this thesis, we first introduce a new feature extraction method called summed component analysis (SCA), which makes use of the structure of eigenvectors of the common covariance matrix to generate new features as sums of certain original features. Secondly, we present a variation of SCA, known as class summed component analysis (CSCA). CSCA takes advantage of the relative ease of computing the class covariance matrices and uses them to determine data transformations. Since the new features consist of simple sums of the original features, we are able to gain a conceptual meaning of the new representation of the data which is appealing for man-machine interface. We evaluate these methods on data sets with varying sample sizes and on financial time series, and are able to show improved classification and prediction accuracies.&lt;/p&gt;","abstract_has_math":false,"creators":["Sofolahan, Mopelola Adediwura"],"institution":null,"degree_name":"Doctor of Philosophy (PhD)","degree_level":"Dissertation","degree_discipline":"Electrical and Computer Engineering","degree_department":null,"school":null,"contributors":["Okan K. Ersoy","Arif Ghafoor","Cordelia M. Brown","Michael D. Zoltowski"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2013,"date_issued":"2013-10-01T07:00:00Z","date_published":"2013-10-01T07:00:00Z","updated_at":"2026-07-24T03:53:11Z","subjects":["classification","hig-dimensional data","machine learning","neural networks","pattern recognition","Electrical and Electronics","Finance and Financial Management"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://docs.lib.purdue.edu/open_access_dissertations/57","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Okan K. Ersoy","Arif Ghafoor","Cordelia M. Brown","Michael D. Zoltowski"]},{"key":"dc:creator","label":"Author","values":["Sofolahan, Mopelola Adediwura"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical and Computer Engineering"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctor of Philosophy (PhD)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["classification","hig-dimensional data","machine learning","neural networks","pattern recognition","Electrical and Electronics","Finance and Financial Management"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://docs.lib.purdue.edu/open_access_dissertations/57"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>When analyzing high dimensional data sets, it is often necessary to implement feature extraction methods in order to capture relevant discriminating information useful for the purposes of classification and prediction. The relevant information can typically be represented in lower-dimensional feature spaces, and a widely used approach for this is the principal component analysis (PCA) method. PCA efficiently compresses information into lower dimensions; however, studies indicate that it is not optimal for feature extraction especially when dealing with classification problems. Furthermore, for high-dimensional data having limited observations, as is typically the case with remote sensing data and nonstationary data such as financial data, covariance matrix estimation becomes unreliable, and this adversely affects the representation of data in the PCA domain. In this thesis, we first introduce a new feature extraction method called summed component analysis (SCA), which makes use of the structure of eigenvectors of the common covariance matrix to generate new features as sums of certain original features. Secondly, we present a variation of SCA, known as class summed component analysis (CSCA). CSCA takes advantage of the relative ease of computing the class covariance matrices and uses them to determine data transformations. Since the new features consist of simple sums of the original features, we are able to gain a conceptual meaning of the new representation of the data which is appealing for man-machine interface. We evaluate these methods on data sets with varying sample sizes and on financial time series, and are able to show improved classification and prediction accuracies.</p>"]},{"key":"dc:title","label":"Title","values":["New Covariance-Based Feature Extraction Methods for Classification and Prediction of High-Dimensional Data"]}]}],"canonical_facts":{"dc:contributor":["Okan K. Ersoy","Arif Ghafoor","Cordelia M. Brown","Michael D. Zoltowski"],"dc:creator":["Sofolahan, Mopelola Adediwura"],"dc:description.abstract":["<p>When analyzing high dimensional data sets, it is often necessary to implement feature extraction methods in order to capture relevant discriminating information useful for the purposes of classification and prediction. The relevant information can typically be represented in lower-dimensional feature spaces, and a widely used approach for this is the principal component analysis (PCA) method. PCA efficiently compresses information into lower dimensions; however, studies indicate that it is not optimal for feature extraction especially when dealing with classification problems. Furthermore, for high-dimensional data having limited observations, as is typically the case with remote sensing data and nonstationary data such as financial data, covariance matrix estimation becomes unreliable, and this adversely affects the representation of data in the PCA domain. In this thesis, we first introduce a new feature extraction method called summed component analysis (SCA), which makes use of the structure of eigenvectors of the common covariance matrix to generate new features as sums of certain original features. Secondly, we present a variation of SCA, known as class summed component analysis (CSCA). CSCA takes advantage of the relative ease of computing the class covariance matrices and uses them to determine data transformations. Since the new features consist of simple sums of the original features, we are able to gain a conceptual meaning of the new representation of the data which is appealing for man-machine interface. We evaluate these methods on data sets with varying sample sizes and on financial time series, and are able to show improved classification and prediction accuracies.</p>"],"dc:identifier":["https://docs.lib.purdue.edu/open_access_dissertations/57"],"dc:subject":["classification","hig-dimensional data","machine learning","neural networks","pattern recognition","Electrical and Electronics","Finance and Financial Management"],"dc:title":["New Covariance-Based Feature Extraction Methods for Classification and Prediction of High-Dimensional Data"],"thesis:degree_discipline":["Electrical and Computer Engineering"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Doctor of Philosophy (PhD)"]},"updated_at":"2026-07-24T03:53:11Z"}