{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/29816"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/29816","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Sparse modeling of high-dimensional data for learning and vision","abstract":"Sparse representations account for most or all of the information of a signal by a linear combination of a few elementary signals called atoms, and have increasingly become recognized as providing high performance for applications as diverse as noise reduction, compression, inpainting, compressive sensing, pattern classification, and blind source separation. In this dissertation, we learn the sparse representations of high-dimensional signals for various learning and vision tasks, including image classification, single image super-resolution, compressive sensing, and graph learning. Based on the bag-of-features (BoF) image representation in a spatial pyramid, we first transform each local image descriptor into a sparse representation, and then these sparse representations are summarized into a fixed-length feature vector over different spatial locations across different spatial scales by max pooling. The proposed generic image feature representation properly handles the large in-class variance problem in image classification, and experiments on object recognition, scene classification, face recognition, gender recognition, and handwritten digit recognition all lead to state-of-the-art performances on the benchmark datasets. We cast the image super-resolution problem as one of recovering a high-resolution image patch for each low-resolution image patch based on recent sparse signal recovery theories, which state that, under mild conditions, a high-resolution signal can be recovered from its low-resolution version if the signal has a sparse representation in terms of some dictionary. We jointly learn the dictionaries for high- and low-resolution image patches and enforce them to have common sparse representations for better recovery. Furthermore, we employ image features and enforce patch overlapping constraints to improve prediction accuracy. Experiments show that the algorithm leads to surprisingly good results. Graph construction is critical for those graph-orientated algorithms designed for the purposes of data clustering, subspace learning, and semi-supervised learning. We model the graph construction problem, including neighbor selection and weight assignment, by finding the sparse representation of a data sample with respect to all other data samples. Since natural signals are high-dimensional signals of a low intrinsic dimension, projecting a signal onto the nearest and lowest dimensional linear subspace is more likely to find its kindred neighbors, and therefore improves the graph quality by avoiding many spurious connections. The proposed graph is informative, sparse, robust to noise, and adaptive to the neighborhood selection; it exhibits exceptionally high performance in various graph-based applications. To this end, we propose a generic dictionary training algorithm that learns more meaningful sparse representations for the above tasks. The dictionary learning algorithm is formulated as a bilevel optimization problem, which we prove can be solved using stochastic gradient descent. Applications of the generic dictionary training algorithm in supervised dictionary training for image classification, super-resolution, and compressive sensing demonstrate its effectiveness in sparse modeling of natural signals.","abstract_html":"Sparse representations account for most or all of the information of a signal by a linear combination of a few elementary signals called atoms, and have increasingly become recognized as providing high performance for applications as diverse as noise reduction, compression, inpainting, compressive sensing, pattern classification, and blind source separation. In this dissertation, we learn the sparse representations of high-dimensional signals for various learning and vision tasks, including image classification, single image super-resolution, compressive sensing, and graph learning. Based on the bag-of-features (BoF) image representation in a spatial pyramid, we first transform each local image descriptor into a sparse representation, and then these sparse representations are summarized into a fixed-length feature vector over different spatial locations across different spatial scales by max pooling. The proposed generic image feature representation properly handles the large in-class variance problem in image classification, and experiments on object recognition, scene classification, face recognition, gender recognition, and handwritten digit recognition all lead to state-of-the-art performances on the benchmark datasets. We cast the image super-resolution problem as one of recovering a high-resolution image patch for each low-resolution image patch based on recent sparse signal recovery theories, which state that, under mild conditions, a high-resolution signal can be recovered from its low-resolution version if the signal has a sparse representation in terms of some dictionary. We jointly learn the dictionaries for high- and low-resolution image patches and enforce them to have common sparse representations for better recovery. Furthermore, we employ image features and enforce patch overlapping constraints to improve prediction accuracy. Experiments show that the algorithm leads to surprisingly good results. Graph construction is critical for those graph-orientated algorithms designed for the purposes of data clustering, subspace learning, and semi-supervised learning. We model the graph construction problem, including neighbor selection and weight assignment, by finding the sparse representation of a data sample with respect to all other data samples. Since natural signals are high-dimensional signals of a low intrinsic dimension, projecting a signal onto the nearest and lowest dimensional linear subspace is more likely to find its kindred neighbors, and therefore improves the graph quality by avoiding many spurious connections. The proposed graph is informative, sparse, robust to noise, and adaptive to the neighborhood selection; it exhibits exceptionally high performance in various graph-based applications. To this end, we propose a generic dictionary training algorithm that learns more meaningful sparse representations for the above tasks. The dictionary learning algorithm is formulated as a bilevel optimization problem, which we prove can be solved using stochastic gradient descent. Applications of the generic dictionary training algorithm in supervised dictionary training for image classification, super-resolution, and compressive sensing demonstrate its effectiveness in sparse modeling of natural signals.","abstract_has_math":false,"creators":["Yang, Jianchao"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Huang, Thomas S.","Ahuja, Narendra","Liang, Zhi-Pei","Ma, Yi","Yu, Kai"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2012,"date_issued":"2012-02-06T20:19:10Z","date_published":"2012-02-06T20:19:10Z","updated_at":"2026-07-22T22:25:29Z","subjects":["sparse coding","sparse representation","sparse modeling","image classification","super-resolution","graph learning","bilevel optimization"],"languages":["en"],"rights":["Copyright 2011 Jianchao Yang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/29816","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Huang, Thomas S.","Ahuja, Narendra","Liang, Zhi-Pei","Ma, Yi","Yu, Kai"]},{"key":"dc:creator","label":"Author","values":["Yang, Jianchao"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2012-02-06T20:19:10Z","2011-12"]},{"key":"dc:type","label":"Dc Type","values":["Dissertation / Thesis","text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["sparse coding","sparse representation","sparse modeling","image classification","super-resolution","graph learning","bilevel optimization"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2011 Jianchao Yang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/29816"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Sparse representations account for most or all of the information of a signal by a linear combination of a few elementary signals called atoms, and have increasingly become recognized as providing high performance for applications as diverse as noise reduction, compression, inpainting, compressive sensing, pattern classification, and blind source separation. In this dissertation, we learn the sparse representations of high-dimensional signals for various learning and vision tasks, including image classification, single image super-resolution, compressive sensing, and graph learning. Based on the bag-of-features (BoF) image representation in a spatial pyramid, we first transform each local image descriptor into a sparse representation, and then these sparse representations are summarized into a fixed-length feature vector over different spatial locations across different spatial scales by max pooling. The proposed generic image feature representation properly handles the large in-class variance problem in image classification, and experiments on object recognition, scene classification, face recognition, gender recognition, and handwritten digit recognition all lead to state-of-the-art performances on the benchmark datasets. We cast the image super-resolution problem as one of recovering a high-resolution image patch for each low-resolution image patch based on recent sparse signal recovery theories, which state that, under mild conditions, a high-resolution signal can be recovered from its low-resolution version if the signal has a sparse representation in terms of some dictionary. We jointly learn the dictionaries for high- and low-resolution image patches and enforce them to have common sparse representations for better recovery. Furthermore, we employ image features and enforce patch overlapping constraints to improve prediction accuracy. Experiments show that the algorithm leads to surprisingly good results. Graph construction is critical for those graph-orientated algorithms designed for the purposes of data clustering, subspace learning, and semi-supervised learning. We model the graph construction problem, including neighbor selection and weight assignment, by finding the sparse representation of a data sample with respect to all other data samples. Since natural signals are high-dimensional signals of a low intrinsic dimension, projecting a signal onto the nearest and lowest dimensional linear subspace is more likely to find its kindred neighbors, and therefore improves the graph quality by avoiding many spurious connections. The proposed graph is informative, sparse, robust to noise, and adaptive to the neighborhood selection; it exhibits exceptionally high performance in various graph-based applications. To this end, we propose a generic dictionary training algorithm that learns more meaningful sparse representations for the above tasks. The dictionary learning algorithm is formulated as a bilevel optimization problem, which we prove can be solved using stochastic gradient descent. Applications of the generic dictionary training algorithm in supervised dictionary training for image classification, super-resolution, and compressive sensing demonstrate its effectiveness in sparse modeling of natural signals.","Item withdrawn by Alexis Thompson (athmpsn1@illinois.edu) on 2011-12-01T22:09:50Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Yang_Jianchao.pdf: 3348207 bytes, checksum: 56b267f303e21a6cc85a5490207c9949 (MD5)","Made available in DSpace on 2012-02-06T20:19:10Z (GMT). No. of bitstreams: 2 Yang_Jianchao.pdf: 3348207 bytes, checksum: 56b267f303e21a6cc85a5490207c9949 (MD5) license.txt: 4062 bytes, checksum: 1a8962f5fd6a9df3263df2411909f799 (MD5)"]},{"key":"dc:title","label":"Title","values":["Sparse modeling of high-dimensional data for learning and vision"]}]}],"canonical_facts":{"dc:contributor":["Huang, Thomas S.","Ahuja, Narendra","Liang, Zhi-Pei","Ma, Yi","Yu, Kai"],"dc:creator":["Yang, Jianchao"],"dc:date":["2012-02-06T20:19:10Z","2011-12"],"dc:description":["Sparse representations account for most or all of the information of a signal by a linear combination of a few elementary signals called atoms, and have increasingly become recognized as providing high performance for applications as diverse as noise reduction, compression, inpainting, compressive sensing, pattern classification, and blind source separation. In this dissertation, we learn the sparse representations of high-dimensional signals for various learning and vision tasks, including image classification, single image super-resolution, compressive sensing, and graph learning. Based on the bag-of-features (BoF) image representation in a spatial pyramid, we first transform each local image descriptor into a sparse representation, and then these sparse representations are summarized into a fixed-length feature vector over different spatial locations across different spatial scales by max pooling. The proposed generic image feature representation properly handles the large in-class variance problem in image classification, and experiments on object recognition, scene classification, face recognition, gender recognition, and handwritten digit recognition all lead to state-of-the-art performances on the benchmark datasets. We cast the image super-resolution problem as one of recovering a high-resolution image patch for each low-resolution image patch based on recent sparse signal recovery theories, which state that, under mild conditions, a high-resolution signal can be recovered from its low-resolution version if the signal has a sparse representation in terms of some dictionary. We jointly learn the dictionaries for high- and low-resolution image patches and enforce them to have common sparse representations for better recovery. Furthermore, we employ image features and enforce patch overlapping constraints to improve prediction accuracy. Experiments show that the algorithm leads to surprisingly good results. Graph construction is critical for those graph-orientated algorithms designed for the purposes of data clustering, subspace learning, and semi-supervised learning. We model the graph construction problem, including neighbor selection and weight assignment, by finding the sparse representation of a data sample with respect to all other data samples. Since natural signals are high-dimensional signals of a low intrinsic dimension, projecting a signal onto the nearest and lowest dimensional linear subspace is more likely to find its kindred neighbors, and therefore improves the graph quality by avoiding many spurious connections. The proposed graph is informative, sparse, robust to noise, and adaptive to the neighborhood selection; it exhibits exceptionally high performance in various graph-based applications. To this end, we propose a generic dictionary training algorithm that learns more meaningful sparse representations for the above tasks. The dictionary learning algorithm is formulated as a bilevel optimization problem, which we prove can be solved using stochastic gradient descent. Applications of the generic dictionary training algorithm in supervised dictionary training for image classification, super-resolution, and compressive sensing demonstrate its effectiveness in sparse modeling of natural signals.","Item withdrawn by Alexis Thompson (athmpsn1@illinois.edu) on 2011-12-01T22:09:50Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Yang_Jianchao.pdf: 3348207 bytes, checksum: 56b267f303e21a6cc85a5490207c9949 (MD5)","Made available in DSpace on 2012-02-06T20:19:10Z (GMT). No. of bitstreams: 2 Yang_Jianchao.pdf: 3348207 bytes, checksum: 56b267f303e21a6cc85a5490207c9949 (MD5) license.txt: 4062 bytes, checksum: 1a8962f5fd6a9df3263df2411909f799 (MD5)"],"dc:identifier":["http://hdl.handle.net/2142/29816"],"dc:language":["en"],"dc:rights":["Copyright 2011 Jianchao Yang"],"dc:subject":["sparse coding","sparse representation","sparse modeling","image classification","super-resolution","graph learning","bilevel optimization"],"dc:title":["Sparse modeling of high-dimensional data for learning and vision"],"dc:type":["Dissertation / Thesis","text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:29Z"}