{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/29785"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/29785","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Multiview feature learning for speech recognition","abstract":"In this thesis, we study the problem of learning a linear transformation of acoustic feature vectors for speech recognition, in a framework where apart from the acoustics, additional views are available at training time. We consider a multiview learning approach based on canonical correlation analysis to learn linear transformations of the acoustic features that are maximally correlated with the data. We propose simple approaches for combining information shared across the views with information that is private to the acoustic view. We apply these methods to a specific scenario in which articulatory data is available at training time. Results of phonetic frame classification on data drawn from the University of Wisconsin X-ray Microbeam Database indicate a small but consistent advantage to the multiview approaches that combine shared and private information, compared to the baseline acoustic features or unsupervised dimensionality reduction using principal component analysis. We then discuss limitations of canonical correlation analysis and possible extensions.","abstract_html":"In this thesis, we study the problem of learning a linear transformation of acoustic feature vectors for speech recognition, in a framework where apart from the acoustics, additional views are available at training time. We consider a multiview learning approach based on canonical correlation analysis to learn linear transformations of the acoustic features that are maximally correlated with the data. We propose simple approaches for combining information shared across the views with information that is private to the acoustic view. We apply these methods to a specific scenario in which articulatory data is available at training time. Results of phonetic frame classification on data drawn from the University of Wisconsin X-ray Microbeam Database indicate a small but consistent advantage to the multiview approaches that combine shared and private information, compared to the baseline acoustic features or unsupervised dimensionality reduction using principal component analysis. We then discuss limitations of canonical correlation analysis and possible extensions.","abstract_has_math":false,"creators":["Bharadwaj, Sujeeth"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Hasegawa-Johnson, Mark A."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2012,"date_issued":"2012-02-06T20:16:38Z","date_published":"2012-02-06T20:16:38Z","updated_at":"2026-07-22T22:25:29Z","subjects":["Multiview learning","canonical correlation analysis","articulatory measurements","dimensionality reduction","acoustic features"],"languages":["en"],"rights":["Copyright 2011 Sujeeth Bharadwaj"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/29785","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Hasegawa-Johnson, Mark A."]},{"key":"dc:creator","label":"Author","values":["Bharadwaj, Sujeeth"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2012-02-06T20:16:38Z","2011-12"]},{"key":"dc:type","label":"Dc Type","values":["Dissertation / Thesis","text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Multiview learning","canonical correlation analysis","articulatory measurements","dimensionality reduction","acoustic features"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2011 Sujeeth Bharadwaj"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/29785"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["In this thesis, we study the problem of learning a linear transformation of acoustic feature vectors for speech recognition, in a framework where apart from the acoustics, additional views are available at training time. We consider a multiview learning approach based on canonical correlation analysis to learn linear transformations of the acoustic features that are maximally correlated with the data. We propose simple approaches for combining information shared across the views with information that is private to the acoustic view. We apply these methods to a specific scenario in which articulatory data is available at training time. Results of phonetic frame classification on data drawn from the University of Wisconsin X-ray Microbeam Database indicate a small but consistent advantage to the multiview approaches that combine shared and private information, compared to the baseline acoustic features or unsupervised dimensionality reduction using principal component analysis. We then discuss limitations of canonical correlation analysis and possible extensions.","Item withdrawn by Katherine Eriksen (eriksen3@illinois.edu) on 2011-12-01T22:28:11Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Bharadwaj_Sujeeth.pdf: 601436 bytes, checksum: 120a3522a9206661a0a5deac4d1f4efa (MD5)","Made available in DSpace on 2012-02-06T20:16:38Z (GMT). No. of bitstreams: 2 Bharadwaj_Sujeeth.pdf: 601436 bytes, checksum: 120a3522a9206661a0a5deac4d1f4efa (MD5) license.txt: 4066 bytes, checksum: 6df73f3f7228d0a5647ebdfda4f3d592 (MD5)"]},{"key":"dc:title","label":"Title","values":["Multiview feature learning for speech recognition"]}]}],"canonical_facts":{"dc:contributor":["Hasegawa-Johnson, Mark A."],"dc:creator":["Bharadwaj, Sujeeth"],"dc:date":["2012-02-06T20:16:38Z","2011-12"],"dc:description":["In this thesis, we study the problem of learning a linear transformation of acoustic feature vectors for speech recognition, in a framework where apart from the acoustics, additional views are available at training time. We consider a multiview learning approach based on canonical correlation analysis to learn linear transformations of the acoustic features that are maximally correlated with the data. We propose simple approaches for combining information shared across the views with information that is private to the acoustic view. We apply these methods to a specific scenario in which articulatory data is available at training time. Results of phonetic frame classification on data drawn from the University of Wisconsin X-ray Microbeam Database indicate a small but consistent advantage to the multiview approaches that combine shared and private information, compared to the baseline acoustic features or unsupervised dimensionality reduction using principal component analysis. We then discuss limitations of canonical correlation analysis and possible extensions.","Item withdrawn by Katherine Eriksen (eriksen3@illinois.edu) on 2011-12-01T22:28:11Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Bharadwaj_Sujeeth.pdf: 601436 bytes, checksum: 120a3522a9206661a0a5deac4d1f4efa (MD5)","Made available in DSpace on 2012-02-06T20:16:38Z (GMT). No. of bitstreams: 2 Bharadwaj_Sujeeth.pdf: 601436 bytes, checksum: 120a3522a9206661a0a5deac4d1f4efa (MD5) license.txt: 4066 bytes, checksum: 6df73f3f7228d0a5647ebdfda4f3d592 (MD5)"],"dc:identifier":["http://hdl.handle.net/2142/29785"],"dc:language":["en"],"dc:rights":["Copyright 2011 Sujeeth Bharadwaj"],"dc:subject":["Multiview learning","canonical correlation analysis","articulatory measurements","dimensionality reduction","acoustic features"],"dc:title":["Multiview feature learning for speech recognition"],"dc:type":["Dissertation / Thesis","text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:29Z"}