{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/78343"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/78343","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"A theory of (almost) zero resource speech recognition","abstract":"Automatic speech recognition has matured into a commercially successful technology, enabling voice-based interfaces for smartphones, smart TVs, and many other consumer devices. The overwhelming popularity, however, is still limited to languages such as English, Japanese, and German, where vast amounts of labeled training data are available. For most other languages, it is prohibitively expensive to 1) collect and transcribe the speech data required to learn good acoustic models; and 2) acquire adequate text to estimate meaningful language models. A theory of unsupervised and semi-supervised techniques for speech recognition is therefore essential. This thesis focuses on HMM-based sequence clustering and examines acoustic modeling, language modeling, and applications beyond the components of an ASR, such as anomaly detection, from the vantage point of PAC-Bayesian theory. The first part of this thesis extends standard PAC-Bayesian bounds to address the sequential nature of speech and language signals. A novel algorithm, based on sparsifying the cluster assignment probabilities with a Renyi entropy prior, is shown to provably minimize the generalization error of any probabilistic model (e.g. HMMs). The second part examines application-specific loss functions such as cluster purity and perplexity. Empirical results on a variety of tasks -- acoustic event detection, class-based language modeling, and unsupervised sequence anomaly detection -- confirm the practicality of the theory and algorithms developed in this thesis.","abstract_html":"Automatic speech recognition has matured into a commercially successful technology, enabling voice-based interfaces for smartphones, smart TVs, and many other consumer devices. The overwhelming popularity, however, is still limited to languages such as English, Japanese, and German, where vast amounts of labeled training data are available. For most other languages, it is prohibitively expensive to 1) collect and transcribe the speech data required to learn good acoustic models; and 2) acquire adequate text to estimate meaningful language models. A theory of unsupervised and semi-supervised techniques for speech recognition is therefore essential. This thesis focuses on HMM-based sequence clustering and examines acoustic modeling, language modeling, and applications beyond the components of an ASR, such as anomaly detection, from the vantage point of PAC-Bayesian theory. The first part of this thesis extends standard PAC-Bayesian bounds to address the sequential nature of speech and language signals. A novel algorithm, based on sparsifying the cluster assignment probabilities with a Renyi entropy prior, is shown to provably minimize the generalization error of any probabilistic model (e.g. HMMs). The second part examines application-specific loss functions such as cluster purity and perplexity. Empirical results on a variety of tasks -- acoustic event detection, class-based language modeling, and unsupervised sequence anomaly detection -- confirm the practicality of the theory and algorithms developed in this thesis.","abstract_has_math":false,"creators":["Bharadwaj, Sujeeth Subramanya"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Hasegawa-Johnson, Mark A.","Levinson, Stephen E.","Liang, Feng","Smaragdis, Paris"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2015,"date_issued":"2015-07-22T22:16:25Z","date_published":"2015-07-22T22:16:25Z","updated_at":"2026-07-22T22:26:11Z","subjects":["Speech recognition","Unsupervised learning","PAC-Bayesian theory","Language Modeling","Acoustic Event Detection","anomaly detection"],"languages":["en"],"rights":["Copyright 2015 Sujeeth Bharadwaj"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/78343","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Hasegawa-Johnson, Mark A.","Levinson, Stephen E.","Liang, Feng","Smaragdis, Paris"]},{"key":"dc:creator","label":"Author","values":["Bharadwaj, Sujeeth Subramanya"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2015-07-22T22:16:25Z","2015-05","2015-03-31","2015-5"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Speech recognition","Unsupervised learning","PAC-Bayesian theory","Language Modeling","Acoustic Event Detection","anomaly detection"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2015 Sujeeth Bharadwaj"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/78343"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Automatic speech recognition has matured into a commercially successful technology, enabling voice-based interfaces for smartphones, smart TVs, and many other consumer devices. The overwhelming popularity, however, is still limited to languages such as English, Japanese, and German, where vast amounts of labeled training data are available. For most other languages, it is prohibitively expensive to 1) collect and transcribe the speech data required to learn good acoustic models; and 2) acquire adequate text to estimate meaningful language models. A theory of unsupervised and semi-supervised techniques for speech recognition is therefore essential. This thesis focuses on HMM-based sequence clustering and examines acoustic modeling, language modeling, and applications beyond the components of an ASR, such as anomaly detection, from the vantage point of PAC-Bayesian theory. The first part of this thesis extends standard PAC-Bayesian bounds to address the sequential nature of speech and language signals. A novel algorithm, based on sparsifying the cluster assignment probabilities with a Renyi entropy prior, is shown to provably minimize the generalization error of any probabilistic model (e.g. HMMs). The second part examines application-specific loss functions such as cluster purity and perplexity. Empirical results on a variety of tasks -- acoustic event detection, class-based language modeling, and unsupervised sequence anomaly detection -- confirm the practicality of the theory and algorithms developed in this thesis.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2015-07-22 without embargo terms","The student, Sujeeth Bharadwaj, accepted the attached license on 2015-03-27 at 03:53.","The student, Sujeeth Bharadwaj, submitted this Dissertation for approval on 2015-03-27 at 04:03.","This Dissertation was approved for publication on 2015-03-31 at 08:46.","DSpace SAF Submission Ingestion Package generated from Vireo submission #7774 on 2015-07-22 at 10:31:21","Made available in DSpace on 2015-07-22T22:16:25Z (GMT). No. of bitstreams: 2 BHARADWAJ-DISSERTATION-2015.pdf: 1367722 bytes, checksum: fd0ea1c00d30c987d48e9e167d94786e (MD5) LICENSE.txt: 4214 bytes, checksum: 605a4d33bb805b4a61bad54cef5ab77d (MD5) Previous issue date: 2015-03-31"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["A theory of (almost) zero resource speech recognition"]}]}],"canonical_facts":{"dc:contributor":["Hasegawa-Johnson, Mark A.","Levinson, Stephen E.","Liang, Feng","Smaragdis, Paris"],"dc:creator":["Bharadwaj, Sujeeth Subramanya"],"dc:date":["2015-07-22T22:16:25Z","2015-05","2015-03-31","2015-5"],"dc:description":["Automatic speech recognition has matured into a commercially successful technology, enabling voice-based interfaces for smartphones, smart TVs, and many other consumer devices. The overwhelming popularity, however, is still limited to languages such as English, Japanese, and German, where vast amounts of labeled training data are available. For most other languages, it is prohibitively expensive to 1) collect and transcribe the speech data required to learn good acoustic models; and 2) acquire adequate text to estimate meaningful language models. A theory of unsupervised and semi-supervised techniques for speech recognition is therefore essential. This thesis focuses on HMM-based sequence clustering and examines acoustic modeling, language modeling, and applications beyond the components of an ASR, such as anomaly detection, from the vantage point of PAC-Bayesian theory. The first part of this thesis extends standard PAC-Bayesian bounds to address the sequential nature of speech and language signals. A novel algorithm, based on sparsifying the cluster assignment probabilities with a Renyi entropy prior, is shown to provably minimize the generalization error of any probabilistic model (e.g. HMMs). The second part examines application-specific loss functions such as cluster purity and perplexity. Empirical results on a variety of tasks -- acoustic event detection, class-based language modeling, and unsupervised sequence anomaly detection -- confirm the practicality of the theory and algorithms developed in this thesis.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2015-07-22 without embargo terms","The student, Sujeeth Bharadwaj, accepted the attached license on 2015-03-27 at 03:53.","The student, Sujeeth Bharadwaj, submitted this Dissertation for approval on 2015-03-27 at 04:03.","This Dissertation was approved for publication on 2015-03-31 at 08:46.","DSpace SAF Submission Ingestion Package generated from Vireo submission #7774 on 2015-07-22 at 10:31:21","Made available in DSpace on 2015-07-22T22:16:25Z (GMT). No. of bitstreams: 2 BHARADWAJ-DISSERTATION-2015.pdf: 1367722 bytes, checksum: fd0ea1c00d30c987d48e9e167d94786e (MD5) LICENSE.txt: 4214 bytes, checksum: 605a4d33bb805b4a61bad54cef5ab77d (MD5) Previous issue date: 2015-03-31"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/78343"],"dc:language":["en"],"dc:rights":["Copyright 2015 Sujeeth Bharadwaj"],"dc:subject":["Speech recognition","Unsupervised learning","PAC-Bayesian theory","Language Modeling","Acoustic Event Detection","anomaly detection"],"dc:title":["A theory of (almost) zero resource speech recognition"],"dc:type":["text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:26:11Z"}