{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/32006"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/32006","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Semi-supervised learning for acoustic and prosodic modeling in speech applications","abstract":"Enormous amounts of audio recordings of human speech are essential ingredients for building reliable statistical models for many speech applications, such as automatic speech recognition and automatic prosody detection. However, most of these speech data are not being utilized because they lack transcriptions. The goal of this thesis is to use untranscribed (unlabeled) data to improve the performance of models trained using only transcribed (labeled) data. We propose a unified semi-supervised learning framework for the problem of phone classification, phone recognition and prosody detection. The proposed approach will be particularly useful in the case where recognition performance is limited by the amount of transcribed data. In the first part of the thesis, we investigate semi-supervised training of Gaussian Mixtures Models (GMMs) and Hidden Markov Models (HMMs) which are the common probabilistic models of acoustic features in a state-of-the-art continuous density HMM based speech recognition system. Specifically, a family of semi-supervised training criteria that reflects reasonable assumptions about labeled and unlabeled data is proposed. Both generative and discriminative kinds of training criteria are explored, and one important proposal of this thesis is to keep the power of discriminative training criteria by using some measures on unlabeled data as regularization to the supervised training objective. Methods are described for the optimization of these criteria, and phone classification experiments show that these criteria reliably give improvements over their supervised versions that use only labeled data. We then extend the proposed semi-supervised training criteria to the phone recognition problem. This problem is novel in the area of semi-supervised learning because there is little research on the use of unlabeled data in the sequence labeling problems. We develop lattice-based approaches for the model optimization that involves both transcribed and untranscribed speech utterances. Experiments for phone recognition show that a maximum mutual information criterion regularized by negative conditional entropy measured using unlabeled data reliably gives better results than other semi-supervised training methods. In the second part of the thesis, we propose to exploit unlabeled data for the task of automatic prosodic event detection. Prosody annotation is even harder to obtain than orthographic text transcription; it usually requires the expert knowledge of phonetics and linguistics. Therefore, we aim at reducing the annotation efforts for building an automatic prosodic event detector. We show that the mixture model has the ability of class discovery when labeled data are available from only one of the two classes and develop the learning algorithm for unsupervised prosodic boundary detection.","abstract_html":"Enormous amounts of audio recordings of human speech are essential ingredients for building reliable statistical models for many speech applications, such as automatic speech recognition and automatic prosody detection. However, most of these speech data are not being utilized because they lack transcriptions. The goal of this thesis is to use untranscribed (unlabeled) data to improve the performance of models trained using only transcribed (labeled) data. We propose a unified semi-supervised learning framework for the problem of phone classification, phone recognition and prosody detection. The proposed approach will be particularly useful in the case where recognition performance is limited by the amount of transcribed data. In the first part of the thesis, we investigate semi-supervised training of Gaussian Mixtures Models (GMMs) and Hidden Markov Models (HMMs) which are the common probabilistic models of acoustic features in a state-of-the-art continuous density HMM based speech recognition system. Specifically, a family of semi-supervised training criteria that reflects reasonable assumptions about labeled and unlabeled data is proposed. Both generative and discriminative kinds of training criteria are explored, and one important proposal of this thesis is to keep the power of discriminative training criteria by using some measures on unlabeled data as regularization to the supervised training objective. Methods are described for the optimization of these criteria, and phone classification experiments show that these criteria reliably give improvements over their supervised versions that use only labeled data. We then extend the proposed semi-supervised training criteria to the phone recognition problem. This problem is novel in the area of semi-supervised learning because there is little research on the use of unlabeled data in the sequence labeling problems. We develop lattice-based approaches for the model optimization that involves both transcribed and untranscribed speech utterances. Experiments for phone recognition show that a maximum mutual information criterion regularized by negative conditional entropy measured using unlabeled data reliably gives better results than other semi-supervised training methods. In the second part of the thesis, we propose to exploit unlabeled data for the task of automatic prosodic event detection. Prosody annotation is even harder to obtain than orthographic text transcription; it usually requires the expert knowledge of phonetics and linguistics. Therefore, we aim at reducing the annotation efforts for building an automatic prosodic event detector. We show that the mixture model has the ability of class discovery when labeled data are available from only one of the two classes and develop the learning algorithm for unsupervised prosodic boundary detection.","abstract_has_math":false,"creators":["Huang, Jui Ting"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Hasegawa-Johnson, Mark A.","Cole, Jennifer S.","Huang, Thomas S.","Levinson, Stephen E."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2012,"date_issued":"2012-06-27T21:24:08Z","date_published":"2012-06-27T21:24:08Z","updated_at":"2026-07-22T22:25:30Z","subjects":["Semi-Supervised Learning","Speech Recognition","Acoustic Modeling","Prosodic Modeling"],"languages":["en"],"rights":["Copyright 2012 Jui Ting Huang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/32006","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Hasegawa-Johnson, Mark A.","Cole, Jennifer S.","Huang, Thomas S.","Levinson, Stephen E."]},{"key":"dc:creator","label":"Author","values":["Huang, Jui Ting"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2012-06-27T21:24:08Z","2014-06-28T10:00:25Z","2012-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Semi-Supervised Learning","Speech Recognition","Acoustic Modeling","Prosodic Modeling"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2012 Jui Ting Huang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/32006"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Enormous amounts of audio recordings of human speech are essential ingredients for building reliable statistical models for many speech applications, such as automatic speech recognition and automatic prosody detection. However, most of these speech data are not being utilized because they lack transcriptions. The goal of this thesis is to use untranscribed (unlabeled) data to improve the performance of models trained using only transcribed (labeled) data. We propose a unified semi-supervised learning framework for the problem of phone classification, phone recognition and prosody detection. The proposed approach will be particularly useful in the case where recognition performance is limited by the amount of transcribed data. In the first part of the thesis, we investigate semi-supervised training of Gaussian Mixtures Models (GMMs) and Hidden Markov Models (HMMs) which are the common probabilistic models of acoustic features in a state-of-the-art continuous density HMM based speech recognition system. Specifically, a family of semi-supervised training criteria that reflects reasonable assumptions about labeled and unlabeled data is proposed. Both generative and discriminative kinds of training criteria are explored, and one important proposal of this thesis is to keep the power of discriminative training criteria by using some measures on unlabeled data as regularization to the supervised training objective. Methods are described for the optimization of these criteria, and phone classification experiments show that these criteria reliably give improvements over their supervised versions that use only labeled data. We then extend the proposed semi-supervised training criteria to the phone recognition problem. This problem is novel in the area of semi-supervised learning because there is little research on the use of unlabeled data in the sequence labeling problems. We develop lattice-based approaches for the model optimization that involves both transcribed and untranscribed speech utterances. Experiments for phone recognition show that a maximum mutual information criterion regularized by negative conditional entropy measured using unlabeled data reliably gives better results than other semi-supervised training methods. In the second part of the thesis, we propose to exploit unlabeled data for the task of automatic prosodic event detection. Prosody annotation is even harder to obtain than orthographic text transcription; it usually requires the expert knowledge of phonetics and linguistics. Therefore, we aim at reducing the annotation efforts for building an automatic prosodic event detector. We show that the mixture model has the ability of class discovery when labeled data are available from only one of the two classes and develop the learning algorithm for unsupervised prosodic boundary detection.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2012-04-13T13:13:33Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Huang_JuiTing.pdf: 1267213 bytes, checksum: d8d34c62c00be7b17fe910612b882764 (MD5)","Made available in DSpace on 2012-06-27T21:24:08Z (GMT). No. of bitstreams: 2 Huang_JuiTing.pdf: 1267213 bytes, checksum: d8d34c62c00be7b17fe910612b882764 (MD5) license.txt: 4059 bytes, checksum: 2d1913a9efdc9fb521fbd5dbede633a8 (MD5)","Item marked as restricted to the 'UIUC Users [automated]' Group (id=2) by William Ingram (wingram2@illinois.edu) on 2012-06-27T21:24:49Z Item is restricted until 2014-06-27T21:24:27Z","Item reinstated by Sarah Shreeves (sshreeve@illinois.edu) on 2014-06-28T10:00:25Z Item was in collections: Graduate Theses and Dissertations at Illinois (ID: 204) No. of bitstreams: 2 Huang_JuiTing.pdf: 1267213 bytes, checksum: d8d34c62c00be7b17fe910612b882764 (MD5) license.txt: 4059 bytes, checksum: 2d1913a9efdc9fb521fbd5dbede633a8 (MD5)","Item released from any restrictions by Sarah Shreeves (sshreeve@illinois.edu) on 2014-06-28T10:00:25Z"]},{"key":"dc:title","label":"Title","values":["Semi-supervised learning for acoustic and prosodic modeling in speech applications"]}]}],"canonical_facts":{"dc:contributor":["Hasegawa-Johnson, Mark A.","Cole, Jennifer S.","Huang, Thomas S.","Levinson, Stephen E."],"dc:creator":["Huang, Jui Ting"],"dc:date":["2012-06-27T21:24:08Z","2014-06-28T10:00:25Z","2012-05"],"dc:description":["Enormous amounts of audio recordings of human speech are essential ingredients for building reliable statistical models for many speech applications, such as automatic speech recognition and automatic prosody detection. However, most of these speech data are not being utilized because they lack transcriptions. The goal of this thesis is to use untranscribed (unlabeled) data to improve the performance of models trained using only transcribed (labeled) data. We propose a unified semi-supervised learning framework for the problem of phone classification, phone recognition and prosody detection. The proposed approach will be particularly useful in the case where recognition performance is limited by the amount of transcribed data. In the first part of the thesis, we investigate semi-supervised training of Gaussian Mixtures Models (GMMs) and Hidden Markov Models (HMMs) which are the common probabilistic models of acoustic features in a state-of-the-art continuous density HMM based speech recognition system. Specifically, a family of semi-supervised training criteria that reflects reasonable assumptions about labeled and unlabeled data is proposed. Both generative and discriminative kinds of training criteria are explored, and one important proposal of this thesis is to keep the power of discriminative training criteria by using some measures on unlabeled data as regularization to the supervised training objective. Methods are described for the optimization of these criteria, and phone classification experiments show that these criteria reliably give improvements over their supervised versions that use only labeled data. We then extend the proposed semi-supervised training criteria to the phone recognition problem. This problem is novel in the area of semi-supervised learning because there is little research on the use of unlabeled data in the sequence labeling problems. We develop lattice-based approaches for the model optimization that involves both transcribed and untranscribed speech utterances. Experiments for phone recognition show that a maximum mutual information criterion regularized by negative conditional entropy measured using unlabeled data reliably gives better results than other semi-supervised training methods. In the second part of the thesis, we propose to exploit unlabeled data for the task of automatic prosodic event detection. Prosody annotation is even harder to obtain than orthographic text transcription; it usually requires the expert knowledge of phonetics and linguistics. Therefore, we aim at reducing the annotation efforts for building an automatic prosodic event detector. We show that the mixture model has the ability of class discovery when labeled data are available from only one of the two classes and develop the learning algorithm for unsupervised prosodic boundary detection.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2012-04-13T13:13:33Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Huang_JuiTing.pdf: 1267213 bytes, checksum: d8d34c62c00be7b17fe910612b882764 (MD5)","Made available in DSpace on 2012-06-27T21:24:08Z (GMT). No. of bitstreams: 2 Huang_JuiTing.pdf: 1267213 bytes, checksum: d8d34c62c00be7b17fe910612b882764 (MD5) license.txt: 4059 bytes, checksum: 2d1913a9efdc9fb521fbd5dbede633a8 (MD5)","Item marked as restricted to the 'UIUC Users [automated]' Group (id=2) by William Ingram (wingram2@illinois.edu) on 2012-06-27T21:24:49Z Item is restricted until 2014-06-27T21:24:27Z","Item reinstated by Sarah Shreeves (sshreeve@illinois.edu) on 2014-06-28T10:00:25Z Item was in collections: Graduate Theses and Dissertations at Illinois (ID: 204) No. of bitstreams: 2 Huang_JuiTing.pdf: 1267213 bytes, checksum: d8d34c62c00be7b17fe910612b882764 (MD5) license.txt: 4059 bytes, checksum: 2d1913a9efdc9fb521fbd5dbede633a8 (MD5)","Item released from any restrictions by Sarah Shreeves (sshreeve@illinois.edu) on 2014-06-28T10:00:25Z"],"dc:identifier":["http://hdl.handle.net/2142/32006"],"dc:language":["en"],"dc:rights":["Copyright 2012 Jui Ting Huang"],"dc:subject":["Semi-Supervised Learning","Speech Recognition","Acoustic Modeling","Prosodic Modeling"],"dc:title":["Semi-supervised learning for acoustic and prosodic modeling in speech applications"],"dc:type":["text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:30Z"}