{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/98268"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/98268","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Application of generative models in speech processing tasks","abstract":"Generative probabilistic and neural models of the speech signal are shown to be effective in speech synthesis and speech enhancement, where generating natural and clean speech is the goal. This thesis develops two probabilistic signal processing algorithms based on the source-filter model of speech production, and two based on neural generative models of the speech signal. They are a model-based speech enhancement algorithm with ad-hoc microphone array, called GRAB; a probabilistic generative model of speech called PAT; a neural generative F0 model called TEReTA; and a Bayesian enhancement network, call BaWN, that incorporates a neural generative model of speech, called WaveNet. PAT and TEReTA aim to develop better generative models for speech synthesis. BaWN and GRAB aim to improve the naturalness and noise robustness of speech enhancement algorithms. Probabilistic Acoustic Tube (PAT) is a probabilistic generative model for speech, whose basis is the source-filter model. The highlights of the model are threefold. First, it is among the very first works to build a complete probabilistic model for speech. Second, it has a well-designed model for the phase spectrum of speech, which has been hard to model and often neglected. Third, it models the AM-FM effects in speech, which are perceptually significant but often ignored in frame-based speech processing algorithms. Experiments show that the proposed model has good potential for a number of speech processing tasks. TEReTA generates pitch contours by incorporating a theoretical model of pitch planning, the piece-wise linear target approximation (TA) model, as the output layer of a deep recurrent neural network. It aims to model semantic variations in the F0 contour, which is challenging for existing network. By combining the TA model, TEReTA is able to memorize semantic context and capture the semantic variations. Experiments on contrastive focus verify TEReTA's ability in semantics modeling. BaWN is a neural network based algorithm for single-channel enhancement. The biggest challenges of the neural network based speech enhancement algorithm are the poor generalizability to unseen noises and unnaturalness of the output speech. By incorporating a neural generative model, WaveNet, in the Bayesian framework, where WaveNet predicts the prior for speech, and where a separate enhancement network incorporates the likelihood function, BaWN is able to achieve satisfactory generalizability and a good intelligibility score of its output, even when the noisy training set is small. GRAB is a beamforming algorithm for ad-hoc microphone arrays. The task of enhancing speech with ad-hoc microphone array is challenging because of the inaccuracy in position and interference calibration. Inspired by the source-filter model, GRAB does not rely on any position or interference calibration. Instead, it incorporates a source-filter speech model and minimizes the energy that cannot be accounted for by the model. Objective and subjective evaluations on both simulated and real-world data show that GRAB is able to suppress noise effectively while keeping the speech natural and dry. Final chapters discuss the implications of this work for future research in speech processing.","abstract_html":"Generative probabilistic and neural models of the speech signal are shown to be effective in speech synthesis and speech enhancement, where generating natural and clean speech is the goal. This thesis develops two probabilistic signal processing algorithms based on the source-filter model of speech production, and two based on neural generative models of the speech signal. They are a model-based speech enhancement algorithm with ad-hoc microphone array, called GRAB; a probabilistic generative model of speech called PAT; a neural generative F0 model called TEReTA; and a Bayesian enhancement network, call BaWN, that incorporates a neural generative model of speech, called WaveNet. PAT and TEReTA aim to develop better generative models for speech synthesis. BaWN and GRAB aim to improve the naturalness and noise robustness of speech enhancement algorithms. Probabilistic Acoustic Tube (PAT) is a probabilistic generative model for speech, whose basis is the source-filter model. The highlights of the model are threefold. First, it is among the very first works to build a complete probabilistic model for speech. Second, it has a well-designed model for the phase spectrum of speech, which has been hard to model and often neglected. Third, it models the AM-FM effects in speech, which are perceptually significant but often ignored in frame-based speech processing algorithms. Experiments show that the proposed model has good potential for a number of speech processing tasks. TEReTA generates pitch contours by incorporating a theoretical model of pitch planning, the piece-wise linear target approximation (TA) model, as the output layer of a deep recurrent neural network. It aims to model semantic variations in the F0 contour, which is challenging for existing network. By combining the TA model, TEReTA is able to memorize semantic context and capture the semantic variations. Experiments on contrastive focus verify TEReTA&#x27;s ability in semantics modeling. BaWN is a neural network based algorithm for single-channel enhancement. The biggest challenges of the neural network based speech enhancement algorithm are the poor generalizability to unseen noises and unnaturalness of the output speech. By incorporating a neural generative model, WaveNet, in the Bayesian framework, where WaveNet predicts the prior for speech, and where a separate enhancement network incorporates the likelihood function, BaWN is able to achieve satisfactory generalizability and a good intelligibility score of its output, even when the noisy training set is small. GRAB is a beamforming algorithm for ad-hoc microphone arrays. The task of enhancing speech with ad-hoc microphone array is challenging because of the inaccuracy in position and interference calibration. Inspired by the source-filter model, GRAB does not rely on any position or interference calibration. Instead, it incorporates a source-filter speech model and minimizes the energy that cannot be accounted for by the model. Objective and subjective evaluations on both simulated and real-world data show that GRAB is able to suppress noise effectively while keeping the speech natural and dry. Final chapters discuss the implications of this work for future research in speech processing.","abstract_has_math":false,"creators":["Zhang, Yang"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Hasegawa-Johnson, Mark A.","Huang, Thomas S.","Levinson, Steven E.","Varshney, Lav R."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2017,"date_issued":"2017-09-29T16:39:40Z","date_published":"2017-09-29T16:39:40Z","updated_at":"2026-07-22T22:24:35Z","subjects":["Generative models","Speech synthesis","Speech enhancement"],"languages":["en"],"rights":["Copyright 2017 Yang Zhang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/98268","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Hasegawa-Johnson, Mark A.","Huang, Thomas S.","Levinson, Steven E.","Varshney, Lav R."]},{"key":"dc:creator","label":"Author","values":["Zhang, Yang"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2017-09-29T16:39:40Z","2019-09-30T09:15:14Z","2017-07-11","2017-08"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Generative models","Speech synthesis","Speech enhancement"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2017 Yang Zhang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/98268"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Generative probabilistic and neural models of the speech signal are shown to be effective in speech synthesis and speech enhancement, where generating natural and clean speech is the goal. This thesis develops two probabilistic signal processing algorithms based on the source-filter model of speech production, and two based on neural generative models of the speech signal. They are a model-based speech enhancement algorithm with ad-hoc microphone array, called GRAB; a probabilistic generative model of speech called PAT; a neural generative F0 model called TEReTA; and a Bayesian enhancement network, call BaWN, that incorporates a neural generative model of speech, called WaveNet. PAT and TEReTA aim to develop better generative models for speech synthesis. BaWN and GRAB aim to improve the naturalness and noise robustness of speech enhancement algorithms. Probabilistic Acoustic Tube (PAT) is a probabilistic generative model for speech, whose basis is the source-filter model. The highlights of the model are threefold. First, it is among the very first works to build a complete probabilistic model for speech. Second, it has a well-designed model for the phase spectrum of speech, which has been hard to model and often neglected. Third, it models the AM-FM effects in speech, which are perceptually significant but often ignored in frame-based speech processing algorithms. Experiments show that the proposed model has good potential for a number of speech processing tasks. TEReTA generates pitch contours by incorporating a theoretical model of pitch planning, the piece-wise linear target approximation (TA) model, as the output layer of a deep recurrent neural network. It aims to model semantic variations in the F0 contour, which is challenging for existing network. By combining the TA model, TEReTA is able to memorize semantic context and capture the semantic variations. Experiments on contrastive focus verify TEReTA's ability in semantics modeling. BaWN is a neural network based algorithm for single-channel enhancement. The biggest challenges of the neural network based speech enhancement algorithm are the poor generalizability to unseen noises and unnaturalness of the output speech. By incorporating a neural generative model, WaveNet, in the Bayesian framework, where WaveNet predicts the prior for speech, and where a separate enhancement network incorporates the likelihood function, BaWN is able to achieve satisfactory generalizability and a good intelligibility score of its output, even when the noisy training set is small. GRAB is a beamforming algorithm for ad-hoc microphone arrays. The task of enhancing speech with ad-hoc microphone array is challenging because of the inaccuracy in position and interference calibration. Inspired by the source-filter model, GRAB does not rely on any position or interference calibration. Instead, it incorporates a source-filter speech model and minimizes the energy that cannot be accounted for by the model. Objective and subjective evaluations on both simulated and real-world data show that GRAB is able to suppress noise effectively while keeping the speech natural and dry. Final chapters discuss the implications of this work for future research in speech processing.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2019-08-01","The student, Yang Zhang, accepted the attached license on 2017-07-10 at 16:26.","The student, Yang Zhang, submitted this Dissertation for approval on 2017-07-10 at 16:39.","This Dissertation was approved for publication on 2017-07-11 at 14:59.","DSpace SAF Submission Ingestion Package generated from Vireo submission #11381 on 2017-09-29 at 11:18:23","Made available in DSpace on 2017-09-29T16:39:40Z (GMT). No. of bitstreams: 14 ZHANG-DISSERTATION-2017.pdf: 2699518 bytes, checksum: f6a4ff37b154acb351cca91eeaa3d664 (MD5) RNN-TA.tex: 55115 bytes, checksum: 374ae69d2d964e634325661db68dad9c (MD5) abs.tex: 3268 bytes, checksum: f1c26a4878018bb33e3f823bf23fff98 (MD5) ack.tex: 787 bytes, checksum: 002b87d69b50476016a7a64772481c79 (MD5) background.tex: 30326 bytes, checksum: cb6180e97422f6e0d05c8f9673220445 (MD5) bawn.tex: 32034 bytes, checksum: 4c1f745a319de01c8debf6c859b9d58f (MD5) conclusion.tex: 3605 bytes, checksum: 2cf41819c573db0f58c19a4b064a200b (MD5) discussion.tex: 14077 bytes, checksum: 1d5631e51de58a03c0ae5333c63425b2 (MD5) ecethesis.tex: 5427 bytes, checksum: 7c27ae8ba3d055e26dab65cd4363ad4a (MD5) grab.tex: 36408 bytes, checksum: 272981d32c2d17dd3db289f601cbaf66 (MD5) motivation.tex: 10022 bytes, checksum: 55c4274a2910308d11c82fd50275f6de (MD5) pat.tex: 64227 bytes, checksum: 364d452d66d2e76576ff695759187c1e (MD5) LICENSE.txt: 4207 bytes, checksum: 518abee862f9c37d7fd11130c21d113a (MD5) PROQUEST_LICENSE.txt: 4553 bytes, checksum: 3b01c71182e971c31ce41123daf602f9 (MD5) Previous issue date: 2017-07-11","Embargo set by: Colleen Fallaw for item 103415 Lift date: 2019-09-29T16:39:52Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Embargo set by: Colleen Fallaw for item 103415 Lift date: 2019-09-29T17:52:45Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 103415 on 2019-09-30T09:15:14Z."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Application of generative models in speech processing tasks"]}]}],"canonical_facts":{"dc:contributor":["Hasegawa-Johnson, Mark A.","Huang, Thomas S.","Levinson, Steven E.","Varshney, Lav R."],"dc:creator":["Zhang, Yang"],"dc:date":["2017-09-29T16:39:40Z","2019-09-30T09:15:14Z","2017-07-11","2017-08"],"dc:description":["Generative probabilistic and neural models of the speech signal are shown to be effective in speech synthesis and speech enhancement, where generating natural and clean speech is the goal. This thesis develops two probabilistic signal processing algorithms based on the source-filter model of speech production, and two based on neural generative models of the speech signal. They are a model-based speech enhancement algorithm with ad-hoc microphone array, called GRAB; a probabilistic generative model of speech called PAT; a neural generative F0 model called TEReTA; and a Bayesian enhancement network, call BaWN, that incorporates a neural generative model of speech, called WaveNet. PAT and TEReTA aim to develop better generative models for speech synthesis. BaWN and GRAB aim to improve the naturalness and noise robustness of speech enhancement algorithms. Probabilistic Acoustic Tube (PAT) is a probabilistic generative model for speech, whose basis is the source-filter model. The highlights of the model are threefold. First, it is among the very first works to build a complete probabilistic model for speech. Second, it has a well-designed model for the phase spectrum of speech, which has been hard to model and often neglected. Third, it models the AM-FM effects in speech, which are perceptually significant but often ignored in frame-based speech processing algorithms. Experiments show that the proposed model has good potential for a number of speech processing tasks. TEReTA generates pitch contours by incorporating a theoretical model of pitch planning, the piece-wise linear target approximation (TA) model, as the output layer of a deep recurrent neural network. It aims to model semantic variations in the F0 contour, which is challenging for existing network. By combining the TA model, TEReTA is able to memorize semantic context and capture the semantic variations. Experiments on contrastive focus verify TEReTA's ability in semantics modeling. BaWN is a neural network based algorithm for single-channel enhancement. The biggest challenges of the neural network based speech enhancement algorithm are the poor generalizability to unseen noises and unnaturalness of the output speech. By incorporating a neural generative model, WaveNet, in the Bayesian framework, where WaveNet predicts the prior for speech, and where a separate enhancement network incorporates the likelihood function, BaWN is able to achieve satisfactory generalizability and a good intelligibility score of its output, even when the noisy training set is small. GRAB is a beamforming algorithm for ad-hoc microphone arrays. The task of enhancing speech with ad-hoc microphone array is challenging because of the inaccuracy in position and interference calibration. Inspired by the source-filter model, GRAB does not rely on any position or interference calibration. Instead, it incorporates a source-filter speech model and minimizes the energy that cannot be accounted for by the model. Objective and subjective evaluations on both simulated and real-world data show that GRAB is able to suppress noise effectively while keeping the speech natural and dry. Final chapters discuss the implications of this work for future research in speech processing.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2019-08-01","The student, Yang Zhang, accepted the attached license on 2017-07-10 at 16:26.","The student, Yang Zhang, submitted this Dissertation for approval on 2017-07-10 at 16:39.","This Dissertation was approved for publication on 2017-07-11 at 14:59.","DSpace SAF Submission Ingestion Package generated from Vireo submission #11381 on 2017-09-29 at 11:18:23","Made available in DSpace on 2017-09-29T16:39:40Z (GMT). No. of bitstreams: 14 ZHANG-DISSERTATION-2017.pdf: 2699518 bytes, checksum: f6a4ff37b154acb351cca91eeaa3d664 (MD5) RNN-TA.tex: 55115 bytes, checksum: 374ae69d2d964e634325661db68dad9c (MD5) abs.tex: 3268 bytes, checksum: f1c26a4878018bb33e3f823bf23fff98 (MD5) ack.tex: 787 bytes, checksum: 002b87d69b50476016a7a64772481c79 (MD5) background.tex: 30326 bytes, checksum: cb6180e97422f6e0d05c8f9673220445 (MD5) bawn.tex: 32034 bytes, checksum: 4c1f745a319de01c8debf6c859b9d58f (MD5) conclusion.tex: 3605 bytes, checksum: 2cf41819c573db0f58c19a4b064a200b (MD5) discussion.tex: 14077 bytes, checksum: 1d5631e51de58a03c0ae5333c63425b2 (MD5) ecethesis.tex: 5427 bytes, checksum: 7c27ae8ba3d055e26dab65cd4363ad4a (MD5) grab.tex: 36408 bytes, checksum: 272981d32c2d17dd3db289f601cbaf66 (MD5) motivation.tex: 10022 bytes, checksum: 55c4274a2910308d11c82fd50275f6de (MD5) pat.tex: 64227 bytes, checksum: 364d452d66d2e76576ff695759187c1e (MD5) LICENSE.txt: 4207 bytes, checksum: 518abee862f9c37d7fd11130c21d113a (MD5) PROQUEST_LICENSE.txt: 4553 bytes, checksum: 3b01c71182e971c31ce41123daf602f9 (MD5) Previous issue date: 2017-07-11","Embargo set by: Colleen Fallaw for item 103415 Lift date: 2019-09-29T16:39:52Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Embargo set by: Colleen Fallaw for item 103415 Lift date: 2019-09-29T17:52:45Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 103415 on 2019-09-30T09:15:14Z."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/98268"],"dc:language":["en"],"dc:rights":["Copyright 2017 Yang Zhang"],"dc:subject":["Generative models","Speech synthesis","Speech enhancement"],"dc:title":["Application of generative models in speech processing tasks"],"dc:type":["text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:35Z"}