{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/109510"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/109510","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Deep generative models for speech editing","abstract":"Generative models are very useful for generating and modifying natural-sounding speech in various speech processing tasks such as speech synthesis, speech enhancement, and voice conversion. There are two ways that the generative models can help in naturalness for speech processing. The first way is to regularize the speech editing process by defining the sample space of natural speech, and the second way is by permitting the separable modification of components of hierarchical speech generative models to modify specified components of natural speech. In particular, four research projects are introduced, where the first two use WaveNet as the clean speech generative model for single-channel and multi-channel speech enhancement; and the last two projects modify different speaking styles by modeling different speech components using autoencoders. Multi-channel speech enhancement with ad-hoc sensors has been a challenging task. Speech model guided beamforming algorithms can recover natural-sounding speech, but the speech models tend to be oversimplified to prevent the inference from becoming too complicated. On the other hand, deep learning-based enhancement approaches can learn complicated speech distributions and perform efficient inference, but they are unable to deal with a variable number of input channels. Also, deep learning approaches introduce many errors, particularly in the presence of unseen noise types and settings. Therefore an enhancement framework called DeepBeam is proposed, which combines the two complementary classes of algorithms. DeepBeam introduces a beamforming filter to produce natural-sounding speech, but the filter coefficients are determined with the help of a WaveNet-based monaural speech enhancement model. Experiments on synthetic and real-world data show that DeepBeam can produce clean, dry, and natural-sounding speech, and is robust against unseen noise. For single-channel speech enhancement, the existing deep learning-based methods still have two limitations. First, the Bayesian framework is not adopted in many such deep-learning-based algorithms. Second, the majority of the existing methods operate on the frequency domain of the noisy speech, such as the spectrogram and its variations. A Bayesian speech enhancement framework, called BaWN (Bayesian WaveNet) is proposed, which directly operates on raw audio samples. It uses the WaveNet as the prior model to regularize the output to be in the speech space and thus improving the performance. Experiments show that BaWN can recover clean and natural speech. Non-parallel many-to-many voice conversion, as well as zero-shot voice conversion, remain under-explored areas. Deep style transfer algorithms, such as generative adversarial networks (GAN) and conditional variational autoencoder (CVAE), are popular solutions in this field. However, GAN training is sophisticated and difficult, and there is no strong evidence that its generated speech is of good perceptual quality. On the other hand, CVAE training is simple but does not come with the distribution-matching property as in GAN. A new style transfer scheme that involves only an autoencoder with a carefully designed bottleneck is proposed. This scheme can achieve distribution-matching style transfer by training only on a self-reconstruction loss. Based on this scheme, AutoVC is proposed, which achieves state-of-the-art results in many-to-many voice conversion with non-parallel data, and which is also the first to perform zero-shot voice conversion. Speech information can be roughly decomposed into four components: language content, timbre, pitch, and rhythm. Obtaining disentangled representations of these components is useful in many speech analysis and generation applications. Recently, state-of-the-art voice conversion systems have led to speech representations that can disentangle speaker-dependent and independent information. However, these systems can only disentangle timbre, while information about pitch, rhythm, and content is still mixed. Further disentangling the remaining speech components is an under-determined problem in the absence of explicit annotations for each component, which are difficult and expensive to obtain. To further explore this problem, we propose SpeechSplit, which can blindly decompose speech into its four components by introducing three carefully designed information bottlenecks. SpeechSplit is among the first algorithms that can separately perform style transfer on timbre, pitch, and rhythm without text labels.","abstract_html":"Generative models are very useful for generating and modifying natural-sounding speech in various speech processing tasks such as speech synthesis, speech enhancement, and voice conversion. There are two ways that the generative models can help in naturalness for speech processing. The first way is to regularize the speech editing process by defining the sample space of natural speech, and the second way is by permitting the separable modification of components of hierarchical speech generative models to modify specified components of natural speech. In particular, four research projects are introduced, where the first two use WaveNet as the clean speech generative model for single-channel and multi-channel speech enhancement; and the last two projects modify different speaking styles by modeling different speech components using autoencoders. Multi-channel speech enhancement with ad-hoc sensors has been a challenging task. Speech model guided beamforming algorithms can recover natural-sounding speech, but the speech models tend to be oversimplified to prevent the inference from becoming too complicated. On the other hand, deep learning-based enhancement approaches can learn complicated speech distributions and perform efficient inference, but they are unable to deal with a variable number of input channels. Also, deep learning approaches introduce many errors, particularly in the presence of unseen noise types and settings. Therefore an enhancement framework called DeepBeam is proposed, which combines the two complementary classes of algorithms. DeepBeam introduces a beamforming filter to produce natural-sounding speech, but the filter coefficients are determined with the help of a WaveNet-based monaural speech enhancement model. Experiments on synthetic and real-world data show that DeepBeam can produce clean, dry, and natural-sounding speech, and is robust against unseen noise. For single-channel speech enhancement, the existing deep learning-based methods still have two limitations. First, the Bayesian framework is not adopted in many such deep-learning-based algorithms. Second, the majority of the existing methods operate on the frequency domain of the noisy speech, such as the spectrogram and its variations. A Bayesian speech enhancement framework, called BaWN (Bayesian WaveNet) is proposed, which directly operates on raw audio samples. It uses the WaveNet as the prior model to regularize the output to be in the speech space and thus improving the performance. Experiments show that BaWN can recover clean and natural speech. Non-parallel many-to-many voice conversion, as well as zero-shot voice conversion, remain under-explored areas. Deep style transfer algorithms, such as generative adversarial networks (GAN) and conditional variational autoencoder (CVAE), are popular solutions in this field. However, GAN training is sophisticated and difficult, and there is no strong evidence that its generated speech is of good perceptual quality. On the other hand, CVAE training is simple but does not come with the distribution-matching property as in GAN. A new style transfer scheme that involves only an autoencoder with a carefully designed bottleneck is proposed. This scheme can achieve distribution-matching style transfer by training only on a self-reconstruction loss. Based on this scheme, AutoVC is proposed, which achieves state-of-the-art results in many-to-many voice conversion with non-parallel data, and which is also the first to perform zero-shot voice conversion. Speech information can be roughly decomposed into four components: language content, timbre, pitch, and rhythm. Obtaining disentangled representations of these components is useful in many speech analysis and generation applications. Recently, state-of-the-art voice conversion systems have led to speech representations that can disentangle speaker-dependent and independent information. However, these systems can only disentangle timbre, while information about pitch, rhythm, and content is still mixed. Further disentangling the remaining speech components is an under-determined problem in the absence of explicit annotations for each component, which are difficult and expensive to obtain. To further explore this problem, we propose SpeechSplit, which can blindly decompose speech into its four components by introducing three carefully designed information bottlenecks. SpeechSplit is among the first algorithms that can separately perform style transfer on timbre, pitch, and rhythm without text labels.","abstract_has_math":false,"creators":["Qian, Kaizhi"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Hasegawa-Johnson, Mark A","Levinson, Steven E","Varshney, Lav R","Chang, Shiyu"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2021,"date_issued":"2021-03-05T21:42:42Z","date_published":"2021-03-05T21:42:42Z","updated_at":"2026-07-22T22:24:50Z","subjects":["Generative model","Speech enhancement","Speech disentanglement"],"languages":["en"],"rights":["Copyright 2020 Kaizhi Qian"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/109510","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Hasegawa-Johnson, Mark A","Levinson, Steven E","Varshney, Lav R","Chang, Shiyu"]},{"key":"dc:creator","label":"Author","values":["Qian, Kaizhi"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2021-03-05T21:42:42Z","2023-03-05T21:43:00Z","2020-12-01","2020-12"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Generative model","Speech enhancement","Speech disentanglement"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2020 Kaizhi Qian"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/109510"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Generative models are very useful for generating and modifying natural-sounding speech in various speech processing tasks such as speech synthesis, speech enhancement, and voice conversion. There are two ways that the generative models can help in naturalness for speech processing. The first way is to regularize the speech editing process by defining the sample space of natural speech, and the second way is by permitting the separable modification of components of hierarchical speech generative models to modify specified components of natural speech. In particular, four research projects are introduced, where the first two use WaveNet as the clean speech generative model for single-channel and multi-channel speech enhancement; and the last two projects modify different speaking styles by modeling different speech components using autoencoders. Multi-channel speech enhancement with ad-hoc sensors has been a challenging task. Speech model guided beamforming algorithms can recover natural-sounding speech, but the speech models tend to be oversimplified to prevent the inference from becoming too complicated. On the other hand, deep learning-based enhancement approaches can learn complicated speech distributions and perform efficient inference, but they are unable to deal with a variable number of input channels. Also, deep learning approaches introduce many errors, particularly in the presence of unseen noise types and settings. Therefore an enhancement framework called DeepBeam is proposed, which combines the two complementary classes of algorithms. DeepBeam introduces a beamforming filter to produce natural-sounding speech, but the filter coefficients are determined with the help of a WaveNet-based monaural speech enhancement model. Experiments on synthetic and real-world data show that DeepBeam can produce clean, dry, and natural-sounding speech, and is robust against unseen noise. For single-channel speech enhancement, the existing deep learning-based methods still have two limitations. First, the Bayesian framework is not adopted in many such deep-learning-based algorithms. Second, the majority of the existing methods operate on the frequency domain of the noisy speech, such as the spectrogram and its variations. A Bayesian speech enhancement framework, called BaWN (Bayesian WaveNet) is proposed, which directly operates on raw audio samples. It uses the WaveNet as the prior model to regularize the output to be in the speech space and thus improving the performance. Experiments show that BaWN can recover clean and natural speech. Non-parallel many-to-many voice conversion, as well as zero-shot voice conversion, remain under-explored areas. Deep style transfer algorithms, such as generative adversarial networks (GAN) and conditional variational autoencoder (CVAE), are popular solutions in this field. However, GAN training is sophisticated and difficult, and there is no strong evidence that its generated speech is of good perceptual quality. On the other hand, CVAE training is simple but does not come with the distribution-matching property as in GAN. A new style transfer scheme that involves only an autoencoder with a carefully designed bottleneck is proposed. This scheme can achieve distribution-matching style transfer by training only on a self-reconstruction loss. Based on this scheme, AutoVC is proposed, which achieves state-of-the-art results in many-to-many voice conversion with non-parallel data, and which is also the first to perform zero-shot voice conversion. Speech information can be roughly decomposed into four components: language content, timbre, pitch, and rhythm. Obtaining disentangled representations of these components is useful in many speech analysis and generation applications. Recently, state-of-the-art voice conversion systems have led to speech representations that can disentangle speaker-dependent and independent information. However, these systems can only disentangle timbre, while information about pitch, rhythm, and content is still mixed. Further disentangling the remaining speech components is an under-determined problem in the absence of explicit annotations for each component, which are difficult and expensive to obtain. To further explore this problem, we propose SpeechSplit, which can blindly decompose speech into its four components by introducing three carefully designed information bottlenecks. SpeechSplit is among the first algorithms that can separately perform style transfer on timbre, pitch, and rhythm without text labels.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2022-12-01","The student, Kaizhi Qian, accepted the attached license on 2020-11-25 at 13:46.","The student, Kaizhi Qian, submitted this Dissertation for approval on 2020-11-25 at 13:56.","This Dissertation was approved for publication on 2020-12-01 at 11:23.","DSpace SAF Submission Ingestion Package generated from Vireo submission #15962 on 2021-03-04 at 16:19:55","Made available in DSpace on 2021-03-05T21:42:42Z (GMT). No. of bitstreams: 2 QIAN-DISSERTATION-2020.pdf: 3643017 bytes, checksum: 84a5c94a00dec7b2959506a633424a0c (MD5) LICENSE.txt: 4208 bytes, checksum: 289087ab5ba6e12c7c4ab8e9122fe0c1 (MD5) Previous issue date: 2020-12-01","Embargo set by: Seth Robbins for item 117215 Lift date: 2023-03-05T21:43:00Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Deep generative models for speech editing"]}]}],"canonical_facts":{"dc:contributor":["Hasegawa-Johnson, Mark A","Levinson, Steven E","Varshney, Lav R","Chang, Shiyu"],"dc:creator":["Qian, Kaizhi"],"dc:date":["2021-03-05T21:42:42Z","2023-03-05T21:43:00Z","2020-12-01","2020-12"],"dc:description":["Generative models are very useful for generating and modifying natural-sounding speech in various speech processing tasks such as speech synthesis, speech enhancement, and voice conversion. There are two ways that the generative models can help in naturalness for speech processing. The first way is to regularize the speech editing process by defining the sample space of natural speech, and the second way is by permitting the separable modification of components of hierarchical speech generative models to modify specified components of natural speech. In particular, four research projects are introduced, where the first two use WaveNet as the clean speech generative model for single-channel and multi-channel speech enhancement; and the last two projects modify different speaking styles by modeling different speech components using autoencoders. Multi-channel speech enhancement with ad-hoc sensors has been a challenging task. Speech model guided beamforming algorithms can recover natural-sounding speech, but the speech models tend to be oversimplified to prevent the inference from becoming too complicated. On the other hand, deep learning-based enhancement approaches can learn complicated speech distributions and perform efficient inference, but they are unable to deal with a variable number of input channels. Also, deep learning approaches introduce many errors, particularly in the presence of unseen noise types and settings. Therefore an enhancement framework called DeepBeam is proposed, which combines the two complementary classes of algorithms. DeepBeam introduces a beamforming filter to produce natural-sounding speech, but the filter coefficients are determined with the help of a WaveNet-based monaural speech enhancement model. Experiments on synthetic and real-world data show that DeepBeam can produce clean, dry, and natural-sounding speech, and is robust against unseen noise. For single-channel speech enhancement, the existing deep learning-based methods still have two limitations. First, the Bayesian framework is not adopted in many such deep-learning-based algorithms. Second, the majority of the existing methods operate on the frequency domain of the noisy speech, such as the spectrogram and its variations. A Bayesian speech enhancement framework, called BaWN (Bayesian WaveNet) is proposed, which directly operates on raw audio samples. It uses the WaveNet as the prior model to regularize the output to be in the speech space and thus improving the performance. Experiments show that BaWN can recover clean and natural speech. Non-parallel many-to-many voice conversion, as well as zero-shot voice conversion, remain under-explored areas. Deep style transfer algorithms, such as generative adversarial networks (GAN) and conditional variational autoencoder (CVAE), are popular solutions in this field. However, GAN training is sophisticated and difficult, and there is no strong evidence that its generated speech is of good perceptual quality. On the other hand, CVAE training is simple but does not come with the distribution-matching property as in GAN. A new style transfer scheme that involves only an autoencoder with a carefully designed bottleneck is proposed. This scheme can achieve distribution-matching style transfer by training only on a self-reconstruction loss. Based on this scheme, AutoVC is proposed, which achieves state-of-the-art results in many-to-many voice conversion with non-parallel data, and which is also the first to perform zero-shot voice conversion. Speech information can be roughly decomposed into four components: language content, timbre, pitch, and rhythm. Obtaining disentangled representations of these components is useful in many speech analysis and generation applications. Recently, state-of-the-art voice conversion systems have led to speech representations that can disentangle speaker-dependent and independent information. However, these systems can only disentangle timbre, while information about pitch, rhythm, and content is still mixed. Further disentangling the remaining speech components is an under-determined problem in the absence of explicit annotations for each component, which are difficult and expensive to obtain. To further explore this problem, we propose SpeechSplit, which can blindly decompose speech into its four components by introducing three carefully designed information bottlenecks. SpeechSplit is among the first algorithms that can separately perform style transfer on timbre, pitch, and rhythm without text labels.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2022-12-01","The student, Kaizhi Qian, accepted the attached license on 2020-11-25 at 13:46.","The student, Kaizhi Qian, submitted this Dissertation for approval on 2020-11-25 at 13:56.","This Dissertation was approved for publication on 2020-12-01 at 11:23.","DSpace SAF Submission Ingestion Package generated from Vireo submission #15962 on 2021-03-04 at 16:19:55","Made available in DSpace on 2021-03-05T21:42:42Z (GMT). No. of bitstreams: 2 QIAN-DISSERTATION-2020.pdf: 3643017 bytes, checksum: 84a5c94a00dec7b2959506a633424a0c (MD5) LICENSE.txt: 4208 bytes, checksum: 289087ab5ba6e12c7c4ab8e9122fe0c1 (MD5) Previous issue date: 2020-12-01","Embargo set by: Seth Robbins for item 117215 Lift date: 2023-03-05T21:43:00Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/109510"],"dc:language":["en"],"dc:rights":["Copyright 2020 Kaizhi Qian"],"dc:subject":["Generative model","Speech enhancement","Speech disentanglement"],"dc:title":["Deep generative models for speech editing"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:50Z"}