{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/89006"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/89006","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Probabilistic generative modeling of speech","abstract":"Speech processing refers to a set of tasks that involve speech analysis and synthesis. Most speech processing algorithms model a subset of speech parameters of interest and blur the rest using signal processing techniques and feature extraction. However, evidence shows that many speech parameters can be more accurately estimated if they are modeled jointly; speech synthesis also benefits from joint modeling. This thesis proposes a probabilistic generative model for speech called the Probabilistic Acoustic Tube (PAT). The highlights of the model are threefold. First, it is among the very first works to build a complete probabilistic model for speech. Second, it has a well-designed model for the phase spectrum of speech, which has been hard to model and often neglected. Third, it models the AM-FM effects in speech, which are perceptually significant but often ignored in frame-based speech processing algorithms. Experiment shows that the proposed model has good potential for a number of speech processing tasks.","abstract_html":"Speech processing refers to a set of tasks that involve speech analysis and synthesis. Most speech processing algorithms model a subset of speech parameters of interest and blur the rest using signal processing techniques and feature extraction. However, evidence shows that many speech parameters can be more accurately estimated if they are modeled jointly; speech synthesis also benefits from joint modeling. This thesis proposes a probabilistic generative model for speech called the Probabilistic Acoustic Tube (PAT). The highlights of the model are threefold. First, it is among the very first works to build a complete probabilistic model for speech. Second, it has a well-designed model for the phase spectrum of speech, which has been hard to model and often neglected. Third, it models the AM-FM effects in speech, which are perceptually significant but often ignored in frame-based speech processing algorithms. Experiment shows that the proposed model has good potential for a number of speech processing tasks.","abstract_has_math":false,"creators":["Zhang, Yang"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Electrical & Computer Engineering","degree_department":null,"school":null,"contributors":["Hasegawa-Johnson, Mark A."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2016,"date_issued":"2016-03-02T19:33:51Z","date_published":"2016-03-02T19:33:51Z","updated_at":"2026-07-22T22:26:32Z","subjects":["Probabilistic acoustic tube","speech modeling","speech analysis","generative model"],"languages":["en"],"rights":["Copyright 2015 Yang Zhang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/89006","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Hasegawa-Johnson, Mark A."]},{"key":"dc:creator","label":"Author","values":["Zhang, Yang"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2016-03-02T19:33:51Z","2015-11-24","2015-12"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engineering"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Probabilistic acoustic tube","speech modeling","speech analysis","generative model"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2015 Yang Zhang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/89006"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Speech processing refers to a set of tasks that involve speech analysis and synthesis. Most speech processing algorithms model a subset of speech parameters of interest and blur the rest using signal processing techniques and feature extraction. However, evidence shows that many speech parameters can be more accurately estimated if they are modeled jointly; speech synthesis also benefits from joint modeling. This thesis proposes a probabilistic generative model for speech called the Probabilistic Acoustic Tube (PAT). The highlights of the model are threefold. First, it is among the very first works to build a complete probabilistic model for speech. Second, it has a well-designed model for the phase spectrum of speech, which has been hard to model and often neglected. Third, it models the AM-FM effects in speech, which are perceptually significant but often ignored in frame-based speech processing algorithms. Experiment shows that the proposed model has good potential for a number of speech processing tasks.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2016-03-02 without embargo terms","The student, Yang Zhang, accepted the attached license on 2015-11-23 at 17:38.","The student, Yang Zhang, submitted this Thesis for approval on 2015-11-23 at 17:50.","This Thesis was approved for publication on 2015-11-24 at 11:11.","DSpace SAF Submission Ingestion Package generated from Vireo submission #8832 on 2016-03-02 at 12:50:23","Made available in DSpace on 2016-03-02T19:33:51Z (GMT). No. of bitstreams: 2 ZHANG-THESIS-2015.pdf: 555655 bytes, checksum: 55561005b4b31b9f6d4d93053c79a006 (MD5) LICENSE.txt: 4207 bytes, checksum: 44f5a50269d4b23b36acfe18ae8dffce (MD5) Previous issue date: 2015-11-24"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Probabilistic generative modeling of speech"]}]}],"canonical_facts":{"dc:contributor":["Hasegawa-Johnson, Mark A."],"dc:creator":["Zhang, Yang"],"dc:date":["2016-03-02T19:33:51Z","2015-11-24","2015-12"],"dc:description":["Speech processing refers to a set of tasks that involve speech analysis and synthesis. Most speech processing algorithms model a subset of speech parameters of interest and blur the rest using signal processing techniques and feature extraction. However, evidence shows that many speech parameters can be more accurately estimated if they are modeled jointly; speech synthesis also benefits from joint modeling. This thesis proposes a probabilistic generative model for speech called the Probabilistic Acoustic Tube (PAT). The highlights of the model are threefold. First, it is among the very first works to build a complete probabilistic model for speech. Second, it has a well-designed model for the phase spectrum of speech, which has been hard to model and often neglected. Third, it models the AM-FM effects in speech, which are perceptually significant but often ignored in frame-based speech processing algorithms. Experiment shows that the proposed model has good potential for a number of speech processing tasks.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2016-03-02 without embargo terms","The student, Yang Zhang, accepted the attached license on 2015-11-23 at 17:38.","The student, Yang Zhang, submitted this Thesis for approval on 2015-11-23 at 17:50.","This Thesis was approved for publication on 2015-11-24 at 11:11.","DSpace SAF Submission Ingestion Package generated from Vireo submission #8832 on 2016-03-02 at 12:50:23","Made available in DSpace on 2016-03-02T19:33:51Z (GMT). No. of bitstreams: 2 ZHANG-THESIS-2015.pdf: 555655 bytes, checksum: 55561005b4b31b9f6d4d93053c79a006 (MD5) LICENSE.txt: 4207 bytes, checksum: 44f5a50269d4b23b36acfe18ae8dffce (MD5) Previous issue date: 2015-11-24"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/89006"],"dc:language":["en"],"dc:rights":["Copyright 2015 Yang Zhang"],"dc:subject":["Probabilistic acoustic tube","speech modeling","speech analysis","generative model"],"dc:title":["Probabilistic generative modeling of speech"],"dc:type":["text"],"thesis:degree_discipline":["Electrical & Computer Engineering"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:26:32Z"}