{"id":{"repo_id":"auckland-ms","oai_identifier":"oai:researchspace.auckland.ac.nz:2292/74747"},"canonical_url":"https://search.dev.ndltd.org/etd/auckland-ms/oai:researchspace.auckland.ac.nz:2292/74747","repository":{"repo_id":"auckland-ms","name":"University of Auckland","base_url":"https://researchspace.auckland.ac.nz/server/oai/request"},"display":{"title":"Assessing the Quality of Synthetic Speech when using Enhanced Speech as Training Data","abstract":"Both speech synthesis and speech enhancement are well researched fields, but their interaction remains under-explored. In particular, the effectiveness of using enhanced speech to train a speech synthesis model is still relatively unknown. This thesis investigates the effects of using enhanced speech to train a speech synthesis model on the overall quality of the synthesised speech, and in doing so, gain a better understanding of the interactions between different speech enhancement algorithms, speech synthesis architectures, gender of the voice and the noise types that degrade the speech used to train the speech synthesis model. A series of large-scale perception tests were conducted whereby over 100 participants evaluated the quality of utterances generated by speech synthesis model trained on noisy or enhanced speech. For this study, 160 synthetic voices were created, one for every combination of four speech enhancement methods, four noise types at two SNRs, two voices, one male and one female for two speech synthesis models – MaryTTS and Tacotron 2. Using linear mixed effect models, the study discerned the interactions between speech enhancement algorithms, speech synthesis architectures, gender of the voice and noise types. The results from this analysis suggest that different genders of voices resulted in similar trends depending on the speech enhancement methods, speech synthesis methods and noise types. However, the speech synthesis methods were affected differently by the various speech enhancement methods and noise types. An investigation into evaluating the quality of synthesised voices created from noisy or enhanced speech using eleven objective quality metrics was also conducted. This investigation found that several common objective quality metrics used to evaluate speech enhancement also have relatively high correlations with subjective quality of utterances generated by speech synthesisers trained on noisy or enhanced speech.","abstract_html":"Both speech synthesis and speech enhancement are well researched fields, but their interaction remains under-explored. In particular, the effectiveness of using enhanced speech to train a speech synthesis model is still relatively unknown. This thesis investigates the effects of using enhanced speech to train a speech synthesis model on the overall quality of the synthesised speech, and in doing so, gain a better understanding of the interactions between different speech enhancement algorithms, speech synthesis architectures, gender of the voice and the noise types that degrade the speech used to train the speech synthesis model. A series of large-scale perception tests were conducted whereby over 100 participants evaluated the quality of utterances generated by speech synthesis model trained on noisy or enhanced speech. For this study, 160 synthetic voices were created, one for every combination of four speech enhancement methods, four noise types at two SNRs, two voices, one male and one female for two speech synthesis models – MaryTTS and Tacotron 2. Using linear mixed effect models, the study discerned the interactions between speech enhancement algorithms, speech synthesis architectures, gender of the voice and noise types. The results from this analysis suggest that different genders of voices resulted in similar trends depending on the speech enhancement methods, speech synthesis methods and noise types. However, the speech synthesis methods were affected differently by the various speech enhancement methods and noise types. An investigation into evaluating the quality of synthesised voices created from noisy or enhanced speech using eleven objective quality metrics was also conducted. This investigation found that several common objective quality metrics used to evaluate speech enhancement also have relatively high correlations with subjective quality of utterances generated by speech synthesisers trained on noisy or enhanced speech.","abstract_has_math":false,"creators":["Eng, Nicholas"],"institution":"ResearchSpace@Auckland","degree_name":"PhD","degree_level":"Doctoral","degree_discipline":"Mechatronics Engineering","degree_department":null,"school":null,"contributors":[],"advisors":["Hioka, Yusuke","Watson, Catherine I"],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025","date_published":"2025","updated_at":"2026-07-24T01:06:29Z","subjects":["Speech Enhancement","Speech Synthesis","Text-To-Speech","Audio Quality","Speech Perception","Speech Quality Assessment","Objective Metrics","Deep Learning"],"languages":[],"rights":["Items in ResearchSpace are protected by copyright, with all rights reserved, unless otherwise indicated."],"rights_urls":["https://researchspace.auckland.ac.nz/docs/uoa-docs/rights.htm"],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2292/74747","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Hioka, Yusuke","Watson, Catherine I"]},{"key":"dc:creator","label":"Author","values":["Eng, Nicholas"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2026-02-24T19:08:34Z"]},{"key":"dc:date.issued","label":"Date","values":["2025"]},{"key":"dc:publisher","label":"Institution","values":["ResearchSpace@Auckland"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Mechatronics Engineering"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Doctoral"]},{"key":"thesis:degree_name","label":"Degree Name","values":["PhD"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["The University of Auckland"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Speech Enhancement","Speech Synthesis","Text-To-Speech","Audio Quality","Speech Perception","Speech Quality Assessment","Objective Metrics","Deep Learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["Items in ResearchSpace are protected by copyright, with all rights reserved, unless otherwise indicated."]},{"key":"dc:rights.uri","label":"Rights URI","values":["https://researchspace.auckland.ac.nz/docs/uoa-docs/rights.htm"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/2292/74747"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Both speech synthesis and speech enhancement are well researched fields, but their interaction remains under-explored. In particular, the effectiveness of using enhanced speech to train a speech synthesis model is still relatively unknown. This thesis investigates the effects of using enhanced speech to train a speech synthesis model on the overall quality of the synthesised speech, and in doing so, gain a better understanding of the interactions between different speech enhancement algorithms, speech synthesis architectures, gender of the voice and the noise types that degrade the speech used to train the speech synthesis model. A series of large-scale perception tests were conducted whereby over 100 participants evaluated the quality of utterances generated by speech synthesis model trained on noisy or enhanced speech. For this study, 160 synthetic voices were created, one for every combination of four speech enhancement methods, four noise types at two SNRs, two voices, one male and one female for two speech synthesis models – MaryTTS and Tacotron 2. Using linear mixed effect models, the study discerned the interactions between speech enhancement algorithms, speech synthesis architectures, gender of the voice and noise types. The results from this analysis suggest that different genders of voices resulted in similar trends depending on the speech enhancement methods, speech synthesis methods and noise types. However, the speech synthesis methods were affected differently by the various speech enhancement methods and noise types. An investigation into evaluating the quality of synthesised voices created from noisy or enhanced speech using eleven objective quality metrics was also conducted. This investigation found that several common objective quality metrics used to evaluate speech enhancement also have relatively high correlations with subjective quality of utterances generated by speech synthesisers trained on noisy or enhanced speech."]},{"key":"dc:title","label":"Title","values":["Assessing the Quality of Synthetic Speech when using Enhanced Speech as Training Data"]}]}],"canonical_facts":{"dc:contributor.advisor":["Hioka, Yusuke","Watson, Catherine I"],"dc:creator":["Eng, Nicholas"],"dc:date.accessioned":["2026-02-24T19:08:34Z"],"dc:date.issued":["2025"],"dc:description.abstract":["Both speech synthesis and speech enhancement are well researched fields, but their interaction remains under-explored. In particular, the effectiveness of using enhanced speech to train a speech synthesis model is still relatively unknown. This thesis investigates the effects of using enhanced speech to train a speech synthesis model on the overall quality of the synthesised speech, and in doing so, gain a better understanding of the interactions between different speech enhancement algorithms, speech synthesis architectures, gender of the voice and the noise types that degrade the speech used to train the speech synthesis model. A series of large-scale perception tests were conducted whereby over 100 participants evaluated the quality of utterances generated by speech synthesis model trained on noisy or enhanced speech. For this study, 160 synthetic voices were created, one for every combination of four speech enhancement methods, four noise types at two SNRs, two voices, one male and one female for two speech synthesis models – MaryTTS and Tacotron 2. Using linear mixed effect models, the study discerned the interactions between speech enhancement algorithms, speech synthesis architectures, gender of the voice and noise types. The results from this analysis suggest that different genders of voices resulted in similar trends depending on the speech enhancement methods, speech synthesis methods and noise types. However, the speech synthesis methods were affected differently by the various speech enhancement methods and noise types. An investigation into evaluating the quality of synthesised voices created from noisy or enhanced speech using eleven objective quality metrics was also conducted. This investigation found that several common objective quality metrics used to evaluate speech enhancement also have relatively high correlations with subjective quality of utterances generated by speech synthesisers trained on noisy or enhanced speech."],"dc:identifier.uri":["https://hdl.handle.net/2292/74747"],"dc:publisher":["ResearchSpace@Auckland"],"dc:rights":["Items in ResearchSpace are protected by copyright, with all rights reserved, unless otherwise indicated."],"dc:rights.uri":["https://researchspace.auckland.ac.nz/docs/uoa-docs/rights.htm"],"dc:subject":["Speech Enhancement","Speech Synthesis","Text-To-Speech","Audio Quality","Speech Perception","Speech Quality Assessment","Objective Metrics","Deep Learning"],"dc:title":["Assessing the Quality of Synthetic Speech when using Enhanced Speech as Training Data"],"dc:type":["Thesis"],"thesis:degree_discipline":["Mechatronics Engineering"],"thesis:degree_level":["Doctoral"],"thesis:degree_name":["PhD"],"thesis:institution_name":["The University of Auckland"]},"updated_at":"2026-07-24T01:06:29Z"}