ResearchSpace@Auckland
Assessing the Quality of Synthetic Speech when using Enhanced Speech as Training Data
Abstract
dc:description.abstractBoth speech synthesis and speech enhancement are well researched fields, but their interaction remains under-explored. In particular, the effectiveness of using enhanced speech to train a speech synthesis model is still relatively unknown. This thesis investigates the effects of using enhanced speech to train a speech synthesis model on the overall quality of the synthesised speech, and in doing so, gain a better understanding of the interactions between different speech enhancement algorithms, speech synthesis architectures, gender of the voice and the noise types that degrade the speech used to train the speech synthesis model. A series of large-scale perception tests were conducted whereby over 100 participants evaluated the quality of utterances generated by speech synthesis model trained on noisy or enhanced speech. For this study, 160 synthetic voices were created, one for every combination of four speech enhancement methods, four noise types at two SNRs, two voices, one male and one female for two speech synthesis models – MaryTTS and Tacotron 2. Using linear mixed effect models, the study discerned the interactions between speech enhancement algorithms, speech synthesis architectures, gender of the voice and noise types. The results from this analysis suggest that different genders of voices resulted in similar trends depending on the speech enhancement methods, speech synthesis methods and noise types. However, the speech synthesis methods were affected differently by the various speech enhancement methods and noise types. An investigation into evaluating the quality of synthesised voices created from noisy or enhanced speech using eleven objective quality metrics was also conducted. This investigation found that several common objective quality metrics used to evaluate speech enhancement also have relatively high correlations with subjective quality of utterances generated by speech synthesisers trained on noisy or enhanced speech.
Degree
thesis:*- Name thesis:degree_name
- PhD
- Level thesis:degree_level
- Doctoral
- Discipline thesis:degree_discipline
- Mechatronics Engineering
- Grantor dc:publisher
- ResearchSpace@Auckland
- Year dc:date.issued
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Eng, Nicholas
- Advisors dc:contributor.advisor
-
- Hioka, Yusuke
- Watson, Catherine I
Subjects
dc:subject × 8Rights
dc:rights- Statement dc:rights
-
- Items in ResearchSpace are protected by copyright, with all rights reserved, unless otherwise indicated.
- Licence dc:rights.uri
Identifiers
dc:identifier.*- Handle dc:identifier.uri
- https://hdl.handle.net/2292/74747
- OAI identifier oai:identifier
- oai:researchspace.auckland.ac.nz:2292/74747