University of Illinois at Urbana-Champaign
SpeechSplit2: An efficient unsupervised speech disentanglement model for multi-aspect voice conversion
Abstract
dc:descriptionSpeechSplit is among the first algorithms that successfully disentangle speech into four components: rhythm, content, pitch, and timbre. However, the model requires exhaustive tuning of the encoder bottlenecks, which can be a daunting task and limits its generalization ability. In this work, we present SpeechSplit2, an improved version of SpeechSplit, in which simple signal processing methods are utilized to alleviate the laborious bottleneck tuning problem. We show that by feeding different inputs to each encoder, we can guide each encoder to only extract one particular aspect of speech and discard the rest, given the bottleneck size is sufficiently large to encode the corresponding information. With the same neural network architecture as SpeechSplit, SpeechSplit2 achieves comparable performance in disentangling speech components when the bottlenecks are carefully tuned and shows superior advantage over the baseline when the bottleneck size varies.
Degree
thesis:*- Name thesis:degree_name
- M.S.
- Level thesis:degree_level
- Thesis
- Discipline thesis:degree_discipline
- Electrical & Computer Engr
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2022
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Chan, Chak Ho
- Contributors dc:contributor
-
- Hasegawa-Johnson, Mark Allan
Subjects
dc:subject × 4Rights
dc:rights- Statement dc:rights
-
- Copyright 2022 Chak Ho Chan
- Language dc:language
- en, eng
Identifiers
dc:identifier.*- Handle dc:identifier
- https://hdl.handle.net/2142/117680