Back to search

University of Illinois at Urbana-Champaign

SpeechSplit2: An efficient unsupervised speech disentanglement model for multi-aspect voice conversion

Abstract

dc:description

SpeechSplit is among the first algorithms that successfully disentangle speech into four components: rhythm, content, pitch, and timbre. However, the model requires exhaustive tuning of the encoder bottlenecks, which can be a daunting task and limits its generalization ability. In this work, we present SpeechSplit2, an improved version of SpeechSplit, in which simple signal processing methods are utilized to alleviate the laborious bottleneck tuning problem. We show that by feeding different inputs to each encoder, we can guide each encoder to only extract one particular aspect of speech and discard the rest, given the bottleneck size is sufficiently large to encode the corresponding information. With the same neural network architecture as SpeechSplit, SpeechSplit2 achieves comparable performance in disentangling speech components when the bottlenecks are carefully tuned and shows superior advantage over the baseline when the bottleneck size varies.

Degree

thesis:*
Name thesis:degree_name
M.S.
Level thesis:degree_level
Thesis
Discipline thesis:degree_discipline
Electrical & Computer Engr
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2022

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Chan, Chak Ho
Contributors dc:contributor
  • Hasegawa-Johnson, Mark Allan

Subjects

dc:subject × 4

Rights

dc:rights
Statement dc:rights
  • Copyright 2022 Chak Ho Chan
Language dc:language
en, eng

Identifiers

dc:identifier.*
Handle dc:identifier
https://hdl.handle.net/2142/117680

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Chan, Chak Ho. SpeechSplit2: An efficient unsupervised speech disentanglement model for multi-aspect voice conversion. Thesis thesis, University of Illinois at Urbana-Champaign, 2022. https://hdl.handle.net/2142/117680