Back to results

University of Illinois at Urbana-Champaign

Unsupervised speech technology for low-resource languages

Abstract

dc:description

Deep neural network based speech processing systems have found widespread applications in daily life, being employed for tasks such as automatic speech recognition (ASR), text-to-speech (TTS) synthesis, spoken language understanding (SLU), etc. With a sufficient amount of parallel speech-text training data, these systems attain performance levels comparable to, or in some cases, even better than human capabilities. However, such sufficient data assumption holds for only resource-rich languages such as English and Mandarin Chinese, and is unrealistic for many existing low-resource languages, posing a challenge for these systems to attain similar high performance. It is therefore meaningful to improve speech processing systems in such conditions to make speech technology accessible to a broader population. Unsupervised learning has been an active research field to mitigate data sparsity of low-resource languages. Depending on different source-target scenarios, unsupervised learning can be classified into four categories: (1) self-supervised learning (SSL), (2) modality matching, (3) unsupervised transfer learning, and (4) unsupervised multimodal learning. This thesis introduces six projects that leverage unsupervised learning methods to improve speech processing systems. The first project pretrains the SSL models on monolingual, cross-lingual, and multimodal data to study the cross-lingual transferability of SSL models. The second project improves the SSL representations using synthetic speech generated by a diffusion-based unit-to-speech synthesizer. The third project falls under modality matching, where we build the first unsupervised speech-to-text system using unsupervised automatic speech recognition technology. The fourth project falls under unsupervised transfer learning, where we improve zero-shot phonetic recognition system using language embeddings derived from external linguistic databases, without requiring any training data from the target languages. The fifth project also falls under transfer learning, where we build a multimodal few-shot SLU system by prompting a frozen pretrained language model with text and acoustic embeddings. The sixth project falls under unsupervised transfer learning, where we improve the current grapheme-to-phoneme (G2P) transducer by integrating the grapheme-to-phoneme model with a unit-to-phoneme (U2P) model, aiming to regularize G2P model outputs without relying on ground truth phoneme transcripts as training labels. This thesis demonstrates that unsupervised learning methods can significantly improve the performance of speech recognition, speech synthesis, and speech understanding in low-resourced application scenarios.

Degree

thesis:*
Name thesis:degree_name
Ph.D.
Level thesis:degree_level
Dissertation
Discipline thesis:degree_discipline
Electrical & Computer Engr
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2024

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Gao, Heting
Contributors dc:contributor
  • Hasegawa-Johnson, Mark
  • Smaragdis, Paris
  • Bhat, Suma P
  • Tang, Yan

Subjects

dc:subject × 7

Rights

dc:rights
Statement dc:rights
  • Copyright 2024 Heting Gao
Language dc:language
en, eng

Identifiers

dc:identifier.*
Handle dc:identifier
https://hdl.handle.net/2142/124233

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Gao, Heting. Unsupervised speech technology for low-resource languages. Dissertation thesis, University of Illinois at Urbana-Champaign, 2024. https://hdl.handle.net/2142/124233