University of Illinois at Urbana-Champaign
Multimodal spoken unit discovery with paired and unpaired modalities
Abstract
dc:descriptionThis thesis addresses the challenge of low-resource speech recognition by formulating it as a multimodal learning problem. The goal is to build a multimodal spoken unit discovery system that does not require any textual transcripts. Instead, it leverages speech and semantically related, multimodal signals such as paired images, unpaired text and unpaired sign language videos. To this end, this thesis proposes several novel algorithms based on neural networks and probabilistic graphical models. Further, it provides theoretical insights and empirical evidence to validate the efficacy of multimodal signals for spoken unit discovery.
Degree
thesis:*- Name thesis:degree_name
- Ph.D.
- Level thesis:degree_level
- Dissertation
- Discipline thesis:degree_discipline
- Electrical & Computer Engr
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2023
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Wang, Liming
- Contributors dc:contributor
-
- Hasegawa-Johnson, Mark
- Smaragdis, Paris
- Schwing, Alexander
- Fleck, Margaret
Subjects
dc:subject × 6Rights
dc:rights- Statement dc:rights
-
- Copyright 2023 by Liming Wang
- Language dc:language
- en, eng
Identifiers
dc:identifier.*- Handle dc:identifier
- https://hdl.handle.net/2142/121497