Back to results

Massachusetts Institute of Technology

Speech processing with less supervision : learning from weak labels and multiple modalities

Abstract

dc:description.abstract

In recent years, supervised learning has achieved great success in speech processing with powerful neural network models and vast quantities of in-domain labeled data. However, collecting a labeled dataset covering all domains can be either expensive due to the diversity of speech or almost impossible for some tasks such as speech-to-speech translation. Such a paradigm limits the applicability of speech technologies to high-resource settings. In sharp contrast, humans are good at reading the training signals from indirect supervision, such as from small amount of explicit labels and from different modalities. This capability enables humans to learn from a wider variety of resources, including better domain coverage. In light of this observation, this thesis focuses on learning algorithms for speech processing that can utilize weak and indirect supervision to overcome the restrictions imposed by the supervised paradigm and make the most out of the data at hand for learning.

Degree

thesis:*
Name thesis:degree_name
Doctoral
Department dc:contributor.department
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
Grantor dc:publisher
Massachusetts Institute of Technology
Year dc:date.issued
2020

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Hsu, Wei-Ning,Ph. D.Massachusetts Institute of Technology.
Advisor dc:contributor.advisor
  • James R. Glass.

Subjects

dc:subject × 1

Rights

dc:rights
Statement dc:rights
  • MIT theses may be protected by copyright. Please reuse MIT thesis content according to the MIT Libraries Permissions Policy, which is available through the URL provided.
Language dc:language.iso
eng

Identifiers

dc:identifier.*
Handle dc:identifier.uri
https://hdl.handle.net/1721.1/127021
OAI identifier oai:identifier
oai:dspace.mit.edu:1721.1/127021

Chain of custody

source
Harvested from
MIT
Base URL
dspace.mit.edu/oai/request
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Hsu, Wei-Ning,Ph. D.Massachusetts Institute of Technology.. Speech processing with less supervision : learning from weak labels and multiple modalities. Massachusetts Institute of Technology, 2020. https://hdl.handle.net/1721.1/127021