Back to results

The Graduate School and University Center of The City University of New York

Data-Centric Machine Learning for Speech and Audio

Abstract

dc:description.abstract

<p>There is growing recognition of the importance of data-centric methods for building machine learning systems. Data-centric methods assume a fixed model and iterate over the data to improve system performance. This is in contrast to traditional model-centric approaches, which assume a fixed dataset and iterate over models for the same ends. Data-centric machine learning is driven by the observation that, beyond the size of the training data, model performance depends on factors such as the quality of the annotations, and whether the data are representative of conditions in which models will be deployed. This is particularly of interest in the domains of speech and audio, where it is relatively cheap to acquire large amounts of data, but highly expensive and laborious to annotate or assess the quality of the training data. In this work, we investigate and develop methods for identifying highly informative subsets of speech or audio for improving system performance while reducing the amount of data required for training.</p> <p>First, we investigate submodular data selection in the context of unsupervised active learning for automatic speech recognition (ASR) of low-resource languages. We demonstrate methods for sampling the acoustic feature space and selecting highly informative and diverse examples to build models with highly limited data. Second, we present data selection methods that apply domain knowledge for a speech enhancement system. We demonstrate that linguistic and acoustic characteristics of speech data can be used to select examples that enable a deep neural network (DNN) to learn a more generalizable similarity metric. Last, we investigate data valuation for curation of training data by assessing data quality. In the context of speech recognition, we develop a method for estimating Shapley values of speech data for training an end-to-end neural ASR model, which performs a structured prediction task. We also demonstrate how data valuation can be used for environmental sound classification. In the context of ecoacoustic monitoring of large data with limited labels, we estimate Shapley values of audio clips with overlapping sounds for a multi-label classifier. We demonstrate that these values identify high quality subsets of the training data for improving model performance. We also demonstrate methods for assessing data quality within individual sound classes and identifying annotation errors.</p>

Degree

thesis:*
Name thesis:degree_name
Doctor of Philosophy
Level thesis:degree_level
Doctoral
Discipline thesis:degree_discipline
Computer Science
Grantor
The Graduate School and University Center of The City University of New York
Year dc:date.available
2022

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Syed, Ali Raza
Advisor dc:contributor.advisor
  • Michael I. Mandel
Committee members dc:contributor.committeemember
  • Rebecca Levitan
  • Alla Rozovskaya
  • Brian Kingsbury

Subjects

dc:subject × 8

Identifiers

dc:identifier.*
Repository record dc:identifier
https://academicworks.cuny.edu/gc_etds/5059
OAI identifier oai:identifier
oai:academicworks.cuny.edu:gc_etds-6215

Chain of custody

source
Harvested from
City University of New York - Graduate Center
Base URL
academicworks.cuny.edu/do/oai/
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Syed, Ali Raza. Data-Centric Machine Learning for Speech and Audio. Doctoral thesis, The Graduate School and University Center of The City University of New York, 2022. https://academicworks.cuny.edu/gc_etds/5059