Back to results

The Graduate School and University Center of The City University of New York

Speech Enhancement Using Speech Synthesis Techniques

Abstract

dc:description.abstract

<p>Traditional speech enhancement systems reduce noise by modifying the noisy signal to make it more like a clean signal, which suffers from two problems: under-suppression of noise and over-suppression of speech. These problems create distortions in enhanced speech and hurt the quality of the enhanced signal. We propose to utilize speech synthesis techniques for a higher quality speech enhancement system. Synthesizing clean speech based on the noisy signal could produce outputs that are both noise-free and high quality. We first show that we can replace the noisy speech with its clean resynthesis from a previously recorded clean speech dictionary from the same speaker (concatenative resynthesis). Next, we show that using a speech synthesizer (vocoder) we can create a clean resynthesis of the noisy speech for more than one speaker. We term this parametric resynthesis (PR). PR can generate better prosody from noisy speech than a TTS system which uses textual information only. Additionally, we can use the high quality speech generation capability of neural vocoders for better quality speech enhancement. When trained on data from enough speakers, these vocoders can generate speech from unseen speakers, both male, and female, with similar quality as seen speakers in training. Finally, we show that using neural vocoders we can achieve better objective signal and overall quality than the state-of-the-art speech enhancement systems and better subjective quality than an oracle mask-based system.</p>

Degree

thesis:*
Name thesis:degree_name
Doctor of Philosophy
Level thesis:degree_level
Doctoral
Discipline thesis:degree_discipline
Computer Science
Grantor
The Graduate School and University Center of The City University of New York
Year dc:date.available
2021

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Maiti, Soumi
Advisor dc:contributor.advisor
  • Michael I. Mandel
Committee members dc:contributor.committeemember
  • Rivka Levitan
  • Kyle Gorman
  • Ron J Weiss

Subjects

dc:subject × 6

Identifiers

dc:identifier.*
Repository record dc:identifier
https://academicworks.cuny.edu/gc_etds/4202
OAI identifier oai:identifier
oai:academicworks.cuny.edu:gc_etds-5276

Chain of custody

source
Harvested from
City University of New York - Graduate Center
Base URL
academicworks.cuny.edu/do/oai/
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Maiti, Soumi. Speech Enhancement Using Speech Synthesis Techniques. Doctoral thesis, The Graduate School and University Center of The City University of New York, 2021. https://academicworks.cuny.edu/gc_etds/4202