Back to results

The Graduate School and University Center of The City University of New York

Toward More Intelligible Simultaneous Multi-channel Speech Enhancement and Recognition

Abstract

dc:description.abstract

<p>Traditional single-channel speech enhancement and separation methods focus on enhancing the target speech signal by suppressing the noise and interfering speech signal. The methods suffer from nonlinear distortion brought by the algorithm, which hurts the intelligibility of the speech and also downstream tasks such as automatic speech recognition (ASR). We propose a method that leverages multi-channel input that robustly reduces the nonlinear speech distortion. We first demonstrate a better time-frequency mask estimation can help improve the mask based MVDR beamforming algorithm. Then we propose a novel mask-dependent training criterion to improve the phase estimation for speech separation. Additionally, we propose an end-to-end multi-channel neural network (WPD++) that can simultaneously separate and dereverberate the multi-channel noisy speech mixture. Finally, we show that integrating self-supervised learning models into the multi-channel speech enhancement and dereverberation network further reduces the word error rate (WER) metric for the downstream ASR task.</p>

Degree

thesis:*
Name thesis:degree_name
Doctor of Philosophy
Level thesis:degree_level
Doctoral
Discipline thesis:degree_discipline
Computer Science
Grantor
The Graduate School and University Center of The City University of New York
Year dc:date.available
2025

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Ni, Zhaoheng
Advisor dc:contributor.advisor
  • Michael I Mandel
Committee members dc:contributor.committeemember
  • Rivka Levitan
  • Lei Xie
  • Keelan Evanini

Subjects

dc:subject × 6

Identifiers

dc:identifier.*
Repository record dc:identifier
https://academicworks.cuny.edu/gc_etds/6163
OAI identifier oai:identifier
oai:academicworks.cuny.edu:gc_etds-7310

Chain of custody

source
Harvested from
City University of New York - Graduate Center
Base URL
academicworks.cuny.edu/do/oai/
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Ni, Zhaoheng. Toward More Intelligible Simultaneous Multi-channel Speech Enhancement and Recognition. Doctoral thesis, The Graduate School and University Center of The City University of New York, 2025. https://academicworks.cuny.edu/gc_etds/6163