The Graduate School and University Center of The City University of New York
Toward More Intelligible Simultaneous Multi-channel Speech Enhancement and Recognition
Abstract
dc:description.abstract<p>Traditional single-channel speech enhancement and separation methods focus on enhancing the target speech signal by suppressing the noise and interfering speech signal. The methods suffer from nonlinear distortion brought by the algorithm, which hurts the intelligibility of the speech and also downstream tasks such as automatic speech recognition (ASR). We propose a method that leverages multi-channel input that robustly reduces the nonlinear speech distortion. We first demonstrate a better time-frequency mask estimation can help improve the mask based MVDR beamforming algorithm. Then we propose a novel mask-dependent training criterion to improve the phase estimation for speech separation. Additionally, we propose an end-to-end multi-channel neural network (WPD++) that can simultaneously separate and dereverberate the multi-channel noisy speech mixture. Finally, we show that integrating self-supervised learning models into the multi-channel speech enhancement and dereverberation network further reduces the word error rate (WER) metric for the downstream ASR task.</p>
Degree
thesis:*- Name thesis:degree_name
- Doctor of Philosophy
- Level thesis:degree_level
- Doctoral
- Discipline thesis:degree_discipline
- Computer Science
- Grantor
- The Graduate School and University Center of The City University of New York
- Year dc:date.available
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Ni, Zhaoheng
- Advisor dc:contributor.advisor
-
- Michael I Mandel
- Committee members dc:contributor.committeemember
-
- Rivka Levitan
- Lei Xie
- Keelan Evanini
Subjects
dc:subject × 6Identifiers
dc:identifier.*- Repository record dc:identifier
- https://academicworks.cuny.edu/gc_etds/6163
- OAI identifier oai:identifier
- oai:academicworks.cuny.edu:gc_etds-7310