{"id":{"repo_id":"cuny-grad","oai_identifier":"oai:academicworks.cuny.edu:gc_etds-7310"},"canonical_url":"https://search.dev.ndltd.org/etd/cuny-grad/oai:academicworks.cuny.edu:gc_etds-7310","repository":{"repo_id":"cuny-grad","name":"City University of New York - Graduate Center","base_url":"https://academicworks.cuny.edu/do/oai/"},"display":{"title":"Toward More Intelligible Simultaneous Multi-channel Speech Enhancement and Recognition","abstract":"<p>Traditional single-channel speech enhancement and separation methods focus on enhancing the target speech signal by suppressing the noise and interfering speech signal. The methods suffer from nonlinear distortion brought by the algorithm, which hurts the intelligibility of the speech and also downstream tasks such as automatic speech recognition (ASR). We propose a method that leverages multi-channel input that robustly reduces the nonlinear speech distortion. We first demonstrate a better time-frequency mask estimation can help improve the mask based MVDR beamforming algorithm. Then we propose a novel mask-dependent training criterion to improve the phase estimation for speech separation. Additionally, we propose an end-to-end multi-channel neural network (WPD++) that can simultaneously separate and dereverberate the multi-channel noisy speech mixture. Finally, we show that integrating self-supervised learning models into the multi-channel speech enhancement and dereverberation network further reduces the word error rate (WER) metric for the downstream ASR task.</p>","abstract_html":"&lt;p&gt;Traditional single-channel speech enhancement and separation methods focus on enhancing the target speech signal by suppressing the noise and interfering speech signal. The methods suffer from nonlinear distortion brought by the algorithm, which hurts the intelligibility of the speech and also downstream tasks such as automatic speech recognition (ASR). We propose a method that leverages multi-channel input that robustly reduces the nonlinear speech distortion. We first demonstrate a better time-frequency mask estimation can help improve the mask based MVDR beamforming algorithm. Then we propose a novel mask-dependent training criterion to improve the phase estimation for speech separation. Additionally, we propose an end-to-end multi-channel neural network (WPD++) that can simultaneously separate and dereverberate the multi-channel noisy speech mixture. Finally, we show that integrating self-supervised learning models into the multi-channel speech enhancement and dereverberation network further reduces the word error rate (WER) metric for the downstream ASR task.&lt;/p&gt;","abstract_has_math":false,"creators":["Ni, Zhaoheng"],"institution":"The Graduate School and University Center of The City University of New York","degree_name":"Doctor of Philosophy","degree_level":"Doctoral","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":[],"advisors":["Michael I Mandel"],"committee_chairs":[],"committee_members":["Rivka Levitan","Lei Xie","Keelan Evanini"],"year":2025,"date_issued":"2025-02-01T08:00:00Z","date_published":"2025-02-01T08:00:00Z","updated_at":"2026-07-24T01:58:57Z","subjects":["Computer Engineering","Signal Processing","speech enhancement","speech dereverberation","multi-channel","beamforming"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://academicworks.cuny.edu/gc_etds/6163","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Michael I Mandel"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Rivka Levitan","Lei Xie","Keelan Evanini"]},{"key":"dc:creator","label":"Author","values":["Ni, Zhaoheng"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2027-02-01T08:00:00Z"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Doctoral"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctor of Philosophy"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["The Graduate School and University Center of The City University of New York"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer Engineering","Signal Processing","speech enhancement","speech dereverberation","multi-channel","beamforming"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://academicworks.cuny.edu/gc_etds/6163"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>Traditional single-channel speech enhancement and separation methods focus on enhancing the target speech signal by suppressing the noise and interfering speech signal. The methods suffer from nonlinear distortion brought by the algorithm, which hurts the intelligibility of the speech and also downstream tasks such as automatic speech recognition (ASR). We propose a method that leverages multi-channel input that robustly reduces the nonlinear speech distortion. We first demonstrate a better time-frequency mask estimation can help improve the mask based MVDR beamforming algorithm. Then we propose a novel mask-dependent training criterion to improve the phase estimation for speech separation. Additionally, we propose an end-to-end multi-channel neural network (WPD++) that can simultaneously separate and dereverberate the multi-channel noisy speech mixture. Finally, we show that integrating self-supervised learning models into the multi-channel speech enhancement and dereverberation network further reduces the word error rate (WER) metric for the downstream ASR task.</p>"]},{"key":"dc:title","label":"Title","values":["Toward More Intelligible Simultaneous Multi-channel Speech Enhancement and Recognition"]}]}],"canonical_facts":{"dc:contributor.advisor":["Michael I Mandel"],"dc:contributor.committeemember":["Rivka Levitan","Lei Xie","Keelan Evanini"],"dc:creator":["Ni, Zhaoheng"],"dc:date.available":["2027-02-01T08:00:00Z"],"dc:description.abstract":["<p>Traditional single-channel speech enhancement and separation methods focus on enhancing the target speech signal by suppressing the noise and interfering speech signal. The methods suffer from nonlinear distortion brought by the algorithm, which hurts the intelligibility of the speech and also downstream tasks such as automatic speech recognition (ASR). We propose a method that leverages multi-channel input that robustly reduces the nonlinear speech distortion. We first demonstrate a better time-frequency mask estimation can help improve the mask based MVDR beamforming algorithm. Then we propose a novel mask-dependent training criterion to improve the phase estimation for speech separation. Additionally, we propose an end-to-end multi-channel neural network (WPD++) that can simultaneously separate and dereverberate the multi-channel noisy speech mixture. Finally, we show that integrating self-supervised learning models into the multi-channel speech enhancement and dereverberation network further reduces the word error rate (WER) metric for the downstream ASR task.</p>"],"dc:identifier":["https://academicworks.cuny.edu/gc_etds/6163"],"dc:subject":["Computer Engineering","Signal Processing","speech enhancement","speech dereverberation","multi-channel","beamforming"],"dc:title":["Toward More Intelligible Simultaneous Multi-channel Speech Enhancement and Recognition"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Doctoral"],"thesis:degree_name":["Doctor of Philosophy"],"thesis:institution_name":["The Graduate School and University Center of The City University of New York"]},"updated_at":"2026-07-24T01:58:57Z"}