Back to search

University of Illinois at Urbana-Champaign

Multi-channel multi-modal speech enhancement and separation

Abstract

dc:description

This thesis explores the topic of multi-channel multi-modal speech enhancement and separation, addressing the challenges associated with diverse acoustic scenarios and different modalities. The research is organized into three interconnected projects/chapters, each contributing to the same goal of improving speech separation and enhancement in various real-world contexts. The first project introduces a novel approach to speech separation utilizing binaural microphone inputs. The system autonomously separates speeches into distinct pre-defined spatial regions, adapting dynamically to different Head-Related Transfer Functions (HRTFs) for individual users by self-supervised fine-tuning. The second project focuses on speech separation and enhancement using microphone arrays integrated into augmented reality (AR) glasses. The system leverages the spatial information provided by the array to effectively separate and enhance speech, accommodating directions of speech arrivals from visual inputs. This project aims to improve the user experience in scenarios where hands-free, unobtrusive speech processing is crucial, such as augmented reality environments. The third project explores audio-visual speech separation by incorporating video data capturing human facial and lip movements. This multi-modal model significantly enhances speech separation performance by integrating visual cues into the audio processing pipeline. The model not only takes the speech mixture as input but also leverages facial expressions and lip movements, resulting in improved accuracy and robustness, especially in challenging acoustic conditions. Collectively, these projects contribute to the advancement of multi-channel multi-modal speech processing, offering solutions for scenarios ranging from personalized spatial audio processing to hands-free augmented reality applications. The findings and methodologies presented in this thesis show a few challenges and opportunities in the field, progressing for better multi-channel multi-modal speech enhancement and separation systems.

Degree

thesis:*
Name thesis:degree_name
M.S.
Level thesis:degree_level
Thesis
Discipline thesis:degree_discipline
Electrical & Computer Engr
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2023

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Xu, Zhongweiyang
Contributors dc:contributor
  • Roy Choudhury, Romit

Subjects

dc:subject × 3

Rights

dc:rights
Statement dc:rights
  • Copyright 2023 Zhongweiyang Xu
Language dc:language
en, eng

Identifiers

dc:identifier.*
Handle dc:identifier
https://hdl.handle.net/2142/122068

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Xu, Zhongweiyang. Multi-channel multi-modal speech enhancement and separation. Thesis thesis, University of Illinois at Urbana-Champaign, 2023. https://hdl.handle.net/2142/122068