{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/122068"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/122068","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Multi-channel multi-modal speech enhancement and separation","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-03-01 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2024-03-01 without embargo terms","abstract_has_math":false,"creators":["Xu, Zhongweiyang"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Roy Choudhury, Romit"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2023,"date_issued":"2023-12","date_published":"2023-12","updated_at":"2026-07-22T22:25:00Z","subjects":["Speech Enhancement/separation","Multi-channel","Multi-modal"],"languages":["en","eng"],"rights":["Copyright 2023 Zhongweiyang Xu"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/122068","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Roy Choudhury, Romit"]},{"key":"dc:creator","label":"Author","values":["Xu, Zhongweiyang"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2023-12","2023-12-06"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Speech Enhancement/separation","Multi-channel","Multi-modal"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2023 Zhongweiyang Xu"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/122068"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-03-01 without embargo terms","The student, Zhongweiyang Xu, accepted the attached license on 2023-12-05 at 17:48.","The student, Zhongweiyang Xu, submitted this Thesis for approval on 2023-12-05 at 17:56.","This Thesis was approved for publication on 2023-12-06 at 11:25.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20149 on 2024-03-01 at 13:15:28","This thesis explores the topic of multi-channel multi-modal speech enhancement and separation, addressing the challenges associated with diverse acoustic scenarios and different modalities. The research is organized into three interconnected projects/chapters, each contributing to the same goal of improving speech separation and enhancement in various real-world contexts. The first project introduces a novel approach to speech separation utilizing binaural microphone inputs. The system autonomously separates speeches into distinct pre-defined spatial regions, adapting dynamically to different Head-Related Transfer Functions (HRTFs) for individual users by self-supervised fine-tuning. The second project focuses on speech separation and enhancement using microphone arrays integrated into augmented reality (AR) glasses. The system leverages the spatial information provided by the array to effectively separate and enhance speech, accommodating directions of speech arrivals from visual inputs. This project aims to improve the user experience in scenarios where hands-free, unobtrusive speech processing is crucial, such as augmented reality environments. The third project explores audio-visual speech separation by incorporating video data capturing human facial and lip movements. This multi-modal model significantly enhances speech separation performance by integrating visual cues into the audio processing pipeline. The model not only takes the speech mixture as input but also leverages facial expressions and lip movements, resulting in improved accuracy and robustness, especially in challenging acoustic conditions. Collectively, these projects contribute to the advancement of multi-channel multi-modal speech processing, offering solutions for scenarios ranging from personalized spatial audio processing to hands-free augmented reality applications. The findings and methodologies presented in this thesis show a few challenges and opportunities in the field, progressing for better multi-channel multi-modal speech enhancement and separation systems."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Multi-channel multi-modal speech enhancement and separation"]}]}],"canonical_facts":{"dc:contributor":["Roy Choudhury, Romit"],"dc:creator":["Xu, Zhongweiyang"],"dc:date":["2023-12","2023-12-06"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-03-01 without embargo terms","The student, Zhongweiyang Xu, accepted the attached license on 2023-12-05 at 17:48.","The student, Zhongweiyang Xu, submitted this Thesis for approval on 2023-12-05 at 17:56.","This Thesis was approved for publication on 2023-12-06 at 11:25.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20149 on 2024-03-01 at 13:15:28","This thesis explores the topic of multi-channel multi-modal speech enhancement and separation, addressing the challenges associated with diverse acoustic scenarios and different modalities. The research is organized into three interconnected projects/chapters, each contributing to the same goal of improving speech separation and enhancement in various real-world contexts. The first project introduces a novel approach to speech separation utilizing binaural microphone inputs. The system autonomously separates speeches into distinct pre-defined spatial regions, adapting dynamically to different Head-Related Transfer Functions (HRTFs) for individual users by self-supervised fine-tuning. The second project focuses on speech separation and enhancement using microphone arrays integrated into augmented reality (AR) glasses. The system leverages the spatial information provided by the array to effectively separate and enhance speech, accommodating directions of speech arrivals from visual inputs. This project aims to improve the user experience in scenarios where hands-free, unobtrusive speech processing is crucial, such as augmented reality environments. The third project explores audio-visual speech separation by incorporating video data capturing human facial and lip movements. This multi-modal model significantly enhances speech separation performance by integrating visual cues into the audio processing pipeline. The model not only takes the speech mixture as input but also leverages facial expressions and lip movements, resulting in improved accuracy and robustness, especially in challenging acoustic conditions. Collectively, these projects contribute to the advancement of multi-channel multi-modal speech processing, offering solutions for scenarios ranging from personalized spatial audio processing to hands-free augmented reality applications. The findings and methodologies presented in this thesis show a few challenges and opportunities in the field, progressing for better multi-channel multi-modal speech enhancement and separation systems."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/122068"],"dc:language":["en","eng"],"dc:rights":["Copyright 2023 Zhongweiyang Xu"],"dc:subject":["Speech Enhancement/separation","Multi-channel","Multi-modal"],"dc:title":["Multi-channel multi-modal speech enhancement and separation"],"dc:type":["text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:00Z"}