{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/127369"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/127369","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Audiovisual processing for generation and enhancement","abstract":"Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2026-12-01","abstract_html":"Submission published under a 24 month embargo labeled &#x27;U of I Access&#x27;, the embargo will last until 2026-12-01","abstract_has_math":false,"creators":["Fan, Xulin"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Hasegawa-Johnson, Mark Allan"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-11-27","date_published":"2024-11-27","updated_at":"2026-07-22T22:25:03Z","subjects":["Multimodal Signal Processing","Speech Processing","Speech Enhancement","Signal Processing"],"languages":["eng","en"],"rights":["Copyright 2024 Xulin Fan"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/127369","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Hasegawa-Johnson, Mark Allan"]},{"key":"dc:creator","label":"Author","values":["Fan, Xulin"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2024-11-27","2024-12"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Multimodal Signal Processing","Speech Processing","Speech Enhancement","Signal Processing"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng","en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2024 Xulin Fan"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/127369"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2026-12-01","The student, Xulin Fan, accepted the attached license on 2024-11-26 at 14:26.","The student, Xulin Fan, submitted this Thesis for approval on 2024-11-26 at 14:39.","This Thesis was approved for publication on 2024-11-27 at 09:26.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21403 on 2025-03-28 at 14:43:27","With advances in deep neural networks and the increasing availability of computational power, machine learning researchers working on unimodal tasks, such as text, audio, or vision, have begun to explore methods that address multimodal problems, which involve inputs from multiple modalities. Each modality carries distinct types of information represented in different formats. For example, audiovisual tasks typically involve an audio-visual aligned video, where the video channel is essentially a sequence of images represented as a 4D tensor (number of frames, RGB channels, height, width), while the audio channel is a 1D signal sampled at a much higher rate. Popular multimodal neural architectures generally consist of three stages: modality-specific encoders, modality fusion, and one or more task-specific decoders. To handle inputs of varying formats, a common design choice is to employ modality-specific encoders, which project each modality into a learnable embedding space. This embedding space is structured to facilitate cross-modal similarity and ease the subsequent fusion process. The modality fusion step, however, is tailored to the requirements of the specific downstream task. For instance, an audiovisual task that outputs an audio signal, such as audiovisual speech enhancement, may require high temporal resolution, whereas tasks that produce textual outputs, such as audiovisual automatic speech recognition, may prioritize different fusion strategies. In this thesis, we explore various design choices for two audiovisual tasks: audio-driven talking head synthesis and audiovisual target speaker extraction. Through extensive experimentation, we identify key considerations for adapting transformer-based and diffusion-based methods to multimodal scenarios."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Audiovisual processing for generation and enhancement"]}]}],"canonical_facts":{"dc:contributor":["Hasegawa-Johnson, Mark Allan"],"dc:creator":["Fan, Xulin"],"dc:date":["2024-11-27","2024-12"],"dc:description":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2026-12-01","The student, Xulin Fan, accepted the attached license on 2024-11-26 at 14:26.","The student, Xulin Fan, submitted this Thesis for approval on 2024-11-26 at 14:39.","This Thesis was approved for publication on 2024-11-27 at 09:26.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21403 on 2025-03-28 at 14:43:27","With advances in deep neural networks and the increasing availability of computational power, machine learning researchers working on unimodal tasks, such as text, audio, or vision, have begun to explore methods that address multimodal problems, which involve inputs from multiple modalities. Each modality carries distinct types of information represented in different formats. For example, audiovisual tasks typically involve an audio-visual aligned video, where the video channel is essentially a sequence of images represented as a 4D tensor (number of frames, RGB channels, height, width), while the audio channel is a 1D signal sampled at a much higher rate. Popular multimodal neural architectures generally consist of three stages: modality-specific encoders, modality fusion, and one or more task-specific decoders. To handle inputs of varying formats, a common design choice is to employ modality-specific encoders, which project each modality into a learnable embedding space. This embedding space is structured to facilitate cross-modal similarity and ease the subsequent fusion process. The modality fusion step, however, is tailored to the requirements of the specific downstream task. For instance, an audiovisual task that outputs an audio signal, such as audiovisual speech enhancement, may require high temporal resolution, whereas tasks that produce textual outputs, such as audiovisual automatic speech recognition, may prioritize different fusion strategies. In this thesis, we explore various design choices for two audiovisual tasks: audio-driven talking head synthesis and audiovisual target speaker extraction. Through extensive experimentation, we identify key considerations for adapting transformer-based and diffusion-based methods to multimodal scenarios."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/127369"],"dc:language":["eng","en"],"dc:rights":["Copyright 2024 Xulin Fan"],"dc:subject":["Multimodal Signal Processing","Speech Processing","Speech Enhancement","Signal Processing"],"dc:title":["Audiovisual processing for generation and enhancement"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:03Z"}