Back to results

University of Illinois at Urbana-Champaign

Audiovisual processing for generation and enhancement

Abstract

dc:description

With advances in deep neural networks and the increasing availability of computational power, machine learning researchers working on unimodal tasks, such as text, audio, or vision, have begun to explore methods that address multimodal problems, which involve inputs from multiple modalities. Each modality carries distinct types of information represented in different formats. For example, audiovisual tasks typically involve an audio-visual aligned video, where the video channel is essentially a sequence of images represented as a 4D tensor (number of frames, RGB channels, height, width), while the audio channel is a 1D signal sampled at a much higher rate. Popular multimodal neural architectures generally consist of three stages: modality-specific encoders, modality fusion, and one or more task-specific decoders. To handle inputs of varying formats, a common design choice is to employ modality-specific encoders, which project each modality into a learnable embedding space. This embedding space is structured to facilitate cross-modal similarity and ease the subsequent fusion process. The modality fusion step, however, is tailored to the requirements of the specific downstream task. For instance, an audiovisual task that outputs an audio signal, such as audiovisual speech enhancement, may require high temporal resolution, whereas tasks that produce textual outputs, such as audiovisual automatic speech recognition, may prioritize different fusion strategies. In this thesis, we explore various design choices for two audiovisual tasks: audio-driven talking head synthesis and audiovisual target speaker extraction. Through extensive experimentation, we identify key considerations for adapting transformer-based and diffusion-based methods to multimodal scenarios.

Degree

thesis:*
Name thesis:degree_name
M.S.
Level thesis:degree_level
Thesis
Discipline thesis:degree_discipline
Electrical & Computer Engr
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2024

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Fan, Xulin
Contributors dc:contributor
  • Hasegawa-Johnson, Mark Allan

Subjects

dc:subject × 4

Rights

dc:rights
Statement dc:rights
  • Copyright 2024 Xulin Fan
Language dc:language
eng, en

Identifiers

dc:identifier.*
Handle dc:identifier
https://hdl.handle.net/2142/127369

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Fan, Xulin. Audiovisual processing for generation and enhancement. Thesis thesis, University of Illinois at Urbana-Champaign, 2024. https://hdl.handle.net/2142/127369