Back to results

University of Illinois at Urbana-Champaign

Multimodal Fusion With Applications to Audio -Visual Speech Recognition

Abstract

dc:description

Differences in the characteristics of the intermodal couplings in audio-visual speech recognition and in multichannel biometrics defy a universal fusion method for both applications. For audio-visual speech modeling, we propose a novel sensory fusion method based on the coupled hidden Markov models (CHMMs). The CHMM framework allows the fusion of two temporally coupled information sources to take place as an integral part of the statistical modeling process. An important advantage of the CHMM-based fusion method lies in its ability to model asynchronies between the audio and visual channels. We describe two approaches to carry out inference and learning in CHMMs. The first is an exact algorithm derived by extending the forward-backward procedure used in hidden Markov model (HMM) inference. The second method relies on the model transformation strategy that maps the state space of a CHMM onto the state space of a classic HMM, and therefore facilitates the development of sophisticated audio-visual speech recognition systems using existing infrastructures. For multichannel biometrics, we introduce a general formulation based on the late integration paradigm and address the environmental robustness issue through multichannel fusion. Based on this formulation, two effective approaches to carry out environment-adaptive decision fusion are developed: the environmental confidence weighting method and the optimal channel weighting method.

Degree

thesis:*
Name thesis:degree_name
Ph.D.
Level thesis:degree_level
Dissertation
Discipline thesis:degree_discipline
Electrical Engineering
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2015

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Chu, Stephen Mingyu
Contributors dc:contributor
  • Huang, Thomas S.

Subjects

dc:subject × 1

Rights

Language dc:language
eng

Identifiers

dc:identifier.*
Identifier
(MiAaPQ)AAI3086036
OAI identifier oai:identifier
oai:www.ideals.illinois.edu:2142/80819

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Chu, Stephen Mingyu. Multimodal Fusion With Applications to Audio -Visual Speech Recognition. Dissertation thesis, University of Illinois at Urbana-Champaign, 2015. http://hdl.handle.net/2142/80819