Back to results

University of Lethbridge

Temporal constraints on human and artificial muliti-sensory speech recognition

Abstract

Audio Visual Speech Recognition (AVSR) is the process of perceiving and understanding speech using audio and visual information. Combining visual information with auditory stimuli has been shown to improve AVSR performance when compared to purely auditory speech recognition when the task is performed in adverse conditions with large amounts of distracting noise. This work examines the relationship of auditory and visual speech information and the effect audio-visual temporary desynchronization has on AVSR performance. Using a whole report task, we show that (1) consistent with prior similar work, performance declines asymmetrically depending on the direction and quantity of a temporal lag, and (2) a common, modern architecture for computational AVSR does not show this asymmetry indicating a fundamental difference in biological and computational AVSR methods.

Author and committee

dc:creator, dc:contributor.*
Authors
  • Perlette, Christopher S.
  • University of Lethbridge. Faculty of Arts and Science

Subjects

dc:subject × 11

Identifiers

dc:identifier.*
Identifier
hdl:10133/6465
OAI identifier oai:identifier
oai:opus.uleth.ca:10133/6465

Chain of custody

source
Harvested from
University of Lethbridge
Base URL
opus.uleth.ca/server/oai/request
Last updated
2026-07-27
Source record
OAI-PMH GetRecord
citation

Perlette, Christopher S.; University of Lethbridge. Faculty of Arts and Science. Temporal constraints on human and artificial muliti-sensory speech recognition. 2023.