Back to results

University of East Anglia

Discovering Dynamic Visemes

Abstract

dc:description.abstract

Abstract This thesis introduces a set of new, dynamic units of visual speech which are learnt using computer vision and machine learning techniques. Rather than clustering phoneme labels as is done traditionally, the visible articulators of a speaker are tracked and automatically segmented into short, visually intuitive speech gestures based on the dynamics of the articulators. The segmented gestures are clustered into dynamic visemes, such that movements relating to the same visual function appear within the same cluster. Speech animation can then be generated on any facial model by mapping a phoneme sequence to a sequence of dynamic visemes, and stitching together an example of each viseme in the sequence. Dynamic visemes model coarticulation and maintain the dynamics of the original speech, so simple blending at the concatenation boundaries ensures a smooth transition. The efficacy of dynamic visemes for computer animation is formally evaluated both objectively and subjectively, and compared with traditional phoneme to static lip-pose interpolation.

Degree

thesis:*
Name dc:type.qualificationname
phd
Level dc:type.qualificationlevel
doctoral
Grantor dc:publisher.institution
University of East Anglia
Year dc:date.issued
2013

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Taylor, Sarah

Rights

Language dc:language
en

Chain of custody

source
Harvested from
University of East Anglia
Base URL
ueaeprints.uea.ac.uk/cgi/oai2
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
related terms
citation

Taylor, Sarah. Discovering Dynamic Visemes. doctoral thesis, University of East Anglia, 2013.