Back to search

Cal Poly

Visual Speech Recognition Using a 3D Convolutional Neural Network

Abstract

dc:description.abstract

<p>Main stream automatic speech recognition (ASR) makes use of audio data to identify spoken words, however visual speech recognition (VSR) has recently been of increased interest to researchers. VSR is used when audio data is corrupted or missing entirely and also to further enhance the accuracy of audio-based ASR systems. In this research, we present both a framework for building 3D feature cubes of lip data from videos and a 3D convolutional neural network (CNN) architecture for performing classification on a dataset of 100 spoken words, recorded in an uncontrolled envi- ronment. Our 3D-CNN architecture achieves a testing accuracy of 64%, comparable with recent works, but using an input data size that is up to 75% smaller. Overall, our research shows that 3D-CNNs can be successful in finding spatial-temporal features using unsupervised feature extraction and are a suitable choice for VSR-based systems.</p>

Degree

thesis:*
Name thesis:degree_name
MS in Electrical Engineering
Discipline thesis:degree_discipline
Electrical Engineering
Year dc:date.available
2019

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Rochford, Matthew
Contributors dc:contributor
  • Jane Zhang
  • Electrical Engineering
  • College of Engineering

Subjects

dc:subject × 7

Identifiers

dc:identifier.*
OAI identifier oai:identifier
oai:digitalcommons.calpoly.edu:theses-3577

Chain of custody

source
Harvested from
Cal Poly
Base URL
digitalcommons.calpoly.edu/do/oai/
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Rochford, Matthew. Visual Speech Recognition Using a 3D Convolutional Neural Network. 2019. https://digitalcommons.calpoly.edu/theses/2109