Back to results

University of Illinois at Urbana-Champaign

Lipreading with convolutional and recurrent neural network models

Abstract

dc:description

Lip reading is the process of speech recognition from solely visual information. The goal of this thesis is to perform a silence vs. speech classification, and to recognize the triphone spoken by a talking head, given only the video using neural network classification models. Two neural network architectures are developed and tested on the AVICAR dataset, including one convolutional neural network (CNN) model with fully connected classification layer, and one recurrent neural network (RNN) model with convolutional layer and one long short-term memory (LSTM) layer to perform the classification on a sequence of input. In both models, the convolutional layers serve as feature extractors. The performance of each model is experimentally evaluated and the detailed network structure and preprocessing pipeline are demonstrated.

Degree

thesis:*
Name thesis:degree_name
M.S.
Level thesis:degree_level
Thesis
Discipline thesis:degree_discipline
Electrical & Computer Engr
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2017

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Zhu, Tianyilin
Contributors dc:contributor
  • Hasegawa-Johnson, Mark

Subjects

dc:subject × 2

Rights

dc:rights
Statement dc:rights
  • Copyright 2017 Tianyilin Zhu
Language dc:language
en

Identifiers

dc:identifier.*
Handle dc:identifier
http://hdl.handle.net/2142/97763
OAI identifier oai:identifier
oai:www.ideals.illinois.edu:2142/97763

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Zhu, Tianyilin. Lipreading with convolutional and recurrent neural network models. Thesis thesis, University of Illinois at Urbana-Champaign, 2017. http://hdl.handle.net/2142/97763