University of Illinois at Urbana-Champaign
Lipreading with convolutional and recurrent neural network models
Abstract
dc:descriptionLip reading is the process of speech recognition from solely visual information. The goal of this thesis is to perform a silence vs. speech classification, and to recognize the triphone spoken by a talking head, given only the video using neural network classification models. Two neural network architectures are developed and tested on the AVICAR dataset, including one convolutional neural network (CNN) model with fully connected classification layer, and one recurrent neural network (RNN) model with convolutional layer and one long short-term memory (LSTM) layer to perform the classification on a sequence of input. In both models, the convolutional layers serve as feature extractors. The performance of each model is experimentally evaluated and the detailed network structure and preprocessing pipeline are demonstrated.
Degree
thesis:*- Name thesis:degree_name
- M.S.
- Level thesis:degree_level
- Thesis
- Discipline thesis:degree_discipline
- Electrical & Computer Engr
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2017
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Zhu, Tianyilin
- Contributors dc:contributor
-
- Hasegawa-Johnson, Mark
Subjects
dc:subject × 2Rights
dc:rights- Statement dc:rights
-
- Copyright 2017 Tianyilin Zhu
- Language dc:language
- en
Identifiers
dc:identifier.*- Handle dc:identifier
- http://hdl.handle.net/2142/97763
- OAI identifier oai:identifier
- oai:www.ideals.illinois.edu:2142/97763