{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/97763"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/97763","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Lipreading with convolutional and recurrent neural network models","abstract":"Lip reading is the process of speech recognition from solely visual information. The goal of this thesis is to perform a silence vs. speech classification, and to recognize the triphone spoken by a talking head, given only the video using neural network classification models. Two neural network architectures are developed and tested on the AVICAR dataset, including one convolutional neural network (CNN) model with fully connected classification layer, and one recurrent neural network (RNN) model with convolutional layer and one long short-term memory (LSTM) layer to perform the classification on a sequence of input. In both models, the convolutional layers serve as feature extractors. The performance of each model is experimentally evaluated and the detailed network structure and preprocessing pipeline are demonstrated.","abstract_html":"Lip reading is the process of speech recognition from solely visual information. The goal of this thesis is to perform a silence vs. speech classification, and to recognize the triphone spoken by a talking head, given only the video using neural network classification models. Two neural network architectures are developed and tested on the AVICAR dataset, including one convolutional neural network (CNN) model with fully connected classification layer, and one recurrent neural network (RNN) model with convolutional layer and one long short-term memory (LSTM) layer to perform the classification on a sequence of input. In both models, the convolutional layers serve as feature extractors. The performance of each model is experimentally evaluated and the detailed network structure and preprocessing pipeline are demonstrated.","abstract_has_math":false,"creators":["Zhu, Tianyilin"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Hasegawa-Johnson, Mark"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2017,"date_issued":"2017-08-10T20:33:19Z","date_published":"2017-08-10T20:33:19Z","updated_at":"2026-07-22T22:24:34Z","subjects":["Lipreading","Convolutional neural network"],"languages":["en"],"rights":["Copyright 2017 Tianyilin Zhu"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/97763","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Hasegawa-Johnson, Mark"]},{"key":"dc:creator","label":"Author","values":["Zhu, Tianyilin"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2017-08-10T20:33:19Z","2019-08-11T09:15:28Z","2017-04-24","2017-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Lipreading","Convolutional neural network"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2017 Tianyilin Zhu"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/97763"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Lip reading is the process of speech recognition from solely visual information. The goal of this thesis is to perform a silence vs. speech classification, and to recognize the triphone spoken by a talking head, given only the video using neural network classification models. Two neural network architectures are developed and tested on the AVICAR dataset, including one convolutional neural network (CNN) model with fully connected classification layer, and one recurrent neural network (RNN) model with convolutional layer and one long short-term memory (LSTM) layer to perform the classification on a sequence of input. In both models, the convolutional layers serve as feature extractors. The performance of each model is experimentally evaluated and the detailed network structure and preprocessing pipeline are demonstrated.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2019-05-01","The student, Tianyilin Zhu, accepted the attached license on 2017-04-20 at 23:56.","The student, Tianyilin Zhu, submitted this Thesis for approval on 2017-04-21 at 00:00.","This Thesis was approved for publication on 2017-04-24 at 12:31.","DSpace SAF Submission Ingestion Package generated from Vireo submission #10955 on 2017-08-10 at 15:06:42","Made available in DSpace on 2017-08-10T20:33:19Z (GMT). No. of bitstreams: 2 ZHU-THESIS-2017.pdf: 1297932 bytes, checksum: 85af1d214423afa5059c70241f2459dc (MD5) LICENSE.txt: 4210 bytes, checksum: 456c6995061c2f56b6fa7571abc74cd0 (MD5) Previous issue date: 2017-04-24","Embargo set by: Colleen Fallaw for item 102816 Lift date: 2019-08-10T21:27:21Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 102816 on 2019-08-11T09:15:28Z."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Lipreading with convolutional and recurrent neural network models"]}]}],"canonical_facts":{"dc:contributor":["Hasegawa-Johnson, Mark"],"dc:creator":["Zhu, Tianyilin"],"dc:date":["2017-08-10T20:33:19Z","2019-08-11T09:15:28Z","2017-04-24","2017-05"],"dc:description":["Lip reading is the process of speech recognition from solely visual information. The goal of this thesis is to perform a silence vs. speech classification, and to recognize the triphone spoken by a talking head, given only the video using neural network classification models. Two neural network architectures are developed and tested on the AVICAR dataset, including one convolutional neural network (CNN) model with fully connected classification layer, and one recurrent neural network (RNN) model with convolutional layer and one long short-term memory (LSTM) layer to perform the classification on a sequence of input. In both models, the convolutional layers serve as feature extractors. The performance of each model is experimentally evaluated and the detailed network structure and preprocessing pipeline are demonstrated.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2019-05-01","The student, Tianyilin Zhu, accepted the attached license on 2017-04-20 at 23:56.","The student, Tianyilin Zhu, submitted this Thesis for approval on 2017-04-21 at 00:00.","This Thesis was approved for publication on 2017-04-24 at 12:31.","DSpace SAF Submission Ingestion Package generated from Vireo submission #10955 on 2017-08-10 at 15:06:42","Made available in DSpace on 2017-08-10T20:33:19Z (GMT). No. of bitstreams: 2 ZHU-THESIS-2017.pdf: 1297932 bytes, checksum: 85af1d214423afa5059c70241f2459dc (MD5) LICENSE.txt: 4210 bytes, checksum: 456c6995061c2f56b6fa7571abc74cd0 (MD5) Previous issue date: 2017-04-24","Embargo set by: Colleen Fallaw for item 102816 Lift date: 2019-08-10T21:27:21Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 102816 on 2019-08-11T09:15:28Z."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/97763"],"dc:language":["en"],"dc:rights":["Copyright 2017 Tianyilin Zhu"],"dc:subject":["Lipreading","Convolutional neural network"],"dc:title":["Lipreading with convolutional and recurrent neural network models"],"dc:type":["text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:34Z"}