{"id":{"repo_id":"calpoly","oai_identifier":"oai:digitalcommons.calpoly.edu:theses-3577"},"canonical_url":"https://search.dev.ndltd.org/etd/calpoly/oai:digitalcommons.calpoly.edu:theses-3577","repository":{"repo_id":"calpoly","name":"Cal Poly","base_url":"https://digitalcommons.calpoly.edu/do/oai/"},"display":{"title":"Visual Speech Recognition Using a 3D Convolutional Neural Network","abstract":"<p>Main stream automatic speech recognition (ASR) makes use of audio data to identify spoken words, however visual speech recognition (VSR) has recently been of increased interest to researchers. VSR is used when audio data is corrupted or missing entirely and also to further enhance the accuracy of audio-based ASR systems. In this research, we present both a framework for building 3D feature cubes of lip data from videos and a 3D convolutional neural network (CNN) architecture for performing classification on a dataset of 100 spoken words, recorded in an uncontrolled envi- ronment. Our 3D-CNN architecture achieves a testing accuracy of 64%, comparable with recent works, but using an input data size that is up to 75% smaller. Overall, our research shows that 3D-CNNs can be successful in finding spatial-temporal features using unsupervised feature extraction and are a suitable choice for VSR-based systems.</p>","abstract_html":"&lt;p&gt;Main stream automatic speech recognition (ASR) makes use of audio data to identify spoken words, however visual speech recognition (VSR) has recently been of increased interest to researchers. VSR is used when audio data is corrupted or missing entirely and also to further enhance the accuracy of audio-based ASR systems. In this research, we present both a framework for building 3D feature cubes of lip data from videos and a 3D convolutional neural network (CNN) architecture for performing classification on a dataset of 100 spoken words, recorded in an uncontrolled envi- ronment. Our 3D-CNN architecture achieves a testing accuracy of 64%, comparable with recent works, but using an input data size that is up to 75% smaller. Overall, our research shows that 3D-CNNs can be successful in finding spatial-temporal features using unsupervised feature extraction and are a suitable choice for VSR-based systems.&lt;/p&gt;","abstract_has_math":false,"creators":["Rochford, Matthew"],"institution":null,"degree_name":"MS in Electrical Engineering","degree_level":null,"degree_discipline":"Electrical Engineering","degree_department":null,"school":null,"contributors":["Jane Zhang","Electrical Engineering","College of Engineering"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019-12-01T08:00:00Z","date_published":"2019-12-01T08:00:00Z","updated_at":"2026-07-24T01:32:13Z","subjects":["Computer Vision","Machine Learning","Artificial Intelligence","Pattern Recognition","Image Processing","Face and Lip Detection","Signal Processing"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["10.15368/theses.2020.7"],"render_values":[{"text":"10.15368/theses.2020.7","href":"https://doi.org/10.15368/theses.2020.7","code":true}]}]},"links":{"outbound_url":"https://digitalcommons.calpoly.edu/theses/2109","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Jane Zhang","Electrical Engineering","College of Engineering"]},{"key":"dc:creator","label":"Author","values":["Rochford, Matthew"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2020-02-04T08:00:00Z"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical Engineering"]},{"key":"thesis:degree_name","label":"Degree Name","values":["MS in Electrical Engineering"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer Vision","Machine Learning","Artificial Intelligence","Pattern Recognition","Image Processing","Face and Lip Detection","Signal Processing"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://digitalcommons.calpoly.edu/theses/2109","10.15368/theses.2020.7"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>Main stream automatic speech recognition (ASR) makes use of audio data to identify spoken words, however visual speech recognition (VSR) has recently been of increased interest to researchers. VSR is used when audio data is corrupted or missing entirely and also to further enhance the accuracy of audio-based ASR systems. In this research, we present both a framework for building 3D feature cubes of lip data from videos and a 3D convolutional neural network (CNN) architecture for performing classification on a dataset of 100 spoken words, recorded in an uncontrolled envi- ronment. Our 3D-CNN architecture achieves a testing accuracy of 64%, comparable with recent works, but using an input data size that is up to 75% smaller. Overall, our research shows that 3D-CNNs can be successful in finding spatial-temporal features using unsupervised feature extraction and are a suitable choice for VSR-based systems.</p>"]},{"key":"dc:title","label":"Title","values":["Visual Speech Recognition Using a 3D Convolutional Neural Network"]}]}],"canonical_facts":{"dc:contributor":["Jane Zhang","Electrical Engineering","College of Engineering"],"dc:creator":["Rochford, Matthew"],"dc:date.available":["2020-02-04T08:00:00Z"],"dc:description.abstract":["<p>Main stream automatic speech recognition (ASR) makes use of audio data to identify spoken words, however visual speech recognition (VSR) has recently been of increased interest to researchers. VSR is used when audio data is corrupted or missing entirely and also to further enhance the accuracy of audio-based ASR systems. In this research, we present both a framework for building 3D feature cubes of lip data from videos and a 3D convolutional neural network (CNN) architecture for performing classification on a dataset of 100 spoken words, recorded in an uncontrolled envi- ronment. Our 3D-CNN architecture achieves a testing accuracy of 64%, comparable with recent works, but using an input data size that is up to 75% smaller. Overall, our research shows that 3D-CNNs can be successful in finding spatial-temporal features using unsupervised feature extraction and are a suitable choice for VSR-based systems.</p>"],"dc:identifier":["https://digitalcommons.calpoly.edu/theses/2109","10.15368/theses.2020.7"],"dc:subject":["Computer Vision","Machine Learning","Artificial Intelligence","Pattern Recognition","Image Processing","Face and Lip Detection","Signal Processing"],"dc:title":["Visual Speech Recognition Using a 3D Convolutional Neural Network"],"thesis:degree_discipline":["Electrical Engineering"],"thesis:degree_name":["MS in Electrical Engineering"]},"updated_at":"2026-07-24T01:32:13Z"}