{"id":{"repo_id":"nus","oai_identifier":"oai:scholarbank.nus.edu.sg:10635/14759"},"canonical_url":"https://search.dev.ndltd.org/etd/nus/oai:scholarbank.nus.edu.sg:10635/14759","repository":{"repo_id":"nus","name":"National University of Singapore","base_url":"https://scholarbank.nus.edu.sg/oai/request"},"display":{"title":"Hidden Markov model based visual speech recognition","abstract":"Speech recognition can be made more accurate if visual speech information such as the movement of the lips is taken into consideration. In this thesis, studies on visual speech processing are presented. Classifiers based on Hidden Markov Model (HMM) are first explored for modeling and identifying the basic visual speech elements. Considering that the temporal features of visual speech elements may be confusable and sensitive to their contexts, three novel training strategies, referred to as two-channel training strategy, Maximum Separable Distance training strategy and HMM Adaptive Boosting strategy, are proposed to improve the discriminative power and robustness of an HMM classifier. Following that, approaches for recognizing words, phrases and connected-digit units in visual speech are presented with exploration of level building method and Viterbi searching algorithm. The thesis also covers the studies of 3D lip tracking and visual speech mapping between different speakers. These approaches may extend the applicability of a visual speech processing system to unfavorable conditions.","abstract_html":"Speech recognition can be made more accurate if visual speech information such as the movement of the lips is taken into consideration. In this thesis, studies on visual speech processing are presented. Classifiers based on Hidden Markov Model (HMM) are first explored for modeling and identifying the basic visual speech elements. Considering that the temporal features of visual speech elements may be confusable and sensitive to their contexts, three novel training strategies, referred to as two-channel training strategy, Maximum Separable Distance training strategy and HMM Adaptive Boosting strategy, are proposed to improve the discriminative power and robustness of an HMM classifier. Following that, approaches for recognizing words, phrases and connected-digit units in visual speech are presented with exploration of level building method and Viterbi searching algorithm. The thesis also covers the studies of 3D lip tracking and visual speech mapping between different speakers. These approaches may extend the applicability of a visual speech processing system to unfavorable conditions.","abstract_has_math":false,"creators":["DONG LIANG"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2005,"date_issued":"2005-06-02","date_published":"2005-06-02","updated_at":"2026-07-24T03:31:26Z","subjects":["Hidden Markov Model, visual speech recognition, viseme, separable distance, adaptive boosting, level building"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":null,"outbound_label":null,"outbound_source":null},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["DONG LIANG"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2005-06-02"]},{"key":"dc:relation.isreferencedby","label":"Dc Relation Isreferencedby","values":["https://scholarbank.nus.edu.sg/handle/10635/14759"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Hidden Markov Model, visual speech recognition, viseme, separable distance, adaptive boosting, level building"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://scholarbank.nus.edu.sg/bitstreams/23ce770a-f2d1-414b-966c-3bf43a202f66/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Speech recognition can be made more accurate if visual speech information such as the movement of the lips is taken into consideration. In this thesis, studies on visual speech processing are presented. Classifiers based on Hidden Markov Model (HMM) are first explored for modeling and identifying the basic visual speech elements. Considering that the temporal features of visual speech elements may be confusable and sensitive to their contexts, three novel training strategies, referred to as two-channel training strategy, Maximum Separable Distance training strategy and HMM Adaptive Boosting strategy, are proposed to improve the discriminative power and robustness of an HMM classifier. Following that, approaches for recognizing words, phrases and connected-digit units in visual speech are presented with exploration of level building method and Viterbi searching algorithm. The thesis also covers the studies of 3D lip tracking and visual speech mapping between different speakers. These approaches may extend the applicability of a visual speech processing system to unfavorable conditions."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["4f5bbeac8c2ed4ddc7f5fae593b19afe","5c523f163c35781da0f43edbae143d97"]},{"key":"dc:title","label":"Title","values":["Hidden Markov model based visual speech recognition"]}]}],"canonical_facts":{"dc:creator":["DONG LIANG"],"dc:date.issued":["2005-06-02"],"dc:description.abstract":["Speech recognition can be made more accurate if visual speech information such as the movement of the lips is taken into consideration. In this thesis, studies on visual speech processing are presented. Classifiers based on Hidden Markov Model (HMM) are first explored for modeling and identifying the basic visual speech elements. Considering that the temporal features of visual speech elements may be confusable and sensitive to their contexts, three novel training strategies, referred to as two-channel training strategy, Maximum Separable Distance training strategy and HMM Adaptive Boosting strategy, are proposed to improve the discriminative power and robustness of an HMM classifier. Following that, approaches for recognizing words, phrases and connected-digit units in visual speech are presented with exploration of level building method and Viterbi searching algorithm. The thesis also covers the studies of 3D lip tracking and visual speech mapping between different speakers. These approaches may extend the applicability of a visual speech processing system to unfavorable conditions."],"dc:format.checksum.md5":["4f5bbeac8c2ed4ddc7f5fae593b19afe","5c523f163c35781da0f43edbae143d97"],"dc:identifier.uri":["https://scholarbank.nus.edu.sg/bitstreams/23ce770a-f2d1-414b-966c-3bf43a202f66/download"],"dc:relation.isreferencedby":["https://scholarbank.nus.edu.sg/handle/10635/14759"],"dc:subject":["Hidden Markov Model, visual speech recognition, viseme, separable distance, adaptive boosting, level building"],"dc:title":["Hidden Markov model based visual speech recognition"],"dc:type":["Thesis"]},"updated_at":"2026-07-24T03:31:26Z"}