{"id":{"repo_id":"rgu","oai_identifier":"oai:rgu-repository.worktribe.com:2807351"},"canonical_url":"https://search.dev.ndltd.org/etd/rgu/oai:rgu-repository.worktribe.com:2807351","repository":{"repo_id":"rgu","name":"Robert Gordon University","base_url":"https://rgu-repository.worktribe.com/oaiprovider"},"display":{"title":"A comparative study of models for automatic speech recognition.","abstract":"This thesis describes a study of the most popular techniques for speech modelling at present, the Dynamic Programming approach, the Hidden Markov Model (HMM), and the Neural Network, which are also evaluated by experiments. The reason why the HMM outperforms the other techniques is examined rigorously in the light of decision theory. The study firstly concludes that the success of the HMM approach is due to its probabilistic representation for the acoustic variability of speech signals, and the transient property of this representation that makes it possible to model the evolution of speech spectra in the course of time. However, the HMM has a limitation when applied to speech in that it assumes that the output observations of this model, which correspond to the spectral vectors of a spoken word, are state-dependent only, implying that these observations are independent of each other when generated by the same state. An effect of this assumption is that the time-ordering information of the spectral vectors within a state is disregarded. As a consequence, the loss of this information limits the performance of HMM-based recognition systems. To account for this time-ordering information, this thesis presents two alternative approaches, namely the Markov model (MM), and the HMM-MM hybrid. Both approaches make use of the Markov property, that the present observation of the Markov process depends on the immediate preceding observation, to model the time-ordering information of speech vectors. Both methods, along with the HMM, have been tested extensively on a task of isolated word recognition using a widely distributed English Alphabet database. The results suggest: (1) the time-ordering of the spectral vectors of a spoken word is important information which may be used to improve the performance of HMM recognition systems; (2) the Markov model attains the performance almost identical to, or better than the HMM (depending on the data used for experiment), and offers a substantial saving in computation time compared with the HMM; and (3) the HMM-MM hybrid has successfully addressed the problem of the temporal information modelling posed by the HMM and the improvement made by this approach is statistically significant. The software developed in this study is included as an appendix.","abstract_html":"This thesis describes a study of the most popular techniques for speech modelling at present, the Dynamic Programming approach, the Hidden Markov Model (HMM), and the Neural Network, which are also evaluated by experiments. The reason why the HMM outperforms the other techniques is examined rigorously in the light of decision theory. The study firstly concludes that the success of the HMM approach is due to its probabilistic representation for the acoustic variability of speech signals, and the transient property of this representation that makes it possible to model the evolution of speech spectra in the course of time. However, the HMM has a limitation when applied to speech in that it assumes that the output observations of this model, which correspond to the spectral vectors of a spoken word, are state-dependent only, implying that these observations are independent of each other when generated by the same state. An effect of this assumption is that the time-ordering information of the spectral vectors within a state is disregarded. As a consequence, the loss of this information limits the performance of HMM-based recognition systems. To account for this time-ordering information, this thesis presents two alternative approaches, namely the Markov model (MM), and the HMM-MM hybrid. Both approaches make use of the Markov property, that the present observation of the Markov process depends on the immediate preceding observation, to model the time-ordering information of speech vectors. Both methods, along with the HMM, have been tested extensively on a task of isolated word recognition using a widely distributed English Alphabet database. The results suggest: (1) the time-ordering of the spectral vectors of a spoken word is important information which may be used to improve the performance of HMM recognition systems; (2) the Markov model attains the performance almost identical to, or better than the HMM (depending on the data used for experiment), and offers a substantial saving in computation time compared with the HMM; and (3) the HMM-MM hybrid has successfully addressed the problem of the temporal information modelling posed by the HMM and the improvement made by this approach is statistically significant. The software developed in this study is included as an appendix.","abstract_has_math":false,"creators":["Dai, Jianing"],"institution":"Robert Gordon's Institute of Technology","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["J. Tyler and I. MacKenzie"],"committee_chairs":[],"committee_members":[],"year":1992,"date_issued":"1992","date_published":"1992","updated_at":"2026-07-24T04:10:09Z","subjects":["Automatic speech recognition","Dynamic programming","Hidden markov model","Neural network","Spectral vectors","Comparative study"],"languages":["en"],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["oai:rgu-repository.worktribe.com:2807351","https://doi.org/10.48526/rgu-wt-2807351"],"render_values":[{"text":"oai:rgu-repository.worktribe.com:2807351","href":null,"code":true},{"text":"https://doi.org/10.48526/rgu-wt-2807351","href":"https://doi.org/10.48526/rgu-wt-2807351","code":true}]}]},"links":{"outbound_url":"https://rgu-repository.worktribe.com/2807351/1/DAI%201992%20A%20comparative%20study%20of%20models","outbound_label":"Repository record","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["J. Tyler and I. MacKenzie"]},{"key":"dc:contributor.sponsor","label":"Sponsor","values":["No Funder Acknowledged (Outputs)"]},{"key":"dc:creator","label":"Author","values":["Dai, Jianing"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["1992-03-31"]},{"key":"dc:date.issued","label":"Date","values":["1992"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["Robert Gordon's Institute of Technology"]},{"key":"dc:relation.isreferencedby","label":"Dc Relation Isreferencedby","values":["https://rgu-repository.worktribe.com/output/2807351"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Automatic speech recognition","Dynamic programming","Hidden markov model","Neural network","Spectral vectors","Comparative study"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["oai:rgu-repository.worktribe.com:2807351","https://doi.org/10.48526/rgu-wt-2807351"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://rgu-repository.worktribe.com/2807351/1/DAI%201992%20A%20comparative%20study%20of%20models"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["This thesis describes a study of the most popular techniques for speech modelling at present, the Dynamic Programming approach, the Hidden Markov Model (HMM), and the Neural Network, which are also evaluated by experiments. The reason why the HMM outperforms the other techniques is examined rigorously in the light of decision theory. The study firstly concludes that the success of the HMM approach is due to its probabilistic representation for the acoustic variability of speech signals, and the transient property of this representation that makes it possible to model the evolution of speech spectra in the course of time. However, the HMM has a limitation when applied to speech in that it assumes that the output observations of this model, which correspond to the spectral vectors of a spoken word, are state-dependent only, implying that these observations are independent of each other when generated by the same state. An effect of this assumption is that the time-ordering information of the spectral vectors within a state is disregarded. As a consequence, the loss of this information limits the performance of HMM-based recognition systems. To account for this time-ordering information, this thesis presents two alternative approaches, namely the Markov model (MM), and the HMM-MM hybrid. Both approaches make use of the Markov property, that the present observation of the Markov process depends on the immediate preceding observation, to model the time-ordering information of speech vectors. Both methods, along with the HMM, have been tested extensively on a task of isolated word recognition using a widely distributed English Alphabet database. The results suggest: (1) the time-ordering of the spectral vectors of a spoken word is important information which may be used to improve the performance of HMM recognition systems; (2) the Markov model attains the performance almost identical to, or better than the HMM (depending on the data used for experiment), and offers a substantial saving in computation time compared with the HMM; and (3) the HMM-MM hybrid has successfully addressed the problem of the temporal information modelling posed by the HMM and the improvement made by this approach is statistically significant. The software developed in this study is included as an appendix."]},{"key":"dc:title","label":"Title","values":["A comparative study of models for automatic speech recognition."]}]}],"canonical_facts":{"dc:contributor.advisor":["J. Tyler and I. MacKenzie"],"dc:contributor.sponsor":["No Funder Acknowledged (Outputs)"],"dc:creator":["Dai, Jianing"],"dc:date":["1992-03-31"],"dc:date.issued":["1992"],"dc:description.abstract":["This thesis describes a study of the most popular techniques for speech modelling at present, the Dynamic Programming approach, the Hidden Markov Model (HMM), and the Neural Network, which are also evaluated by experiments. The reason why the HMM outperforms the other techniques is examined rigorously in the light of decision theory. The study firstly concludes that the success of the HMM approach is due to its probabilistic representation for the acoustic variability of speech signals, and the transient property of this representation that makes it possible to model the evolution of speech spectra in the course of time. However, the HMM has a limitation when applied to speech in that it assumes that the output observations of this model, which correspond to the spectral vectors of a spoken word, are state-dependent only, implying that these observations are independent of each other when generated by the same state. An effect of this assumption is that the time-ordering information of the spectral vectors within a state is disregarded. As a consequence, the loss of this information limits the performance of HMM-based recognition systems. To account for this time-ordering information, this thesis presents two alternative approaches, namely the Markov model (MM), and the HMM-MM hybrid. Both approaches make use of the Markov property, that the present observation of the Markov process depends on the immediate preceding observation, to model the time-ordering information of speech vectors. Both methods, along with the HMM, have been tested extensively on a task of isolated word recognition using a widely distributed English Alphabet database. The results suggest: (1) the time-ordering of the spectral vectors of a spoken word is important information which may be used to improve the performance of HMM recognition systems; (2) the Markov model attains the performance almost identical to, or better than the HMM (depending on the data used for experiment), and offers a substantial saving in computation time compared with the HMM; and (3) the HMM-MM hybrid has successfully addressed the problem of the temporal information modelling posed by the HMM and the improvement made by this approach is statistically significant. The software developed in this study is included as an appendix."],"dc:identifier":["oai:rgu-repository.worktribe.com:2807351","https://doi.org/10.48526/rgu-wt-2807351"],"dc:identifier.uri":["https://rgu-repository.worktribe.com/2807351/1/DAI%201992%20A%20comparative%20study%20of%20models"],"dc:language":["en"],"dc:publisher.institution":["Robert Gordon's Institute of Technology"],"dc:relation.isreferencedby":["https://rgu-repository.worktribe.com/output/2807351"],"dc:subject":["Automatic speech recognition","Dynamic programming","Hidden markov model","Neural network","Spectral vectors","Comparative study"],"dc:title":["A comparative study of models for automatic speech recognition."],"dc:type":["Thesis"]},"updated_at":"2026-07-24T04:10:09Z"}