{"id":{"repo_id":"rgu","oai_identifier":"oai:rgu-repository.worktribe.com:2807323"},"canonical_url":"https://search.dev.ndltd.org/etd/rgu/oai:rgu-repository.worktribe.com:2807323","repository":{"repo_id":"rgu","name":"Robert Gordon University","base_url":"https://rgu-repository.worktribe.com/oaiprovider"},"display":{"title":"Speaker model adaptation in automatic speech recognition.","abstract":"One of the main obstacles of automatic speech recognition is to achieve speaker independence. It is generally believed that the main difficulty is the inter-speaker variability in which the acoustic characteristics of different speakers are not the same. There are mainly three approaches to overcome this problem; extract invariant information, multiple template representation and speaker adaptation. This thesis describes a study of employing different speaker adaptation techniques in an attempt to improve the performance of a continuous density HMM-based speaker independent speech recogniser. Two alternative approaches which compensate for speaker difference are described. Speaker normalisation in which different speakers are transformed into a common parameter space is briefly outlined. Speaker adaptation in which a speech recogniser is tuned into the new speaker’s acoustic characteristics is examined in detail and some of the previously proposed techniques are outlined. Several different speaker adaptation techniques are investigated in detail. Three of these techniques are based on transformation of spectral parameters of the recogniser into the new speaker’s domain. The first method uses a spectral probabilistic mapping, the second method uses a neural network architecture as non-linear transform. The third method is based on statistical regression analysis. Apart from transform adaptation, two more adaptation techniques which are based on statistical estimation are also described in detail. The first technique is based on Bayesian inference estimation. The second technique, not previously applied in speaker adaptation, is based on statistical time series analysis via a set of recursive Kalman filtering equations. The speaker adaptation techniques are evaluated using two large population speech databases. Each of the techniques is compared with a speaker independent recogniser and a comparison between different techniques is also made. The experimental results have shown that recognition performance improves when only a small amount of data is given from the new speaker for adaptation.","abstract_html":"One of the main obstacles of automatic speech recognition is to achieve speaker independence. It is generally believed that the main difficulty is the inter-speaker variability in which the acoustic characteristics of different speakers are not the same. There are mainly three approaches to overcome this problem; extract invariant information, multiple template representation and speaker adaptation. This thesis describes a study of employing different speaker adaptation techniques in an attempt to improve the performance of a continuous density HMM-based speaker independent speech recogniser. Two alternative approaches which compensate for speaker difference are described. Speaker normalisation in which different speakers are transformed into a common parameter space is briefly outlined. Speaker adaptation in which a speech recogniser is tuned into the new speaker’s acoustic characteristics is examined in detail and some of the previously proposed techniques are outlined. Several different speaker adaptation techniques are investigated in detail. Three of these techniques are based on transformation of spectral parameters of the recogniser into the new speaker’s domain. The first method uses a spectral probabilistic mapping, the second method uses a neural network architecture as non-linear transform. The third method is based on statistical regression analysis. Apart from transform adaptation, two more adaptation techniques which are based on statistical estimation are also described in detail. The first technique is based on Bayesian inference estimation. The second technique, not previously applied in speaker adaptation, is based on statistical time series analysis via a set of recursive Kalman filtering equations. The speaker adaptation techniques are evaluated using two large population speech databases. Each of the techniques is compared with a speaker independent recogniser and a comparison between different techniques is also made. The experimental results have shown that recognition performance improves when only a small amount of data is given from the new speaker for adaptation.","abstract_has_math":false,"creators":["Chan, Carlos Chun Ming"],"institution":"Robert Gordon University","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["J. Tyler, I. McKenzie and S. Cox"],"committee_chairs":[],"committee_members":[],"year":1993,"date_issued":"1993","date_published":"1993","updated_at":"2026-07-24T04:10:09Z","subjects":["Automatic speech recognition","Multiple template representation","Speaker adaptation","Acoustic characteristics","Spectral probabilistic mapping","Neural network","Bayesian inference estimation","Kalman filtering equations"],"languages":["en"],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["oai:rgu-repository.worktribe.com:2807323","https://doi.org/10.48526/rgu-wt-2807323"],"render_values":[{"text":"oai:rgu-repository.worktribe.com:2807323","href":null,"code":true},{"text":"https://doi.org/10.48526/rgu-wt-2807323","href":"https://doi.org/10.48526/rgu-wt-2807323","code":true}]}]},"links":{"outbound_url":"https://rgu-repository.worktribe.com/2807323/1/CHAN%201993%20Speaker%20model%20adaptation%20in","outbound_label":"Repository record","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["J. Tyler, I. McKenzie and S. Cox"]},{"key":"dc:contributor.sponsor","label":"Sponsor","values":["British Telecom"]},{"key":"dc:creator","label":"Author","values":["Chan, Carlos Chun Ming"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["1993-09-29"]},{"key":"dc:date.issued","label":"Date","values":["1993"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["Robert Gordon University"]},{"key":"dc:relation.isreferencedby","label":"Dc Relation Isreferencedby","values":["https://rgu-repository.worktribe.com/output/2807323"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Automatic speech recognition","Multiple template representation","Speaker adaptation","Acoustic characteristics","Spectral probabilistic mapping","Neural network","Bayesian inference estimation","Kalman filtering equations"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["oai:rgu-repository.worktribe.com:2807323","https://doi.org/10.48526/rgu-wt-2807323"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://rgu-repository.worktribe.com/2807323/1/CHAN%201993%20Speaker%20model%20adaptation%20in"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["One of the main obstacles of automatic speech recognition is to achieve speaker independence. It is generally believed that the main difficulty is the inter-speaker variability in which the acoustic characteristics of different speakers are not the same. There are mainly three approaches to overcome this problem; extract invariant information, multiple template representation and speaker adaptation. This thesis describes a study of employing different speaker adaptation techniques in an attempt to improve the performance of a continuous density HMM-based speaker independent speech recogniser. Two alternative approaches which compensate for speaker difference are described. Speaker normalisation in which different speakers are transformed into a common parameter space is briefly outlined. Speaker adaptation in which a speech recogniser is tuned into the new speaker’s acoustic characteristics is examined in detail and some of the previously proposed techniques are outlined. Several different speaker adaptation techniques are investigated in detail. Three of these techniques are based on transformation of spectral parameters of the recogniser into the new speaker’s domain. The first method uses a spectral probabilistic mapping, the second method uses a neural network architecture as non-linear transform. The third method is based on statistical regression analysis. Apart from transform adaptation, two more adaptation techniques which are based on statistical estimation are also described in detail. The first technique is based on Bayesian inference estimation. The second technique, not previously applied in speaker adaptation, is based on statistical time series analysis via a set of recursive Kalman filtering equations. The speaker adaptation techniques are evaluated using two large population speech databases. Each of the techniques is compared with a speaker independent recogniser and a comparison between different techniques is also made. The experimental results have shown that recognition performance improves when only a small amount of data is given from the new speaker for adaptation."]},{"key":"dc:title","label":"Title","values":["Speaker model adaptation in automatic speech recognition."]}]}],"canonical_facts":{"dc:contributor.advisor":["J. Tyler, I. McKenzie and S. Cox"],"dc:contributor.sponsor":["British Telecom"],"dc:creator":["Chan, Carlos Chun Ming"],"dc:date":["1993-09-29"],"dc:date.issued":["1993"],"dc:description.abstract":["One of the main obstacles of automatic speech recognition is to achieve speaker independence. It is generally believed that the main difficulty is the inter-speaker variability in which the acoustic characteristics of different speakers are not the same. There are mainly three approaches to overcome this problem; extract invariant information, multiple template representation and speaker adaptation. This thesis describes a study of employing different speaker adaptation techniques in an attempt to improve the performance of a continuous density HMM-based speaker independent speech recogniser. Two alternative approaches which compensate for speaker difference are described. Speaker normalisation in which different speakers are transformed into a common parameter space is briefly outlined. Speaker adaptation in which a speech recogniser is tuned into the new speaker’s acoustic characteristics is examined in detail and some of the previously proposed techniques are outlined. Several different speaker adaptation techniques are investigated in detail. Three of these techniques are based on transformation of spectral parameters of the recogniser into the new speaker’s domain. The first method uses a spectral probabilistic mapping, the second method uses a neural network architecture as non-linear transform. The third method is based on statistical regression analysis. Apart from transform adaptation, two more adaptation techniques which are based on statistical estimation are also described in detail. The first technique is based on Bayesian inference estimation. The second technique, not previously applied in speaker adaptation, is based on statistical time series analysis via a set of recursive Kalman filtering equations. The speaker adaptation techniques are evaluated using two large population speech databases. Each of the techniques is compared with a speaker independent recogniser and a comparison between different techniques is also made. The experimental results have shown that recognition performance improves when only a small amount of data is given from the new speaker for adaptation."],"dc:identifier":["oai:rgu-repository.worktribe.com:2807323","https://doi.org/10.48526/rgu-wt-2807323"],"dc:identifier.uri":["https://rgu-repository.worktribe.com/2807323/1/CHAN%201993%20Speaker%20model%20adaptation%20in"],"dc:language":["en"],"dc:publisher.institution":["Robert Gordon University"],"dc:relation.isreferencedby":["https://rgu-repository.worktribe.com/output/2807323"],"dc:subject":["Automatic speech recognition","Multiple template representation","Speaker adaptation","Acoustic characteristics","Spectral probabilistic mapping","Neural network","Bayesian inference estimation","Kalman filtering equations"],"dc:title":["Speaker model adaptation in automatic speech recognition."],"dc:type":["Thesis"]},"updated_at":"2026-07-24T04:10:09Z"}