{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/105187"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/105187","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Dealing with linguistic mismatches for automatic speech recognition","abstract":"Recent breakthroughs in automatic speech recognition (ASR) have resulted in a word error rate (WER) on par with human transcribers on the English Switchboard benchmark. However, dealing with linguistic mismatches between the training and testing data is still a significant challenge that remains unsolved. Under the monolingual environment, it is well-known that the performance of ASR systems degrades significantly when presented with the speech from speakers with different accents, dialects, and speaking styles than those encountered during system training. Under the multi-lingual environment, ASR systems trained on a source language achieve even worse performance when tested on another target language because of mismatches in terms of the number of phonemes, lexical ambiguity, and power of phonotactic constraints provided by phone-level n-grams. In order to address the issues of linguistic mismatches for current ASR systems, my dissertation investigates both knowledge-gnostic and knowledge-agnostic solutions. In the first part, classic theories relevant to acoustics and articulatory phonetics that present capability of being transferred across a dialect continuum from local dialects to another standardized language are re-visited. Experiments demonstrate the potentials that acoustic correlates in the vicinity of landmarks could help to build a bridge for dealing with mismatches across difference local or global varieties in a dialect continuum. In the second part, we design an end-to-end acoustic modeling approach based on connectionist temporal classification loss and propose to link the training of acoustics and accent altogether in a manner similar to the learning process in human speech perception. This joint model not only performed well on ASR with multiple accents but also boosted accuracies of accent identification task in comparison to separately-trained models.","abstract_html":"Recent breakthroughs in automatic speech recognition (ASR) have resulted in a word error rate (WER) on par with human transcribers on the English Switchboard benchmark. However, dealing with linguistic mismatches between the training and testing data is still a significant challenge that remains unsolved. Under the monolingual environment, it is well-known that the performance of ASR systems degrades significantly when presented with the speech from speakers with different accents, dialects, and speaking styles than those encountered during system training. Under the multi-lingual environment, ASR systems trained on a source language achieve even worse performance when tested on another target language because of mismatches in terms of the number of phonemes, lexical ambiguity, and power of phonotactic constraints provided by phone-level n-grams. In order to address the issues of linguistic mismatches for current ASR systems, my dissertation investigates both knowledge-gnostic and knowledge-agnostic solutions. In the first part, classic theories relevant to acoustics and articulatory phonetics that present capability of being transferred across a dialect continuum from local dialects to another standardized language are re-visited. Experiments demonstrate the potentials that acoustic correlates in the vicinity of landmarks could help to build a bridge for dealing with mismatches across difference local or global varieties in a dialect continuum. In the second part, we design an end-to-end acoustic modeling approach based on connectionist temporal classification loss and propose to link the training of acoustics and accent altogether in a manner similar to the learning process in human speech perception. This joint model not only performed well on ASR with multiple accents but also boosted accuracies of accent identification task in comparison to separately-trained models.","abstract_has_math":false,"creators":["Yang, Xuesong"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Informatics","degree_department":null,"school":null,"contributors":["Hasegawa-Johnson, Mark","Huang, Thomas S.","Smaragdis, Paris","Shih, Chilin"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019-08-23T20:44:44Z","date_published":"2019-08-23T20:44:44Z","updated_at":"2026-07-22T22:24:44Z","subjects":["Automatic Speech Recognition","Acoustic Modeling","Multi-Accents","Multi-Lingual","Acoustic Phonetics","Distinctive Features","Acoustic Landmarks","End-to-End","Multi-Task Learning","Model Compression","Deep Learning","Pronunciation Error Detection","Connectionist Temporal Classification"],"languages":["en"],"rights":["Copyright 2019 Xuesong Yang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/105187","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Hasegawa-Johnson, Mark","Huang, Thomas S.","Smaragdis, Paris","Shih, Chilin"]},{"key":"dc:creator","label":"Author","values":["Yang, Xuesong"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2019-08-23T20:44:44Z","2021-08-24T09:15:20Z","2019-04-15","2019-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Informatics"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Automatic Speech Recognition","Acoustic Modeling","Multi-Accents","Multi-Lingual","Acoustic Phonetics","Distinctive Features","Acoustic Landmarks","End-to-End","Multi-Task Learning","Model Compression","Deep Learning","Pronunciation Error Detection","Connectionist Temporal Classification"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2019 Xuesong Yang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/105187"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Recent breakthroughs in automatic speech recognition (ASR) have resulted in a word error rate (WER) on par with human transcribers on the English Switchboard benchmark. However, dealing with linguistic mismatches between the training and testing data is still a significant challenge that remains unsolved. Under the monolingual environment, it is well-known that the performance of ASR systems degrades significantly when presented with the speech from speakers with different accents, dialects, and speaking styles than those encountered during system training. Under the multi-lingual environment, ASR systems trained on a source language achieve even worse performance when tested on another target language because of mismatches in terms of the number of phonemes, lexical ambiguity, and power of phonotactic constraints provided by phone-level n-grams. In order to address the issues of linguistic mismatches for current ASR systems, my dissertation investigates both knowledge-gnostic and knowledge-agnostic solutions. In the first part, classic theories relevant to acoustics and articulatory phonetics that present capability of being transferred across a dialect continuum from local dialects to another standardized language are re-visited. Experiments demonstrate the potentials that acoustic correlates in the vicinity of landmarks could help to build a bridge for dealing with mismatches across difference local or global varieties in a dialect continuum. In the second part, we design an end-to-end acoustic modeling approach based on connectionist temporal classification loss and propose to link the training of acoustics and accent altogether in a manner similar to the learning process in human speech perception. This joint model not only performed well on ASR with multiple accents but also boosted accuracies of accent identification task in comparison to separately-trained models.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2021-05-01","The student, Xuesong Yang, accepted the attached license on 2019-04-12 at 16:15.","The student, Xuesong Yang, submitted this Dissertation for approval on 2019-04-12 at 16:32.","This Dissertation was approved for publication on 2019-04-15 at 07:56.","DSpace SAF Submission Ingestion Package generated from Vireo submission #13594 on 2019-08-22 at 16:21:11","Made available in DSpace on 2019-08-23T20:44:44Z (GMT). No. of bitstreams: 3 YANG-DISSERTATION-2019.pdf: 4002927 bytes, checksum: 1726669a2ef5b581ca5be57181c17f6f (MD5) LICENSE.txt: 4209 bytes, checksum: c970d1155109328014035b22f5f9e2a4 (MD5) PROQUEST_LICENSE.txt: 4555 bytes, checksum: 9bf4935e8d3934b6834b4ae4530834ba (MD5) Previous issue date: 2019-04-15","Embargo set by: Seth Robbins for item 112306 Lift date: 2021-08-23T20:44:50Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 112306 Lift date: 2021-08-23T20:46:41Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 112306 Lift date: 2021-08-23T20:47:38Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 112306 Lift date: 2021-08-23T20:48:32Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 112306 on 2021-08-24T09:15:20Z."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Dealing with linguistic mismatches for automatic speech recognition"]}]}],"canonical_facts":{"dc:contributor":["Hasegawa-Johnson, Mark","Huang, Thomas S.","Smaragdis, Paris","Shih, Chilin"],"dc:creator":["Yang, Xuesong"],"dc:date":["2019-08-23T20:44:44Z","2021-08-24T09:15:20Z","2019-04-15","2019-05"],"dc:description":["Recent breakthroughs in automatic speech recognition (ASR) have resulted in a word error rate (WER) on par with human transcribers on the English Switchboard benchmark. However, dealing with linguistic mismatches between the training and testing data is still a significant challenge that remains unsolved. Under the monolingual environment, it is well-known that the performance of ASR systems degrades significantly when presented with the speech from speakers with different accents, dialects, and speaking styles than those encountered during system training. Under the multi-lingual environment, ASR systems trained on a source language achieve even worse performance when tested on another target language because of mismatches in terms of the number of phonemes, lexical ambiguity, and power of phonotactic constraints provided by phone-level n-grams. In order to address the issues of linguistic mismatches for current ASR systems, my dissertation investigates both knowledge-gnostic and knowledge-agnostic solutions. In the first part, classic theories relevant to acoustics and articulatory phonetics that present capability of being transferred across a dialect continuum from local dialects to another standardized language are re-visited. Experiments demonstrate the potentials that acoustic correlates in the vicinity of landmarks could help to build a bridge for dealing with mismatches across difference local or global varieties in a dialect continuum. In the second part, we design an end-to-end acoustic modeling approach based on connectionist temporal classification loss and propose to link the training of acoustics and accent altogether in a manner similar to the learning process in human speech perception. This joint model not only performed well on ASR with multiple accents but also boosted accuracies of accent identification task in comparison to separately-trained models.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2021-05-01","The student, Xuesong Yang, accepted the attached license on 2019-04-12 at 16:15.","The student, Xuesong Yang, submitted this Dissertation for approval on 2019-04-12 at 16:32.","This Dissertation was approved for publication on 2019-04-15 at 07:56.","DSpace SAF Submission Ingestion Package generated from Vireo submission #13594 on 2019-08-22 at 16:21:11","Made available in DSpace on 2019-08-23T20:44:44Z (GMT). No. of bitstreams: 3 YANG-DISSERTATION-2019.pdf: 4002927 bytes, checksum: 1726669a2ef5b581ca5be57181c17f6f (MD5) LICENSE.txt: 4209 bytes, checksum: c970d1155109328014035b22f5f9e2a4 (MD5) PROQUEST_LICENSE.txt: 4555 bytes, checksum: 9bf4935e8d3934b6834b4ae4530834ba (MD5) Previous issue date: 2019-04-15","Embargo set by: Seth Robbins for item 112306 Lift date: 2021-08-23T20:44:50Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 112306 Lift date: 2021-08-23T20:46:41Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 112306 Lift date: 2021-08-23T20:47:38Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 112306 Lift date: 2021-08-23T20:48:32Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 112306 on 2021-08-24T09:15:20Z."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/105187"],"dc:language":["en"],"dc:rights":["Copyright 2019 Xuesong Yang"],"dc:subject":["Automatic Speech Recognition","Acoustic Modeling","Multi-Accents","Multi-Lingual","Acoustic Phonetics","Distinctive Features","Acoustic Landmarks","End-to-End","Multi-Task Learning","Model Compression","Deep Learning","Pronunciation Error Detection","Connectionist Temporal Classification"],"dc:title":["Dealing with linguistic mismatches for automatic speech recognition"],"dc:type":["text"],"thesis:degree_discipline":["Informatics"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:44Z"}