{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/104888"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/104888","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"The benefits of acoustic perceptual information for speech processing systems","abstract":"The frame-synchronized framework has dominated many speech processing systems, such as ASR and AED targeting human speech activities. These systems have little consideration for the science behind speech and treat the task as a simple statistical classification. The framework also assumes each feature vector to be equally important to the task. However, through some preliminary experiments, this study has found evidence that some concepts defined in speech perception theories such as auditory roughness and acoustic landmarks can act as heuristics to these systems and benefit them in multiple ways. Findings of acoustic landmarks hint that the idea of treating each frame equally might not be optimal. In some cases, landmark information can improve system accuracy through highlighting the more significant frames, or improve the acoustic model accuracy by training through MTL. Further investigation into the topic found experimental evidence suggesting that acoustic landmark information can also benefit end-to-end acoustic models trained through CTC loss. With the help of acoustic landmarks, CTC models can converge with less training data and achieve lower error rate. For the first time, positive results were collected on a mid-size ASR corpus (WSJ) for acoustic landmarks. The results indicate that audio perception information can benefit a broad range of audio processing systems.","abstract_html":"The frame-synchronized framework has dominated many speech processing systems, such as ASR and AED targeting human speech activities. These systems have little consideration for the science behind speech and treat the task as a simple statistical classification. The framework also assumes each feature vector to be equally important to the task. However, through some preliminary experiments, this study has found evidence that some concepts defined in speech perception theories such as auditory roughness and acoustic landmarks can act as heuristics to these systems and benefit them in multiple ways. Findings of acoustic landmarks hint that the idea of treating each frame equally might not be optimal. In some cases, landmark information can improve system accuracy through highlighting the more significant frames, or improve the acoustic model accuracy by training through MTL. Further investigation into the topic found experimental evidence suggesting that acoustic landmark information can also benefit end-to-end acoustic models trained through CTC loss. With the help of acoustic landmarks, CTC models can converge with less training data and achieve lower error rate. For the first time, positive results were collected on a mid-size ASR corpus (WSJ) for acoustic landmarks. The results indicate that audio perception information can benefit a broad range of audio processing systems.","abstract_has_math":false,"creators":["He, Di"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Chen, Deming","Hasegawa-Johnson, Mark","Wong, Martin","Lim, Boon Pang"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019-08-23T20:00:08Z","date_published":"2019-08-23T20:00:08Z","updated_at":"2026-07-22T22:24:42Z","subjects":["ASR","AED","Acoustic Landmark","Auditory Roughness","IoT","FPGA","MTL","CTC"],"languages":["en"],"rights":["Copyright 2019 Di He"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/104888","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Chen, Deming","Hasegawa-Johnson, Mark","Wong, Martin","Lim, Boon Pang"]},{"key":"dc:creator","label":"Author","values":["He, Di"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2019-08-23T20:00:08Z","2019-04-19","2019-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["ASR","AED","Acoustic Landmark","Auditory Roughness","IoT","FPGA","MTL","CTC"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2019 Di He"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/104888"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The frame-synchronized framework has dominated many speech processing systems, such as ASR and AED targeting human speech activities. These systems have little consideration for the science behind speech and treat the task as a simple statistical classification. The framework also assumes each feature vector to be equally important to the task. However, through some preliminary experiments, this study has found evidence that some concepts defined in speech perception theories such as auditory roughness and acoustic landmarks can act as heuristics to these systems and benefit them in multiple ways. Findings of acoustic landmarks hint that the idea of treating each frame equally might not be optimal. In some cases, landmark information can improve system accuracy through highlighting the more significant frames, or improve the acoustic model accuracy by training through MTL. Further investigation into the topic found experimental evidence suggesting that acoustic landmark information can also benefit end-to-end acoustic models trained through CTC loss. With the help of acoustic landmarks, CTC models can converge with less training data and achieve lower error rate. For the first time, positive results were collected on a mid-size ASR corpus (WSJ) for acoustic landmarks. The results indicate that audio perception information can benefit a broad range of audio processing systems.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2019-08-22 without embargo terms","The student, Di He, accepted the attached license on 2019-04-19 at 15:55.","The student, Di He, submitted this Dissertation for approval on 2019-04-19 at 16:05.","This Dissertation was approved for publication on 2019-04-19 at 17:36.","DSpace SAF Submission Ingestion Package generated from Vireo submission #13794 on 2019-08-22 at 14:45:11","Made available in DSpace on 2019-08-23T20:00:08Z (GMT). No. of bitstreams: 3 HE-DISSERTATION-2019.pdf: 3612262 bytes, checksum: e589ec3451c53220a1f0b837f16227b7 (MD5) LICENSE.txt: 4202 bytes, checksum: f6ca289987e697ace633362d46e0cce2 (MD5) PROQUEST_LICENSE.txt: 4548 bytes, checksum: f0eec8de3bc97044b499dfeb10918a3c (MD5) Previous issue date: 2019-04-19"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["The benefits of acoustic perceptual information for speech processing systems"]}]}],"canonical_facts":{"dc:contributor":["Chen, Deming","Hasegawa-Johnson, Mark","Wong, Martin","Lim, Boon Pang"],"dc:creator":["He, Di"],"dc:date":["2019-08-23T20:00:08Z","2019-04-19","2019-05"],"dc:description":["The frame-synchronized framework has dominated many speech processing systems, such as ASR and AED targeting human speech activities. These systems have little consideration for the science behind speech and treat the task as a simple statistical classification. The framework also assumes each feature vector to be equally important to the task. However, through some preliminary experiments, this study has found evidence that some concepts defined in speech perception theories such as auditory roughness and acoustic landmarks can act as heuristics to these systems and benefit them in multiple ways. Findings of acoustic landmarks hint that the idea of treating each frame equally might not be optimal. In some cases, landmark information can improve system accuracy through highlighting the more significant frames, or improve the acoustic model accuracy by training through MTL. Further investigation into the topic found experimental evidence suggesting that acoustic landmark information can also benefit end-to-end acoustic models trained through CTC loss. With the help of acoustic landmarks, CTC models can converge with less training data and achieve lower error rate. For the first time, positive results were collected on a mid-size ASR corpus (WSJ) for acoustic landmarks. The results indicate that audio perception information can benefit a broad range of audio processing systems.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2019-08-22 without embargo terms","The student, Di He, accepted the attached license on 2019-04-19 at 15:55.","The student, Di He, submitted this Dissertation for approval on 2019-04-19 at 16:05.","This Dissertation was approved for publication on 2019-04-19 at 17:36.","DSpace SAF Submission Ingestion Package generated from Vireo submission #13794 on 2019-08-22 at 14:45:11","Made available in DSpace on 2019-08-23T20:00:08Z (GMT). No. of bitstreams: 3 HE-DISSERTATION-2019.pdf: 3612262 bytes, checksum: e589ec3451c53220a1f0b837f16227b7 (MD5) LICENSE.txt: 4202 bytes, checksum: f6ca289987e697ace633362d46e0cce2 (MD5) PROQUEST_LICENSE.txt: 4548 bytes, checksum: f0eec8de3bc97044b499dfeb10918a3c (MD5) Previous issue date: 2019-04-19"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/104888"],"dc:language":["en"],"dc:rights":["Copyright 2019 Di He"],"dc:subject":["ASR","AED","Acoustic Landmark","Auditory Roughness","IoT","FPGA","MTL","CTC"],"dc:title":["The benefits of acoustic perceptual information for speech processing systems"],"dc:type":["text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:42Z"}