{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/121386"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/121386","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Explainable artificial intelligence for inclusive automatic speech recognition","abstract":"Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2025-08-01","abstract_html":"Submission published under a 24 month embargo labeled &#x27;U of I Access&#x27;, the embargo will last until 2025-08-01","abstract_has_math":false,"creators":["Lee, Seunghyun"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Hasegawa-Johnson, Mark A"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2023,"date_issued":"2023-08","date_published":"2023-08","updated_at":"2026-07-22T22:24:57Z","subjects":["Inclusive Asr","Explainable Ai","Asr Visualization"],"languages":["en","eng"],"rights":["Copyright 2023 Seunghyun Lee"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/121386","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Hasegawa-Johnson, Mark A"]},{"key":"dc:creator","label":"Author","values":["Lee, Seunghyun"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2023-08","2023-07-21"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Inclusive Asr","Explainable Ai","Asr Visualization"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2023 Seunghyun Lee"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/121386"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2025-08-01","The student, Seunghyun Lee, accepted the attached license on 2023-07-21 at 13:20.","The student, Seunghyun Lee, submitted this Thesis for approval on 2023-07-21 at 13:26.","This Thesis was approved for publication on 2023-07-21 at 14:47.","DSpace SAF Submission Ingestion Package generated from Vireo submission #19777 on 2023-12-04 at 17:18:55","While the widespread adoption of automatic speech recognition (ASR) technology has brought significant benefits to society, it has also highlighted a persistent issue of inequality in access and utilization of technology. Furthermore, in response to the increasing prevalence of artificial intelligence applications, there has been a growing demand for explainable artificial intelligence (XAI). To address the need for interpretability and explainability in ASR, particularly in the context of inclusiveness, this paper aims to visualize the inner workings of the convolutional neural network (CNN) layer and Transformer block in Wav2Vec2.0. This is achieved by calculating the weighted relevance of the connectionist temporal classification (CTC) with respect to the attention and convolutional layers. Leveraging a Wav2Vec2.0 model pre-trained and fine-tuned on LibriSpeech, and testing the model using the Speech Accent Archive, we discovered that the Transformer exhibits a focus on other vowel transcriptions when encountering vowels within a word, whereas it exhibits a more localized attention when transcribing consonants or vowels in non-words absent from its learned vocabulary. Analysis of the weighted convolutional relevance in the first layer of the CNN revealed that different channels concentrate on distinct frequency and time sequences to capture the overall input characteristics. By obtaining a comprehensive understanding of the underlying causes and dynamics behind performance disparities, we can strive to mitigate these disparities and promote a more inclusive ASR technology."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Explainable artificial intelligence for inclusive automatic speech recognition"]}]}],"canonical_facts":{"dc:contributor":["Hasegawa-Johnson, Mark A"],"dc:creator":["Lee, Seunghyun"],"dc:date":["2023-08","2023-07-21"],"dc:description":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2025-08-01","The student, Seunghyun Lee, accepted the attached license on 2023-07-21 at 13:20.","The student, Seunghyun Lee, submitted this Thesis for approval on 2023-07-21 at 13:26.","This Thesis was approved for publication on 2023-07-21 at 14:47.","DSpace SAF Submission Ingestion Package generated from Vireo submission #19777 on 2023-12-04 at 17:18:55","While the widespread adoption of automatic speech recognition (ASR) technology has brought significant benefits to society, it has also highlighted a persistent issue of inequality in access and utilization of technology. Furthermore, in response to the increasing prevalence of artificial intelligence applications, there has been a growing demand for explainable artificial intelligence (XAI). To address the need for interpretability and explainability in ASR, particularly in the context of inclusiveness, this paper aims to visualize the inner workings of the convolutional neural network (CNN) layer and Transformer block in Wav2Vec2.0. This is achieved by calculating the weighted relevance of the connectionist temporal classification (CTC) with respect to the attention and convolutional layers. Leveraging a Wav2Vec2.0 model pre-trained and fine-tuned on LibriSpeech, and testing the model using the Speech Accent Archive, we discovered that the Transformer exhibits a focus on other vowel transcriptions when encountering vowels within a word, whereas it exhibits a more localized attention when transcribing consonants or vowels in non-words absent from its learned vocabulary. Analysis of the weighted convolutional relevance in the first layer of the CNN revealed that different channels concentrate on distinct frequency and time sequences to capture the overall input characteristics. By obtaining a comprehensive understanding of the underlying causes and dynamics behind performance disparities, we can strive to mitigate these disparities and promote a more inclusive ASR technology."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/121386"],"dc:language":["en","eng"],"dc:rights":["Copyright 2023 Seunghyun Lee"],"dc:subject":["Inclusive Asr","Explainable Ai","Asr Visualization"],"dc:title":["Explainable artificial intelligence for inclusive automatic speech recognition"],"dc:type":["text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:57Z"}