{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/124335"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/124335","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Speech classification and lexical semantic modeling via self-supervision and knowledge transfer","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-09-16 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2024-09-16 without embargo terms","abstract_has_math":false,"creators":["Harvill, John Bowman"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Hasegawa-Johnson, Mark","Ahuja, Narendra","Hockenmaier, Julia","Shomorony, Ilan"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-05","date_published":"2024-05","updated_at":"2026-07-22T22:25:00Z","subjects":["Self-supervision","Knowledge Transfer","Speech Classification","Lexical Semantics"],"languages":["en","eng"],"rights":["Copyright 2024 John Harvill"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/124335","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Hasegawa-Johnson, Mark","Ahuja, Narendra","Hockenmaier, Julia","Shomorony, Ilan"]},{"key":"dc:creator","label":"Author","values":["Harvill, John Bowman"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2024-05","2024-04-22"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Self-supervision","Knowledge Transfer","Speech Classification","Lexical Semantics"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2024 John Harvill"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/124335"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-09-16 without embargo terms","The student, John Harvill, accepted the attached license on 2024-04-18 at 22:10.","The student, John Harvill, submitted this Dissertation for approval on 2024-04-18 at 22:20.","This Dissertation was approved for publication on 2024-04-22 at 11:12.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20490 on 2024-09-16 at 00:35:26","The field of speech and natural language processing has experienced dramatic progress over the past decade due to a major paradigm shift. Instead of using training data for a target task only, modern speech and text applications rely on pretraining as the first step. After pretraining a model on a certain task, or potentially multiple tasks, the knowledge that was learned can be transferred to a downstream task and lead to large performance gains. Given that labeled data is much more challenging to collect than raw speech or text, the most explosive growth in the field has come from discovering effective ways to perform pretraining in a self-supervised fashion. By cleverly manipulating a raw speech waveform or raw text, it is possible to learn an immense amount of information without requiring annotations from humans. In this dissertation, I explore several speech and text tasks that benefit from self-supervision and knowledge transfer. For speech, I demonstrate that for both the stutter detection and device arbitration problems, tailored self-supervised pretraining schemes can be developed that lead to significant performance gains compared to relying on labeled data only. For stutter detection, I propose the idea of creating artificial stuttered speech from healthy speech and using it for pretraining. I also show that knowledge of whether stuttering occurs somewhere within a window of several seconds of speech audio can be used to learn the location of stuttering to a much finer degree via multiple instance learning. For device arbitration, I show that contrastive learning and autoencoding can both create useful representations of acoustic information that improve the ability of an arbitration system to determine which voice assistant is closest to a user. In the text domain, I explore lexical semantic modeling, exemplification modeling, and Automatic Speech Recognition (ASR) error detection and correction. Similar to the speech tasks, I find that all text-based tasks can be improved via knowledge transfer, self-supervision, or a combination of the two. For lexical semantic modeling, I propose a graph-based solution and find that knowledge from many languages is required to perform well on any single language. For exemplification modeling, I propose an autoencoding technique that can effectively isolate information related to contextual meaning of a target polysemous word and generate new, diverse sentences using that word with the intended meaning. For ASR error detection and correction, I show that significant errors can be detected to a high degree of accuracy by combining knowledge from both a sentence-level semantic encoder and Large Language Model (LLM) and highlight the existence of statistical bias within correction and detection models."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Speech classification and lexical semantic modeling via self-supervision and knowledge transfer"]}]}],"canonical_facts":{"dc:contributor":["Hasegawa-Johnson, Mark","Ahuja, Narendra","Hockenmaier, Julia","Shomorony, Ilan"],"dc:creator":["Harvill, John Bowman"],"dc:date":["2024-05","2024-04-22"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-09-16 without embargo terms","The student, John Harvill, accepted the attached license on 2024-04-18 at 22:10.","The student, John Harvill, submitted this Dissertation for approval on 2024-04-18 at 22:20.","This Dissertation was approved for publication on 2024-04-22 at 11:12.","DSpace SAF Submission Ingestion Package generated from Vireo submission #20490 on 2024-09-16 at 00:35:26","The field of speech and natural language processing has experienced dramatic progress over the past decade due to a major paradigm shift. Instead of using training data for a target task only, modern speech and text applications rely on pretraining as the first step. After pretraining a model on a certain task, or potentially multiple tasks, the knowledge that was learned can be transferred to a downstream task and lead to large performance gains. Given that labeled data is much more challenging to collect than raw speech or text, the most explosive growth in the field has come from discovering effective ways to perform pretraining in a self-supervised fashion. By cleverly manipulating a raw speech waveform or raw text, it is possible to learn an immense amount of information without requiring annotations from humans. In this dissertation, I explore several speech and text tasks that benefit from self-supervision and knowledge transfer. For speech, I demonstrate that for both the stutter detection and device arbitration problems, tailored self-supervised pretraining schemes can be developed that lead to significant performance gains compared to relying on labeled data only. For stutter detection, I propose the idea of creating artificial stuttered speech from healthy speech and using it for pretraining. I also show that knowledge of whether stuttering occurs somewhere within a window of several seconds of speech audio can be used to learn the location of stuttering to a much finer degree via multiple instance learning. For device arbitration, I show that contrastive learning and autoencoding can both create useful representations of acoustic information that improve the ability of an arbitration system to determine which voice assistant is closest to a user. In the text domain, I explore lexical semantic modeling, exemplification modeling, and Automatic Speech Recognition (ASR) error detection and correction. Similar to the speech tasks, I find that all text-based tasks can be improved via knowledge transfer, self-supervision, or a combination of the two. For lexical semantic modeling, I propose a graph-based solution and find that knowledge from many languages is required to perform well on any single language. For exemplification modeling, I propose an autoencoding technique that can effectively isolate information related to contextual meaning of a target polysemous word and generate new, diverse sentences using that word with the intended meaning. For ASR error detection and correction, I show that significant errors can be detected to a high degree of accuracy by combining knowledge from both a sentence-level semantic encoder and Large Language Model (LLM) and highlight the existence of statistical bias within correction and detection models."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/124335"],"dc:language":["en","eng"],"dc:rights":["Copyright 2024 John Harvill"],"dc:subject":["Self-supervision","Knowledge Transfer","Speech Classification","Lexical Semantics"],"dc:title":["Speech classification and lexical semantic modeling via self-supervision and knowledge transfer"],"dc:type":["text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:00Z"}