{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/129378"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/129378","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Towards unsupervised speech technology with fewer resources","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2025-10-19 without embargo terms","abstract_has_math":false,"creators":["Ni, Junrui"],"institution":"University of Illinois Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Hasegawa-Johnson, Mark","Schwing, Alexander","Bhat, Suma","Shomorony, Ilan"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-03-31","date_published":"2025-03-31","updated_at":"2026-07-22T22:25:05Z","subjects":["Unsupervised speech processing","Automatic speech recognition","Text-to-speech synthesis"],"languages":["en","eng"],"rights":["Copyright 2025 Junrui Ni"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/129378","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Hasegawa-Johnson, Mark","Schwing, Alexander","Bhat, Suma","Shomorony, Ilan"]},{"key":"dc:creator","label":"Author","values":["Ni, Junrui"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-03-31","2025-05"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Unsupervised speech processing","Automatic speech recognition","Text-to-speech synthesis"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Junrui Ni"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/129378"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","The student, Junrui Ni, accepted the attached license on 2025-03-28 at 09:45.","The student, Junrui Ni, submitted this Dissertation for approval on 2025-03-28 at 10:08.","This Dissertation was approved for publication on 2025-03-31 at 11:08.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21695 on 2025-10-19 at 18:17:55","Recent advancements in supervised automatic speech recognition (ASR) and text-to-speech synthesis (TTS) have achieved remarkable performance, primarily driven by the increasing availability of large transcribed speech corpora. In this work, we propose an unsupervised TTS system and a whole-word unsupervised ASR system, aiming to advance fully unsupervised speech technology that can be developed with minimal supervision. An unsupervised TTS system learns to generate the speech waveform corresponding to any written sentence in a language by observing: 1) a collection of untranscribed speech waveforms in that language; 2) a collection of texts written in that language without access to any transcribed speech. Developing such a system can significantly improve the availability of speech technology to languages without a large amount of parallel speech and text data. This work proposes an unsupervised TTS system that trains on the pseudo-transcripts from an unsupervised ASR system. Our unsupervised system can achieve comparable performance to the supervised system in seven languages with about 10-20 hours of speech each. A careful study on the effect of text units and vocoders has also been conducted to better understand what factors may affect unsupervised TTS performance. We further tackle the existing challenge of developing ASR systems without paired speech and text corpora by proposing the removal of reliance on a phoneme lexicon. We explore a new research direction: word-level unsupervised ASR and experimentally demonstrate that an unsupervised speech recognizer can emerge from joint speech-to-speech and text-to-text masked token-infilling. Using a curated speech corpus containing a fixed number of English words, our system iteratively refines the word segmentation structure and achieves a word error rate of between 20-23%, depending on the vocabulary size, without parallel transcripts, oracle word boundaries, or a pronunciation lexicon. This innovative model surpasses the performance of previous unsupervised ASR models under the lexicon-free setting."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Towards unsupervised speech technology with fewer resources"]}]}],"canonical_facts":{"dc:contributor":["Hasegawa-Johnson, Mark","Schwing, Alexander","Bhat, Suma","Shomorony, Ilan"],"dc:creator":["Ni, Junrui"],"dc:date":["2025-03-31","2025-05"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-19 without embargo terms","The student, Junrui Ni, accepted the attached license on 2025-03-28 at 09:45.","The student, Junrui Ni, submitted this Dissertation for approval on 2025-03-28 at 10:08.","This Dissertation was approved for publication on 2025-03-31 at 11:08.","DSpace SAF Submission Ingestion Package generated from Vireo submission #21695 on 2025-10-19 at 18:17:55","Recent advancements in supervised automatic speech recognition (ASR) and text-to-speech synthesis (TTS) have achieved remarkable performance, primarily driven by the increasing availability of large transcribed speech corpora. In this work, we propose an unsupervised TTS system and a whole-word unsupervised ASR system, aiming to advance fully unsupervised speech technology that can be developed with minimal supervision. An unsupervised TTS system learns to generate the speech waveform corresponding to any written sentence in a language by observing: 1) a collection of untranscribed speech waveforms in that language; 2) a collection of texts written in that language without access to any transcribed speech. Developing such a system can significantly improve the availability of speech technology to languages without a large amount of parallel speech and text data. This work proposes an unsupervised TTS system that trains on the pseudo-transcripts from an unsupervised ASR system. Our unsupervised system can achieve comparable performance to the supervised system in seven languages with about 10-20 hours of speech each. A careful study on the effect of text units and vocoders has also been conducted to better understand what factors may affect unsupervised TTS performance. We further tackle the existing challenge of developing ASR systems without paired speech and text corpora by proposing the removal of reliance on a phoneme lexicon. We explore a new research direction: word-level unsupervised ASR and experimentally demonstrate that an unsupervised speech recognizer can emerge from joint speech-to-speech and text-to-text masked token-infilling. Using a curated speech corpus containing a fixed number of English words, our system iteratively refines the word segmentation structure and achieves a word error rate of between 20-23%, depending on the vocabulary size, without parallel transcripts, oracle word boundaries, or a pronunciation lexicon. This innovative model surpasses the performance of previous unsupervised ASR models under the lexicon-free setting."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/129378"],"dc:language":["en","eng"],"dc:rights":["Copyright 2025 Junrui Ni"],"dc:subject":["Unsupervised speech processing","Automatic speech recognition","Text-to-speech synthesis"],"dc:title":["Towards unsupervised speech technology with fewer resources"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:05Z"}