{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/120363"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/120363","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Learning shared semantic space for speech-to-text translation","abstract":"Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2025-05-01","abstract_html":"Submission published under a 24 month embargo labeled &#x27;U of I Access&#x27;, the embargo will last until 2025-05-01","abstract_has_math":false,"creators":["Han, Chi"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Ji, Heng"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2023,"date_issued":"2023-05","date_published":"2023-05","updated_at":"2026-07-22T22:24:57Z","subjects":["Speech-to-text Translation","Natural Language Processing","Representation Learning"],"languages":["en","eng"],"rights":["Copyright 2023 Chi Han"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/120363","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Ji, Heng"]},{"key":"dc:creator","label":"Author","values":["Han, Chi"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2023-05","2023-04-13"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Speech-to-text Translation","Natural Language Processing","Representation Learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2023 Chi Han"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/120363"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2025-05-01","The student, Chi Han, accepted the attached license on 2023-04-12 at 18:51.","The student, Chi Han, submitted this Thesis for approval on 2023-04-12 at 19:24.","This Thesis was approved for publication on 2023-04-13 at 13:39.","DSpace SAF Submission Ingestion Package generated from Vireo submission #18971 on 2023-09-01 at 17:13:23","End-to-end speech translation (ST) has far-reaching implications and numerous potential applications, making it an area of significant interest and impact. Despite its importance, ST has traditionally been treated as a separate task, failing to fully leverage the rapid ad- vancements in its closely related sibling - text machine translation (MT). This separation is due to the modality gap, which results from the different representations of text and audio inputs, rendering MT data and end-to-end models incompatible with their ST counterparts. In light of this challenge, we present Chimera, a novel approach designed to bridge the rep- resentation gap between these two modalities. Chimera achieves this by projecting audio and text features onto a common semantic representation, effectively unifying the MT and ST tasks. Consequently, Chimera enhances the performance on ST benchmarks, such as MuST-C and Augmented Librispeech, setting new state-of-the-art results. More specifically, Chimera attains a 27.1 BLEU score on the MuST-C EN-DE benchmark, improving the existing state-of-the-art by a substantial margin of +1.9 BLEU. Further experimental anal- yses substantiate that the shared semantic space indeed facilitates the exchange of common knowledge between the MT and ST tasks. We discovered identifiable semantic regions within the shared joint speech-text encoding space, highlighting the effective integration of both modalities. By plotting neural activation maps between parallel speech and text, we were able to visualize the convergence of semantic information, further demonstrating the success of our approach in bridging the modality gap and fostering a more robust understanding of the underlying linguistic structures. This finding paves the way for augmenting training resources across modalities and opens up new avenues for exploration in the field of speech translation."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Learning shared semantic space for speech-to-text translation"]}]}],"canonical_facts":{"dc:contributor":["Ji, Heng"],"dc:creator":["Han, Chi"],"dc:date":["2023-05","2023-04-13"],"dc:description":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2025-05-01","The student, Chi Han, accepted the attached license on 2023-04-12 at 18:51.","The student, Chi Han, submitted this Thesis for approval on 2023-04-12 at 19:24.","This Thesis was approved for publication on 2023-04-13 at 13:39.","DSpace SAF Submission Ingestion Package generated from Vireo submission #18971 on 2023-09-01 at 17:13:23","End-to-end speech translation (ST) has far-reaching implications and numerous potential applications, making it an area of significant interest and impact. Despite its importance, ST has traditionally been treated as a separate task, failing to fully leverage the rapid ad- vancements in its closely related sibling - text machine translation (MT). This separation is due to the modality gap, which results from the different representations of text and audio inputs, rendering MT data and end-to-end models incompatible with their ST counterparts. In light of this challenge, we present Chimera, a novel approach designed to bridge the rep- resentation gap between these two modalities. Chimera achieves this by projecting audio and text features onto a common semantic representation, effectively unifying the MT and ST tasks. Consequently, Chimera enhances the performance on ST benchmarks, such as MuST-C and Augmented Librispeech, setting new state-of-the-art results. More specifically, Chimera attains a 27.1 BLEU score on the MuST-C EN-DE benchmark, improving the existing state-of-the-art by a substantial margin of +1.9 BLEU. Further experimental anal- yses substantiate that the shared semantic space indeed facilitates the exchange of common knowledge between the MT and ST tasks. We discovered identifiable semantic regions within the shared joint speech-text encoding space, highlighting the effective integration of both modalities. By plotting neural activation maps between parallel speech and text, we were able to visualize the convergence of semantic information, further demonstrating the success of our approach in bridging the modality gap and fostering a more robust understanding of the underlying linguistic structures. This finding paves the way for augmenting training resources across modalities and opens up new avenues for exploration in the field of speech translation."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/120363"],"dc:language":["en","eng"],"dc:rights":["Copyright 2023 Chi Han"],"dc:subject":["Speech-to-text Translation","Natural Language Processing","Representation Learning"],"dc:title":["Learning shared semantic space for speech-to-text translation"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:57Z"}