{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/122004"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/122004","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Exploiting monolingual data for neural machine translation","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-03-01 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2024-03-01 without embargo terms","abstract_has_math":false,"creators":["Wang, Yiren"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Zhai, ChengXiang","Hockenmaier, Julia","Ji, Heng","Awadalla, Hany Hassan"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2023,"date_issued":"2023-12","date_published":"2023-12","updated_at":"2026-07-22T22:25:00Z","subjects":["Neural Machine Translation","Dual Learning","Multi-task Learning"],"languages":["en","eng"],"rights":["Copyright 2023 Yiren Wang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/122004","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Zhai, ChengXiang","Hockenmaier, Julia","Ji, Heng","Awadalla, Hany Hassan"]},{"key":"dc:creator","label":"Author","values":["Wang, Yiren"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2023-12","2023-11-21"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Neural Machine Translation","Dual Learning","Multi-task Learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2023 Yiren Wang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/122004"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-03-01 without embargo terms","The student, Yiren Wang, accepted the attached license on 2023-11-20 at 17:20.","The student, Yiren Wang, submitted this Dissertation for approval on 2023-11-20 at 17:27.","This Dissertation was approved for publication on 2023-11-21 at 12:11.","DSpace SAF Submission Ingestion Package generated from Vireo submission #19982 on 2024-03-01 at 13:14:51","Neural Machine Translation (NMT) has made rapid progress in recent years and achieved remarkable translation quality. NMT systems heavily rely on large-scale and high-quality parallel training data of source and target languages (bitext), which is limited and costly to collect. In the meantime, large amount of unlabeled monolingual data is available in different languages. While the target-side monolingual data has been proven useful through Back Translation, the source-side data is less investigated. Effectively exploiting both source-side and target-side monolingual data could not only help boost NMT performance, but also improve data efficiency. In this thesis, we focus on exploring novel approaches that effectively utilize monolingual data to improve the translation quality of NMT models. First, we exploit the potential of monolingual data with dual learning. Dual learning effectively utilize both bitext and monolingual data by leveraging the primal-dual structure of artificial intelligence tasks to generate informative feedback signals to regularize training. We extend the standard dual learning framework by introducing multiple primal and dual models, and propose the Multi-Agent Dual Learning (MADL) to further boost data utilization and translation quality. We show the effectiveness of our proposed approach on multiple benchmark NMT tasks, including both supervised and unsupervised NMT. Second, we study the effect of forward and back translation data at scale, and propose a large-scale noisy training pipeline to effectively leverage both of them. First, we generate synthetic bitext with both forward and back translation. Next, we train the NMT model on a noisy version of this synthetic corpus, where each source sentence is randomly corrupted. Finally, the model is fine-tuned on the genuine bitext and a clean, high-quality subset of the synthetic bitext without any noise. With our proposed strategy, we achieve the state-of-the-art performances on various benchmark translation tasks. Third, we explore more general strategies to utilize monolingual data for not only high-resource language pairs in bilingual NMT, but also low-resource language pairs and multilingual NMT as well. We propose a multi-task learning (MTL) framework, which jointly trains the model with the translation task on bitext data, and two auxiliary self-supervised learning tasks on the source-side and target-side monolingual data respectively. We show that our approach can effectively improve translation quality in the multilingual setting for both high-resource and low-resource languages. We also demonstrate the effectiveness of our method over pre-training approaches for both NMT and cross-lingual transfer learning natural language understanding tasks."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Exploiting monolingual data for neural machine translation"]}]}],"canonical_facts":{"dc:contributor":["Zhai, ChengXiang","Hockenmaier, Julia","Ji, Heng","Awadalla, Hany Hassan"],"dc:creator":["Wang, Yiren"],"dc:date":["2023-12","2023-11-21"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-03-01 without embargo terms","The student, Yiren Wang, accepted the attached license on 2023-11-20 at 17:20.","The student, Yiren Wang, submitted this Dissertation for approval on 2023-11-20 at 17:27.","This Dissertation was approved for publication on 2023-11-21 at 12:11.","DSpace SAF Submission Ingestion Package generated from Vireo submission #19982 on 2024-03-01 at 13:14:51","Neural Machine Translation (NMT) has made rapid progress in recent years and achieved remarkable translation quality. NMT systems heavily rely on large-scale and high-quality parallel training data of source and target languages (bitext), which is limited and costly to collect. In the meantime, large amount of unlabeled monolingual data is available in different languages. While the target-side monolingual data has been proven useful through Back Translation, the source-side data is less investigated. Effectively exploiting both source-side and target-side monolingual data could not only help boost NMT performance, but also improve data efficiency. In this thesis, we focus on exploring novel approaches that effectively utilize monolingual data to improve the translation quality of NMT models. First, we exploit the potential of monolingual data with dual learning. Dual learning effectively utilize both bitext and monolingual data by leveraging the primal-dual structure of artificial intelligence tasks to generate informative feedback signals to regularize training. We extend the standard dual learning framework by introducing multiple primal and dual models, and propose the Multi-Agent Dual Learning (MADL) to further boost data utilization and translation quality. We show the effectiveness of our proposed approach on multiple benchmark NMT tasks, including both supervised and unsupervised NMT. Second, we study the effect of forward and back translation data at scale, and propose a large-scale noisy training pipeline to effectively leverage both of them. First, we generate synthetic bitext with both forward and back translation. Next, we train the NMT model on a noisy version of this synthetic corpus, where each source sentence is randomly corrupted. Finally, the model is fine-tuned on the genuine bitext and a clean, high-quality subset of the synthetic bitext without any noise. With our proposed strategy, we achieve the state-of-the-art performances on various benchmark translation tasks. Third, we explore more general strategies to utilize monolingual data for not only high-resource language pairs in bilingual NMT, but also low-resource language pairs and multilingual NMT as well. We propose a multi-task learning (MTL) framework, which jointly trains the model with the translation task on bitext data, and two auxiliary self-supervised learning tasks on the source-side and target-side monolingual data respectively. We show that our approach can effectively improve translation quality in the multilingual setting for both high-resource and low-resource languages. We also demonstrate the effectiveness of our method over pre-training approaches for both NMT and cross-lingual transfer learning natural language understanding tasks."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/122004"],"dc:language":["en","eng"],"dc:rights":["Copyright 2023 Yiren Wang"],"dc:subject":["Neural Machine Translation","Dual Learning","Multi-task Learning"],"dc:title":["Exploiting monolingual data for neural machine translation"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:00Z"}