{"id":{"repo_id":"iu","oai_identifier":"oai:scholarworks.iu.edu:2022/22672"},"canonical_url":"https://search.dev.ndltd.org/etd/iu/oai:scholarworks.iu.edu:2022/22672","repository":{"repo_id":"iu","name":"Indiana University","base_url":"https://scholarworks.iu.edu/iuswrrest/oai/request"},"display":{"title":"Cross-Lingual Word Sense Disambiguation for Low-Resource Hybrid Machine Translation","abstract":"This thesis argues that cross-lingual word sense disambiguation (CL-WSD) can be used to improve lexical selection for machine translation when translating from a resource- rich language into an under-resourced one, especially when relatively little bitext is avail- able. In CL-WSD, we perform word sense disambiguation, considering the senses of a word to be its possible translations into some target language, rather than using a sense inventory developed manually by lexicographers. Using explicitly trained classifiers that make use of source-language context and of resources for the source language can help machine translation systems make better decisions when selecting target-language words. This is especially the case when the alternative is hand-written lexical selection rules developed by researchers with linguistic knowledge of the source and target languages, but also true when lexical selection would be performed by a statistical machine translation system, when there is a relatively small amount of available target-language text for training language models. In this work, I present the Chipa system for CL-WSD and apply it to the task of translating from Spanish to Guarani and Quechua, two indigenous languages of South America. I demonstrate several extensions to the basic Chipa system, including tech- niques that allow us to benefit from the wealth of available unannotated Spanish text and existing text analysis tools for Spanish, as well as approaches for learning from bitext resources that pair Spanish with languages unrelated to our intended target lan- guages. Finally, I provide proof-of-concept integrations of Chipa with existing machine translation systems, of two completely different architectures.","abstract_html":"This thesis argues that cross-lingual word sense disambiguation (CL-WSD) can be used to improve lexical selection for machine translation when translating from a resource- rich language into an under-resourced one, especially when relatively little bitext is avail- able. In CL-WSD, we perform word sense disambiguation, considering the senses of a word to be its possible translations into some target language, rather than using a sense inventory developed manually by lexicographers. Using explicitly trained classifiers that make use of source-language context and of resources for the source language can help machine translation systems make better decisions when selecting target-language words. This is especially the case when the alternative is hand-written lexical selection rules developed by researchers with linguistic knowledge of the source and target languages, but also true when lexical selection would be performed by a statistical machine translation system, when there is a relatively small amount of available target-language text for training language models. In this work, I present the Chipa system for CL-WSD and apply it to the task of translating from Spanish to Guarani and Quechua, two indigenous languages of South America. I demonstrate several extensions to the basic Chipa system, including tech- niques that allow us to benefit from the wealth of available unannotated Spanish text and existing text analysis tools for Spanish, as well as approaches for learning from bitext resources that pair Spanish with languages unrelated to our intended target lan- guages. Finally, I provide proof-of-concept integrations of Chipa with existing machine translation systems, of two completely different architectures.","abstract_has_math":false,"creators":["Rudnick, Alexander James"],"institution":"[Bloomington, Ind.] : Indiana University","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019-01","date_published":"2019-01","updated_at":"2026-07-24T02:40:07Z","subjects":["machine translation","artificial intelligence","computational linguistics"],"languages":["en"],"rights":["Creative Commons Attribution 4.0 International"],"rights_urls":["https://creativecommons.org/licenses/by/4.0/"],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2022/22672","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["Rudnick, Alexander James"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2019-01-22T17:50:32Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2019-01-22T17:50:32Z"]},{"key":"dc:date.issued","label":"Date","values":["2019-01"]},{"key":"dc:publisher","label":"Institution","values":["[Bloomington, Ind.] : Indiana University"]},{"key":"dc:type","label":"Dc Type","values":["Doctoral Dissertation"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["machine translation","artificial intelligence","computational linguistics"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Creative Commons Attribution 4.0 International"]},{"key":"dc:rights.uri","label":"Rights URI","values":["https://creativecommons.org/licenses/by/4.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/2022/22672"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Thesis (Ph.D.) - Indiana University, School of Informatics, Computing, and Engineering, 2019"]},{"key":"dc:description.abstract","label":"Abstract","values":["This thesis argues that cross-lingual word sense disambiguation (CL-WSD) can be used to improve lexical selection for machine translation when translating from a resource- rich language into an under-resourced one, especially when relatively little bitext is avail- able. In CL-WSD, we perform word sense disambiguation, considering the senses of a word to be its possible translations into some target language, rather than using a sense inventory developed manually by lexicographers. Using explicitly trained classifiers that make use of source-language context and of resources for the source language can help machine translation systems make better decisions when selecting target-language words. This is especially the case when the alternative is hand-written lexical selection rules developed by researchers with linguistic knowledge of the source and target languages, but also true when lexical selection would be performed by a statistical machine translation system, when there is a relatively small amount of available target-language text for training language models. In this work, I present the Chipa system for CL-WSD and apply it to the task of translating from Spanish to Guarani and Quechua, two indigenous languages of South America. I demonstrate several extensions to the basic Chipa system, including tech- niques that allow us to benefit from the wealth of available unannotated Spanish text and existing text analysis tools for Spanish, as well as approaches for learning from bitext resources that pair Spanish with languages unrelated to our intended target lan- guages. Finally, I provide proof-of-concept integrations of Chipa with existing machine translation systems, of two completely different architectures."]},{"key":"dc:title","label":"Title","values":["Cross-Lingual Word Sense Disambiguation for Low-Resource Hybrid Machine Translation"]}]}],"canonical_facts":{"dc:creator":["Rudnick, Alexander James"],"dc:date.accessioned":["2019-01-22T17:50:32Z"],"dc:date.available":["2019-01-22T17:50:32Z"],"dc:date.issued":["2019-01"],"dc:description":["Thesis (Ph.D.) - Indiana University, School of Informatics, Computing, and Engineering, 2019"],"dc:description.abstract":["This thesis argues that cross-lingual word sense disambiguation (CL-WSD) can be used to improve lexical selection for machine translation when translating from a resource- rich language into an under-resourced one, especially when relatively little bitext is avail- able. In CL-WSD, we perform word sense disambiguation, considering the senses of a word to be its possible translations into some target language, rather than using a sense inventory developed manually by lexicographers. Using explicitly trained classifiers that make use of source-language context and of resources for the source language can help machine translation systems make better decisions when selecting target-language words. This is especially the case when the alternative is hand-written lexical selection rules developed by researchers with linguistic knowledge of the source and target languages, but also true when lexical selection would be performed by a statistical machine translation system, when there is a relatively small amount of available target-language text for training language models. In this work, I present the Chipa system for CL-WSD and apply it to the task of translating from Spanish to Guarani and Quechua, two indigenous languages of South America. I demonstrate several extensions to the basic Chipa system, including tech- niques that allow us to benefit from the wealth of available unannotated Spanish text and existing text analysis tools for Spanish, as well as approaches for learning from bitext resources that pair Spanish with languages unrelated to our intended target lan- guages. Finally, I provide proof-of-concept integrations of Chipa with existing machine translation systems, of two completely different architectures."],"dc:identifier.uri":["https://hdl.handle.net/2022/22672"],"dc:language.iso":["en"],"dc:publisher":["[Bloomington, Ind.] : Indiana University"],"dc:rights":["Creative Commons Attribution 4.0 International"],"dc:rights.uri":["https://creativecommons.org/licenses/by/4.0/"],"dc:subject":["machine translation","artificial intelligence","computational linguistics"],"dc:title":["Cross-Lingual Word Sense Disambiguation for Low-Resource Hybrid Machine Translation"],"dc:type":["Doctoral Dissertation"]},"updated_at":"2026-07-24T02:40:07Z"}