{"id":{"repo_id":"trento","oai_identifier":"oai:iris.unitn.it:11572/368927"},"canonical_url":"https://search.dev.ndltd.org/etd/trento/oai:iris.unitn.it:11572/368927","repository":{"repo_id":"trento","name":"Università degli Studi di Trento","base_url":"https://iris.unitn.it/oai/request"},"display":{"title":"Learning Morphology for Open-Vocabulary Neural Machine Translation","abstract":"State-of-the-art neural machine translation systems typically have low accuracy in translating rare or unseen words due to the requirement of using a fixed-size word vocabulary during training. In addition to controlling the model complexity, this limitation is also related to the difficulty of learning accurate word representations under conditions of high data sparsity. This problem is an important bottleneck on performance, especially in morphologically-rich languages, where the word vocabulary tends to be huge and sparse. In this dissertation, we propose to solve the vocabulary limitation problem in neural machine translation by integrating morphology learning within the translation model, aiding to learn richer word representations in terms of phonological and morphological information. Our model improves the accuracy while translating into low-resource and morphologically-rich languages and shows better generalization capability over varieties of languages with different morphological characteristics.","abstract_html":"State-of-the-art neural machine translation systems typically have low accuracy in translating rare or unseen words due to the requirement of using a fixed-size word vocabulary during training. In addition to controlling the model complexity, this limitation is also related to the difficulty of learning accurate word representations under conditions of high data sparsity. This problem is an important bottleneck on performance, especially in morphologically-rich languages, where the word vocabulary tends to be huge and sparse. In this dissertation, we propose to solve the vocabulary limitation problem in neural machine translation by integrating morphology learning within the translation model, aiding to learn richer word representations in terms of phonological and morphological information. Our model improves the accuracy while translating into low-resource and morphologically-rich languages and shows better generalization capability over varieties of languages with different morphological characteristics.","abstract_has_math":false,"creators":["Ataman, Duygu"],"institution":"Università degli studi di Trento","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":["Federico, Marcello"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019","date_published":"2019","updated_at":"2026-07-24T05:04:19Z","subjects":["Settore INF/01 - Informatica","Settore ING-INF/05 - Sistemi di Elaborazione delle Informazioni"],"languages":["eng"],"rights":["info:eu-repo/semantics/openAccess","license:Tutti i diritti riservati (All rights reserved)","license uri:iris.PRI01"],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["http://dx.doi.org/10.15168/11572_368927","10.15168/11572_368927"],"render_values":[{"text":"http://dx.doi.org/10.15168/11572_368927","href":"http://dx.doi.org/10.15168/11572_368927","code":true},{"text":"10.15168/11572_368927","href":"https://doi.org/10.15168/11572_368927","code":true}]}]},"links":{"outbound_url":"https://hdl.handle.net/11572/368927","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Ataman, Duygu","Federico, Marcello"]},{"key":"dc:creator","label":"Author","values":["Ataman, Duygu"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2019"]},{"key":"dc:publisher","label":"Institution","values":["Università degli studi di Trento","place:TRENTO"]},{"key":"dc:relation","label":"Dc Relation","values":["firstpage:1","lastpage:164","numberofpages:164"]},{"key":"dc:type","label":"Dc Type","values":["info:eu-repo/semantics/doctoralThesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Settore INF/01 - Informatica","Settore ING-INF/05 - Sistemi di Elaborazione delle Informazioni"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["info:eu-repo/semantics/openAccess","license:Tutti i diritti riservati (All rights reserved)","license uri:iris.PRI01"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/11572/368927","http://dx.doi.org/10.15168/11572_368927","10.15168/11572_368927"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["State-of-the-art neural machine translation systems typically have low accuracy in translating rare or unseen words due to the requirement of using a fixed-size word vocabulary during training. In addition to controlling the model complexity, this limitation is also related to the difficulty of learning accurate word representations under conditions of high data sparsity. This problem is an important bottleneck on performance, especially in morphologically-rich languages, where the word vocabulary tends to be huge and sparse. In this dissertation, we propose to solve the vocabulary limitation problem in neural machine translation by integrating morphology learning within the translation model, aiding to learn richer word representations in terms of phonological and morphological information. Our model improves the accuracy while translating into low-resource and morphologically-rich languages and shows better generalization capability over varieties of languages with different morphological characteristics."]},{"key":"dc:title","label":"Title","values":["Learning Morphology for Open-Vocabulary Neural Machine Translation"]}]}],"canonical_facts":{"dc:contributor":["Ataman, Duygu","Federico, Marcello"],"dc:creator":["Ataman, Duygu"],"dc:date":["2019"],"dc:description":["State-of-the-art neural machine translation systems typically have low accuracy in translating rare or unseen words due to the requirement of using a fixed-size word vocabulary during training. In addition to controlling the model complexity, this limitation is also related to the difficulty of learning accurate word representations under conditions of high data sparsity. This problem is an important bottleneck on performance, especially in morphologically-rich languages, where the word vocabulary tends to be huge and sparse. In this dissertation, we propose to solve the vocabulary limitation problem in neural machine translation by integrating morphology learning within the translation model, aiding to learn richer word representations in terms of phonological and morphological information. Our model improves the accuracy while translating into low-resource and morphologically-rich languages and shows better generalization capability over varieties of languages with different morphological characteristics."],"dc:identifier":["https://hdl.handle.net/11572/368927","http://dx.doi.org/10.15168/11572_368927","10.15168/11572_368927"],"dc:language":["eng"],"dc:publisher":["Università degli studi di Trento","place:TRENTO"],"dc:relation":["firstpage:1","lastpage:164","numberofpages:164"],"dc:rights":["info:eu-repo/semantics/openAccess","license:Tutti i diritti riservati (All rights reserved)","license uri:iris.PRI01"],"dc:subject":["Settore INF/01 - Informatica","Settore ING-INF/05 - Sistemi di Elaborazione delle Informazioni"],"dc:title":["Learning Morphology for Open-Vocabulary Neural Machine Translation"],"dc:type":["info:eu-repo/semantics/doctoralThesis"]},"updated_at":"2026-07-24T05:04:19Z"}