{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/101374"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/101374","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Multimodal machine translation","abstract":"Over the past few years, there has been a lot of progress being made in machine translation through deep learning networks. But there is relatively lesser progress made in using images to catalyze the translation tasks. In this study, we explore various models to incorporate the image features in the machine translation models. We start with a monomodal translation model which uses only textual features. We extend this model to develop the multimodal system which incorporates the visual features related to the source sentence. We also propose a multitask system which uses image captioning task to aid the translation task. Our models are tested on multiple datasets using the automatic evaluation metrics like METEOR and BLEU. The experiments show that the proposed models outperform the text-only baseline model.","abstract_html":"Over the past few years, there has been a lot of progress being made in machine translation through deep learning networks. But there is relatively lesser progress made in using images to catalyze the translation tasks. In this study, we explore various models to incorporate the image features in the machine translation models. We start with a monomodal translation model which uses only textual features. We extend this model to develop the multimodal system which incorporates the visual features related to the source sentence. We also propose a multitask system which uses image captioning task to aid the translation task. Our models are tested on multiple datasets using the automatic evaluation metrics like METEOR and BLEU. The experiments show that the proposed models outperform the text-only baseline model.","abstract_has_math":false,"creators":["Dave, Mihika"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Hockenmaier, Julia"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2018,"date_issued":"2018-09-04T20:47:30Z","date_published":"2018-09-04T20:47:30Z","updated_at":"2026-07-22T22:24:40Z","subjects":["multimodal machine translation","neural machine translation","multi-task learning","image captioning"],"languages":["en"],"rights":["Copyright 2018 Mihika Dave"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/101374","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Hockenmaier, Julia"]},{"key":"dc:creator","label":"Author","values":["Dave, Mihika"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2018-09-04T20:47:30Z","2020-09-05T09:15:16Z","2018-04-25","2018-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["multimodal machine translation","neural machine translation","multi-task learning","image captioning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2018 Mihika Dave"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/101374"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Over the past few years, there has been a lot of progress being made in machine translation through deep learning networks. But there is relatively lesser progress made in using images to catalyze the translation tasks. In this study, we explore various models to incorporate the image features in the machine translation models. We start with a monomodal translation model which uses only textual features. We extend this model to develop the multimodal system which incorporates the visual features related to the source sentence. We also propose a multitask system which uses image captioning task to aid the translation task. Our models are tested on multiple datasets using the automatic evaluation metrics like METEOR and BLEU. The experiments show that the proposed models outperform the text-only baseline model.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2020-05-01","The student, Mihika Dave, accepted the attached license on 2018-04-24 at 23:35.","The student, Mihika Dave, submitted this Thesis for approval on 2018-04-24 at 23:38.","This Thesis was approved for publication on 2018-04-25 at 15:11.","DSpace SAF Submission Ingestion Package generated from Vireo submission #12469 on 2018-08-31 at 17:30:24","Made available in DSpace on 2018-09-04T20:47:30Z (GMT). No. of bitstreams: 2 DAVE-THESIS-2018.pdf: 10520808 bytes, checksum: 896bcb1751935803be2f4385ca47e9e8 (MD5) LICENSE.txt: 4208 bytes, checksum: 545bfba9860cdc0ef1f83960ee1de8ca (MD5) Previous issue date: 2018-04-25","Embargo set by: Seth Robbins for item 107459 Lift date: 2020-09-04T20:47:38Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 107459 Lift date: 2020-09-04T20:50:11Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 107459 on 2020-09-05T09:15:16Z."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Multimodal machine translation"]}]}],"canonical_facts":{"dc:contributor":["Hockenmaier, Julia"],"dc:creator":["Dave, Mihika"],"dc:date":["2018-09-04T20:47:30Z","2020-09-05T09:15:16Z","2018-04-25","2018-05"],"dc:description":["Over the past few years, there has been a lot of progress being made in machine translation through deep learning networks. But there is relatively lesser progress made in using images to catalyze the translation tasks. In this study, we explore various models to incorporate the image features in the machine translation models. We start with a monomodal translation model which uses only textual features. We extend this model to develop the multimodal system which incorporates the visual features related to the source sentence. We also propose a multitask system which uses image captioning task to aid the translation task. Our models are tested on multiple datasets using the automatic evaluation metrics like METEOR and BLEU. The experiments show that the proposed models outperform the text-only baseline model.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2020-05-01","The student, Mihika Dave, accepted the attached license on 2018-04-24 at 23:35.","The student, Mihika Dave, submitted this Thesis for approval on 2018-04-24 at 23:38.","This Thesis was approved for publication on 2018-04-25 at 15:11.","DSpace SAF Submission Ingestion Package generated from Vireo submission #12469 on 2018-08-31 at 17:30:24","Made available in DSpace on 2018-09-04T20:47:30Z (GMT). No. of bitstreams: 2 DAVE-THESIS-2018.pdf: 10520808 bytes, checksum: 896bcb1751935803be2f4385ca47e9e8 (MD5) LICENSE.txt: 4208 bytes, checksum: 545bfba9860cdc0ef1f83960ee1de8ca (MD5) Previous issue date: 2018-04-25","Embargo set by: Seth Robbins for item 107459 Lift date: 2020-09-04T20:47:38Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 107459 Lift date: 2020-09-04T20:50:11Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 107459 on 2020-09-05T09:15:16Z."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/101374"],"dc:language":["en"],"dc:rights":["Copyright 2018 Mihika Dave"],"dc:subject":["multimodal machine translation","neural machine translation","multi-task learning","image captioning"],"dc:title":["Multimodal machine translation"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:40Z"}