{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/108055"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/108055","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"A translation framework for discovering word-like units from visual scenes and spoken descriptions","abstract":"In the absence of dictionaries, translators, or grammars, it is still possible to learn some of the words of a new language by listening to spoken descriptions of images. If several images, each containing a particular visually salient object, each co-occur with a particular sequence of speech sounds, we can infer that those speech sounds are a word whose definition is the visible object. A multimodal word discovery system accepts, as input, a database of spoken descriptions of images (or a set of corresponding phone transcriptions) and learns a mapping from waveform segments (or phone strings) to their associated image concepts. In this thesis, we propose a novel framework for multimodal word discovery systems based on statistical machine translation (SMT) and neural machine translation (NMT). We extend the existing theoretical frameworks on unsupervised word discovery and demonstrate a class of effective models for end-to-end word discovery from image regions and spoken descriptions. Finally, we provide a careful ablation study on components of my system and present some of the challenges in multimodal spoken word discovery.","abstract_html":"In the absence of dictionaries, translators, or grammars, it is still possible to learn some of the words of a new language by listening to spoken descriptions of images. If several images, each containing a particular visually salient object, each co-occur with a particular sequence of speech sounds, we can infer that those speech sounds are a word whose definition is the visible object. A multimodal word discovery system accepts, as input, a database of spoken descriptions of images (or a set of corresponding phone transcriptions) and learns a mapping from waveform segments (or phone strings) to their associated image concepts. In this thesis, we propose a novel framework for multimodal word discovery systems based on statistical machine translation (SMT) and neural machine translation (NMT). We extend the existing theoretical frameworks on unsupervised word discovery and demonstrate a class of effective models for end-to-end word discovery from image regions and spoken descriptions. Finally, we provide a careful ablation study on components of my system and present some of the challenges in multimodal spoken word discovery.","abstract_has_math":false,"creators":["Wang, Liming"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Hasegawa-Johnson, Mark A"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2020,"date_issued":"2020-08-26T21:58:07Z","date_published":"2020-08-26T21:58:07Z","updated_at":"2026-07-22T22:24:47Z","subjects":["machine translation","multimodal learning","low-resource speech technology"],"languages":["en"],"rights":["Copyright 2020 Liming Wang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/108055","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Hasegawa-Johnson, Mark A"]},{"key":"dc:creator","label":"Author","values":["Wang, Liming"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2020-08-26T21:58:07Z","2020-05-14","2020-05"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["machine translation","multimodal learning","low-resource speech technology"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2020 Liming Wang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/108055"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["In the absence of dictionaries, translators, or grammars, it is still possible to learn some of the words of a new language by listening to spoken descriptions of images. If several images, each containing a particular visually salient object, each co-occur with a particular sequence of speech sounds, we can infer that those speech sounds are a word whose definition is the visible object. A multimodal word discovery system accepts, as input, a database of spoken descriptions of images (or a set of corresponding phone transcriptions) and learns a mapping from waveform segments (or phone strings) to their associated image concepts. In this thesis, we propose a novel framework for multimodal word discovery systems based on statistical machine translation (SMT) and neural machine translation (NMT). We extend the existing theoretical frameworks on unsupervised word discovery and demonstrate a class of effective models for end-to-end word discovery from image regions and spoken descriptions. Finally, we provide a careful ablation study on components of my system and present some of the challenges in multimodal spoken word discovery.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2020-08-25 without embargo terms","The student, Liming Wang, accepted the attached license on 2020-05-13 at 10:26.","The student, Liming Wang, submitted this Thesis for approval on 2020-05-13 at 10:28.","This Thesis was approved for publication on 2020-05-14 at 09:24.","DSpace SAF Submission Ingestion Package generated from Vireo submission #15375 on 2020-08-25 at 17:14:33","Made available in DSpace on 2020-08-26T21:58:07Z (GMT). No. of bitstreams: 2 WANG-THESIS-2020.pdf: 1955236 bytes, checksum: d466468eb8c76cf7cef7185d035e1cb1 (MD5) LICENSE.txt: 4208 bytes, checksum: 404bb44630b880d96166085d5cde918c (MD5) Previous issue date: 2020-05-14"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["A translation framework for discovering word-like units from visual scenes and spoken descriptions"]}]}],"canonical_facts":{"dc:contributor":["Hasegawa-Johnson, Mark A"],"dc:creator":["Wang, Liming"],"dc:date":["2020-08-26T21:58:07Z","2020-05-14","2020-05"],"dc:description":["In the absence of dictionaries, translators, or grammars, it is still possible to learn some of the words of a new language by listening to spoken descriptions of images. If several images, each containing a particular visually salient object, each co-occur with a particular sequence of speech sounds, we can infer that those speech sounds are a word whose definition is the visible object. A multimodal word discovery system accepts, as input, a database of spoken descriptions of images (or a set of corresponding phone transcriptions) and learns a mapping from waveform segments (or phone strings) to their associated image concepts. In this thesis, we propose a novel framework for multimodal word discovery systems based on statistical machine translation (SMT) and neural machine translation (NMT). We extend the existing theoretical frameworks on unsupervised word discovery and demonstrate a class of effective models for end-to-end word discovery from image regions and spoken descriptions. Finally, we provide a careful ablation study on components of my system and present some of the challenges in multimodal spoken word discovery.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2020-08-25 without embargo terms","The student, Liming Wang, accepted the attached license on 2020-05-13 at 10:26.","The student, Liming Wang, submitted this Thesis for approval on 2020-05-13 at 10:28.","This Thesis was approved for publication on 2020-05-14 at 09:24.","DSpace SAF Submission Ingestion Package generated from Vireo submission #15375 on 2020-08-25 at 17:14:33","Made available in DSpace on 2020-08-26T21:58:07Z (GMT). No. of bitstreams: 2 WANG-THESIS-2020.pdf: 1955236 bytes, checksum: d466468eb8c76cf7cef7185d035e1cb1 (MD5) LICENSE.txt: 4208 bytes, checksum: 404bb44630b880d96166085d5cde918c (MD5) Previous issue date: 2020-05-14"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/108055"],"dc:language":["en"],"dc:rights":["Copyright 2020 Liming Wang"],"dc:subject":["machine translation","multimodal learning","low-resource speech technology"],"dc:title":["A translation framework for discovering word-like units from visual scenes and spoken descriptions"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:47Z"}