Back to results

University of Illinois at Urbana-Champaign

A translation framework for discovering word-like units from visual scenes and spoken descriptions

Abstract

dc:description

In the absence of dictionaries, translators, or grammars, it is still possible to learn some of the words of a new language by listening to spoken descriptions of images. If several images, each containing a particular visually salient object, each co-occur with a particular sequence of speech sounds, we can infer that those speech sounds are a word whose definition is the visible object. A multimodal word discovery system accepts, as input, a database of spoken descriptions of images (or a set of corresponding phone transcriptions) and learns a mapping from waveform segments (or phone strings) to their associated image concepts. In this thesis, we propose a novel framework for multimodal word discovery systems based on statistical machine translation (SMT) and neural machine translation (NMT). We extend the existing theoretical frameworks on unsupervised word discovery and demonstrate a class of effective models for end-to-end word discovery from image regions and spoken descriptions. Finally, we provide a careful ablation study on components of my system and present some of the challenges in multimodal spoken word discovery.

Degree

thesis:*
Name thesis:degree_name
M.S.
Level thesis:degree_level
Thesis
Discipline thesis:degree_discipline
Electrical & Computer Engr
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2020

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Wang, Liming
Contributors dc:contributor
  • Hasegawa-Johnson, Mark A

Subjects

dc:subject × 3

Rights

dc:rights
Statement dc:rights
  • Copyright 2020 Liming Wang
Language dc:language
en

Identifiers

dc:identifier.*
Handle dc:identifier
http://hdl.handle.net/2142/108055
OAI identifier oai:identifier
oai:www.ideals.illinois.edu:2142/108055

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Wang, Liming. A translation framework for discovering word-like units from visual scenes and spoken descriptions. Thesis thesis, University of Illinois at Urbana-Champaign, 2020. http://hdl.handle.net/2142/108055