Back to results

Massachusetts Institute of Technology

Incremental speech understanding in a multimodal web-based spoken dialogue system

Abstract

dc:description.abstract

In most spoken dialogue systems, the human speaker interacting with the system must wait until after finishing speaking to find out whether his or her speech has been accurately understood. The verbal and nonverbal indicators of understanding typical in human-to-human interaction are generally nonexistent in automated systems, resulting in an interaction that feels unnatural to the human user. However, as automatic speech recognition gets incorporated into web-based and portable interfaces, there are now graphical means in addition to the verbal means by which a spoken dialogue system can communicate to the user. In this thesis, we present a multimodal web-based spoken dialogue system that incorporates incremental understanding of human speech. Through incremental understanding, the system can display to the user its current understanding of specific concepts in real-time while the user is still in the process of uttering a sentence. In addition, the user can interact with the system through nonverbal input modalities such as typing and mouse clicking. We evaluate the results of a comparative user study in which one group uses a configuration that receives incremental concept understanding, while another group uses a configuration that lacks this feature. We found that the group receiving incremental updates had a greater task completion rate and overall user satisfaction.

Degree

thesis:*
Department dc:contributor.department
Massachusetts Institute of Technology. Dept. of Electrical Engineering and Computer Science.
Grantor dc:publisher
Massachusetts Institute of Technology
Year dc:date.issued
2009

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Matthias Gary M. (Gary Michael)
Advisor dc:contributor.advisor
  • James R. Glass and Stephanie Seneff.

Subjects

dc:subject × 1

Rights

dc:rights
Statement dc:rights
  • M.I.T. theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. See provided URL for inquiries about permission.
Language dc:language.iso
eng

Identifiers

dc:identifier.*
Handle dc:identifier.uri
http://hdl.handle.net/1721.1/53170
OAI identifier oai:identifier
oai:dspace.mit.edu:1721.1/53170

Chain of custody

source
Harvested from
MIT
Base URL
dspace.mit.edu/oai/request
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Matthias Gary M. (Gary Michael). Incremental speech understanding in a multimodal web-based spoken dialogue system. Massachusetts Institute of Technology, 2009. http://hdl.handle.net/1721.1/53170