Back to results

University of Illinois at Urbana-Champaign

Generative Models for Retrieval of Video, Audio and Text Data

Abstract

dc:description

"We propose a general approach to audio segment retrieval, which allows a user to query audio data by an example audio segment of a short duration and to find similar segments. The basic idea of our approach is to first train a hidden Markov model (HMM) using the given example, it is called the theme HMM. The total audio data available is used to train a background HMM. We combine these individual HMMs to form a synthesized ""background-theme-background"" HMM. This synthesized HMM can then be applied to any audio stream as a parser to detect the most likely theme segment. A major advantage of this approach is that it does not assume any predefined segment boundaries as in previous work and thus can be expected to retrieve theme segments with more accurate boundaries. We overcome the problem of only being able to use a short duration query to train a theme HMM by using the maximum a posteriori estimator with the background model as a prior model. Evaluation of the proposed retrieval scheme using short duration example audio clips of narration as queries gives quite promising results."

Degree

thesis:*
Name thesis:degree_name
Ph.D.
Level thesis:degree_level
Dissertation
Discipline thesis:degree_discipline
Electrical and Computer Engineering
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2015

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Velivelli, Atulya
Contributors dc:contributor
  • Huang, Thomas S.

Subjects

dc:subject × 1

Rights

Language dc:language
eng

Identifiers

dc:identifier.*
Identifier
(MiAaPQ)AAI3425399
OAI identifier oai:identifier
oai:www.ideals.illinois.edu:2142/81157

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Velivelli, Atulya. Generative Models for Retrieval of Video, Audio and Text Data. Dissertation thesis, University of Illinois at Urbana-Champaign, 2015. http://hdl.handle.net/2142/81157