Back to results

Georgia Institute of Technology

Estimation of glottal source features from the spectral envelope of the acoustic speech signal

Abstract

dc:description.abstract

Speech communication encompasses diverse types of information, including phonetics, affective state, voice quality, and speaker identity. From a speech production standpoint, the acoustic speech signal can be mainly divided into glottal source and vocal tract components, which play distinct roles in rendering the various types of information it contains. Most deployed speech analysis systems, however, do not explicitly represent these two components as distinct entities, as their joint estimation from the acoustic speech signal becomes an ill-defined blind deconvolution problem. Nevertheless, because of the desire to understand glottal behavior and how it relates to perceived voice quality, there has been continued interest in explicitly estimating the glottal component of the speech signal. To this end, several inverse filtering (IF) algorithms have been proposed, but they are unreliable in practice because of the blind formulation of the separation problem. In an effort to develop a method that can bypass the challenging IF process, this thesis proposes a new glottal source information extraction method that relies on supervised machine learning to transform smoothed spectral representations of speech, which are already used in some of the most widely deployed and successful speech analysis applications, into a set of glottal source features. A transformation method based on Gaussian mixture regression (GMR) is presented and compared to current IF methods in terms of feature similarity, reliability, and speaker discrimination capability on a large speech corpus, and potential representations of the spectral envelope of speech are investigated for their ability represent glottal source variation in a predictable manner. The proposed system was found to produce glottal source features that reasonably matched their IF counterparts in many cases, while being less susceptible to spurious errors. The development of the proposed method entailed a study into the aspects of glottal source information that are already contained within the spectral features commonly used in speech analysis, yielding an objective assessment regarding the expected advantages of explicitly using glottal information extracted from the speech signal via currently available IF methods, versus the alternative of relying on the glottal source information that is implicitly contained in spectral envelope representations.

Degree

thesis:*
Department dc:contributor.department
Electrical and Computer Engineering
Grantor dc:publisher
Georgia Institute of Technology
Year dc:date.issued
2010

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Torres, Juan Félix
Advisor dc:contributor.advisor
  • Moore, Elliot
Committee members dc:contributor.committeemember
  • Haas, Kevin
  • Hayes, Monson
  • Lee, Chin-Hui
  • Wu, Hongwei

Subjects

dc:subject × 4

Identifiers

dc:identifier.*
Handle dc:identifier.uri
http://hdl.handle.net/1853/34736
OAI identifier oai:identifier
oai:repository.gatech.edu:1853/34736

Chain of custody

source
Harvested from
Georgia Tech
Base URL
repository.gatech.edu/server/oai/request
Last updated
2026-07-27
Source record
OAI-PMH GetRecord
citation

Torres, Juan Félix. Estimation of glottal source features from the spectral envelope of the acoustic speech signal. Georgia Institute of Technology, 2010. http://hdl.handle.net/1853/34736