Publikationsserver der RWTH Aachen University
Acoustic feature combination for speech recognition
Abstract
dc:descriptionIn this thesis, the use of multiple acoustic features of the speech signal is considered for speech recognition. The goals of this thesis are twofold: on the one hand, new acoustic features are developed, on the other hand, feature combination methods are investigated in order to find an effective integration of the newly developed features into state-of-the-art speech recognition systems. The most commonly used feature extraction methods are the Mel Frequency Cepstrum Coefficients (MFCC), Perceptual Linear Prediction (PLP), and variations of these techniques. These methods are mainly based on the models of the human auditory system. A detailed review of the implementation of these features is presented in this thesis. There have also been attempts at using articulatory motivated acoustic features for speech recognition which are motivated by models of the human speech production system. This thesis focuses partially on the development of new articulatory motivated acoustic features. The voicing information is one of the most commonly used articulatory features. Three voicing extraction methods are presented in this work followed by a systematic comparison. Besides the analysis of the voicing feature, the novel spectrum derivative feature is introduced which aims to capture the differences between magnitude spectra produced by obstruent and sonant consonants. The articulatory motivated features are tested in combinations with state-of-the-art acoustic features based on auditory models mainly. The features are combined both directly using Linear Discriminant Analysis (LDA) as well as indirectly on model level using Discriminative Model Combination (DMC). Both methods have already been used successfully in automatic speech recognition systems. In this work, a comparative study is presented which describes and analyzes the application of these methods to feature combination. Robustness issues of the LDA based method are addressed which are induced by increasing the amount of acoustic features coefficients. An application of DMC to feature combination is introduced based on the splitting of the acoustic model into separate scalable knowledge sources. After the analysis of the individual methods, a comparison is carried out on the basis of the underlying acoustic emission models. Experimental results are presented for small- and large-vocabulary tasks. The results show that the accuracy of automatic speech recognition systems can be significantly improved by the combination of auditory and articulatory motivated features. The combination of the Vocal Tract Length Normalized MFCC and articulatory motivated features demonstrates that additional articulatory information can even improve the performance of speaker adapted systems. The word error rate is reduced from 1.8% to 1.5% on the SieTill, a German digit string recognition task. Consistent improvements in word error rate have been obtained on two large-vocabulary corpora. The word error rate is reduced from 19.1% to 18.2% on the VerbMobil II, a German large vocabulary conversational speech task, and from 14.1% to 13.5% on the European Parliament Plenary Sessions task.
Degree
thesis:*- Grantor dc:publisher
- Publikationsserver der RWTH Aachen University
- Year dc:date
- 2006
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Zolnay, András
- Contributors dc:contributor
-
- Ney, Hermann
Subjects
dc:subject × 11Rights
dc:rights- Statement dc:rights
-
- info:eu-repo/semantics/openAccess
- Language dc:language
- eng
Identifiers
dc:identifier.*- OAI identifier oai:identifier
- oai:publications.rwth-aachen.de:61493