Publikationsserver der RWTH Aachen University
Normalization in the acoustic feature space for improved speech recognition
Abstract
dc:descriptionIn this work, normalization techniques in the acoustic feature space are studied which improve the robustness of automatic speech recognition systems. It is shown that there is a fundamental mismatch between training and test data which causes degraded recognition performance. Adaptation and normalization, basic strategies to reduce the mismatch, are introduced and placed into the framework of statistical speech recognition. A classification scheme for different normalization techniques is introduced. Common normalization schemes proposed in the literature are motivated and discussed, and two promising techniques are implemented and studied in detail. Vocal tract length normalization relies on frequency axis warping during signal analysis to reduce inter-speaker variability. The baseline procedure for training and test data normalization is introduced and optimized so that consistently large improvements in recognition performance are achieved under a variety of acoustic conditions. A technique for fast parameter estimation is developed that gives the same improvements as the baseline technique without an increase in computation time. It is shown that vocal tract length normalization can be applied successfully in online applications. A novel approach for integrated frequency axis warping is developed that merges successive signal analysis steps into a single one. It simplifies signal analysis and gives a better control over the amount of spectral smoothing. The second set of techniques explored in detail are histogram normalization and feature space rotation. They aim at reducing the mismatch between training and test by matching the distributions of the training and test data. The effect of histogram normalization at different signal analysis stages, as well as training and test data normalization are investigated in detail. One of the basic assumptions of histogram normalization is relaxed by taking care of the variable silence fraction. Feature space rotation is introduced to account for undesired variations in the speech signal that are correlated in the feature space dimensions. The interaction of histogram normalization and feature space rotation is analyzed, and it is shown that both techniques significantly improve the recognition accuracy in scenarios with different degrees of mismatch. Finally, it is demonstrated how the application of several normalization schemes in presence of large mismatch between training and test data can make the difference from essentially zero recognition accuracy to a high level of 90%. Experimental results are reported for corpora with different acoustic conditions, vocabulary sizes, languages, and speaking styles: North American Business News is a large vocabulary task of English read speech, VerbMobil II is a German large vocabulary conversational speech task, EuTrans II is an Italian speech corpus of conversational speech over telephone, and CarNavigation a German isolated-word recognition task recorded partly in noisy car environments.
Degree
thesis:*- Grantor dc:publisher
- Publikationsserver der RWTH Aachen University
- Year dc:date
- 2003
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Molau, Sirko
- Contributors dc:contributor
-
- Ney, Hermann
Subjects
dc:subject × 9Rights
dc:rights- Statement dc:rights
-
- info:eu-repo/semantics/openAccess
- Language dc:language
- eng