National University of Singapore
Noise-Robust Speech Recognition Using Deep Neural Network
Abstract
dc:description.abstractThis thesis addresses the noise robustness of the recently developed Deep Neural Networks (DNNs) based speech recognition systems. Five techniques have been proposed. Firstly, a Mean Variance Normalization technique was developed to integrate noise statistics using Vector Taylor Series. Secondly, a Deep Split Temporal Context system was proposed to separately model sub-contexts of a long temporal span of speech. Thirdly, we revisited the missing feature theory and developed a DNN-based spectral masking system. Fourthly, an Ideal Hidden-activation Mask (IHM) was proposed to remove noise-prone latent detectors. Lastly, a noise code technique was developed to simulate IHM with reduced computational costs. Improved noise robustness has been obtained using the proposed techniques on benchmark tasks, Aurora-2 and Aurora-4. The spectral masking approach successfully ranks first in the literature on both tasks at the time of writing and is one of the most promising noise-robust techniques for DNNs.
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- LI BO