Abstract
dc:description.abstractInstrumental speech quality prediction is a long-studied field in which many models have been presented. However, in particular, the single-ended prediction without the use of a clean reference signal remains challenging. This thesis studies how recent developments in machine learning can be leveraged to improve the quality prediction of transmitted speech and additionally provide diagnostic information through the prediction of speech quality dimensions. In particular, different deep learning architectures were analysed towards their suitability to predict speech quality. To this end, a large dataset with distorted speech files and crowdsourced subjective ratings was created. A number of deep learning architectures, such as CNNs, LSTM networks, and Transformer/self-attention networks were combined and compared. It was found that a network with CNN, Self-Attention, and a proposed attention-pooling delivers the best single-ended speech quality predictions on the considered dataset. Furthermore, a double-ended speech quality prediction model based on a Siamese neural network is presented. However, it could be shown that, in contrast to traditional models, deep learning models only slightly benefit from including the clean reference signal. For the prediction of perceptual speech quality dimensions, a multi-task learning based model is presented that predicts the overall speech quality and the quality dimensions Noisiness, Coloration, Discontinuity, and Loudness in parallel, where most of the neural network layers are shared between the individual tasks. Finally, the single-ended speech quality prediction model NISQA is presented that was trained on a large variety of 59 different datasets. Because the training datasets come from a variety of sources and contain different quality ranges, they are exposed to subjective biases. Therefore, the same speech distortions can lead to very different quality ratings in two datasets. To prevent a negative influence of this effect, a bias-aware loss function is proposed that estimates and considers the biases during the training of the neural network weights. The final model was tested on a live-talking test set with real recorded phone calls, on which it achieved a Pearson's correlation of 0.90 for the overall speech quality prediction.
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Mittag, Gabriel
- Advisor dc:contributor.advisor
-
- Möller, Sebastian
Rights
- Licence dc:rights.uri
- Language dc:language.iso
- en
Identifiers
dc:identifier.*- Identifier URI
- http://dx.doi.org/10.14279/depositonce-11873
- OAI identifier oai:identifier
- oai:depositonce.tu-berlin.de:11303/13077