University of Cambridge
Speech-Based Emotion Modelling and Mental Disorder Detection
Abstract
dc:description.abstractEmotion modelling and understanding are crucial for artificial intelligence (AI) systems to achieve enhanced contextual understanding and adaptive, personalised human-AI interaction. Speech contains important clues for detecting emotion through a variety of vocal characteristics such as prosody, along with speech patterns such as hesitation and laughter. This thesis first explores automatic emotion recognition (AER) from speech input. Current AER systems face two primary challenges: (i) the mismatch between research experiments and practical applications such as the use of reference transcriptions and sentence segmentation; (ii) the inconsistency of emotion annotations due to ambiguous expressions and subjective perception. To tackle the first challenge, an integrated system is developed which integrates AER with speaker diarisation and speech recognition in a jointly-trained system. Compared to separately optimised cascaded systems, the proposed system achieves not only improved efficiency but also reduced recognition errors for emotional speech. In addition, two novel metrics are introduced to evaluate AER performance with automatic segmentation based on time-weighted emotion classification errors. In response to the second AER challenge, it is proposed to represent emotion as a distribution rather than a single class. Different emotion annotations provided by human annotators are treated as samples drawn from the emotion distribution. Evidential deep learning (EDL) is used to quantify the uncertainty in emotion distribution estimation by learning an utterance-specific prior distribution. Representing emotion as a distribution offers not only a more comprehensive representation of emotional content but also an inclusive representation of human opinions. The challenge of inconsistent human opinions extends beyond emotion annotation and affects various subjective tasks such as speech quality assessment and toxic speech detection. A general framework for human annotator simulation is introduced, which accounts for the variability in human judgements. The framework meta-learns a conditional flow model, which demonstrates superior capability and efficiency in predicting the aggregated behaviour of human annotators, matching the distribution of human annotations, and simulating inter-annotator disagreements. It is hoped that the proposed methods could contribute to the promotion of inclusivity and fairness in ethical AI practices. Furthermore, emotion is closely linked with mental wellbeing. A speech-based automatic depression detection system is introduced which uses foundation models pretrained on large speech datasets to alleviate the data sparsity issue of medical datasets. It is shown that incorporating emotion information is useful for depression detection. Integrating representations from multiple foundation models achieves state-of-the-art results without requiring oracle transcriptions. To enhance the reliability of automatic diagnosis systems, confidence estimation methods are studied. The proposed method builds upon the EDL approach introduced previously for emotion distribution estimation, adapting it to learn the predictive distribution of mental illness detection. This method aims to foster reliable and trustworthy automatic diagnostic systems.
Degree
thesis:*- Name dc:type.qualificationname
- Doctor of Philosophy (PhD)
- Level dc:type.qualificationlevel
- Doctoral
- Grantor dc:publisher.institution
- University of Cambridge
- Year dc:date.issued
- 2024
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Wu, Wen
- Advisor dc:contributor.advisor
-
- Woodland, Phil
Subjects
dc:subject × 6Rights
dc:rightsIdentifiers
dc:identifier.*- DOI dc:identifier.doi
- https://doi.org/10.17863/CAM.112567
- OAI identifier oai:identifier
- oai:www.repository.cam.ac.uk:1810/374527