Back to search

Queen's University Belfast

Speech enhancement for real-time applications

Abstract

dc:description.abstract

Segmental exemplar-based speech enhancement algorithms are promising for real-world applications as they can function well without noise training, meaning they can easily generalize to real-world applications. They are, however, typically slow due to needing to search through a corpus that is sufficiently large to be representative. Another key issue is the large size of a representative corpus. This thesis explores two adaptations necessary to meet the requirements of real-time enhancement for lower-powered devices by resolving these two issues. <br/><br/>Firstly, in Chapter 3, the nature of speech is exploited to impose a hierarchical structure on the clean speech corpus, to facilitate a tree-based search of the corpus. A baseline test system is used to demonstrate the concept, which matches fixed length test segments with clean speech corpus segments. These are then used to synthesize output speech. This approach is evaluated and compared with a linear search of the corpus - finding that the algorithm can now function 20x faster than real-time. This is due to the search space for a corpus of n segments being dramatically reduced - from O(n) to O(log(n)). <br/><br/>Secondly, in Chapter 4 we consider the second key issue. The size of the corpus is too large for many low-powered devices. Clustering can be used to obtain a lossy compression of the speech corpus by replacing original segments with codewords. Several different means of clustering for this application are evaluated. It is shown that this results in a corpus a tenth of the size of the original corpus while maintaining the representiveness of the corpus. How compression and tree-based search function together is explored. A technique aimed at reducing the loss in quality when using both techniques together is also explored. <br/><br/>A third study, in Chapter 5, attempts to make use of the lingua-acoustic data present in a sentence to improve the accuracy of automatic speech recognition (ASR) rather than the perceived speech quality. The basic approach leads to a improvement of up to 40% in speech recognition accuracy. Application of the long-term lingua-acoustic data yielded a small but noticeable improvement of around 2% overall. Some of the lingua-acoustic constraints used to exploit this data contributed to the accuracy improvement, while some did not contribute at all. These results are obtained while functioning under real-time on a non-streaming basis. This suggests that the system could be useful for an ASR application that tolerates a small amount of latency. <br/><br/>Finally, future directions for the work are identified.

Degree

thesis:*
Name dc:type.qualificationname
Doctor of Philosophy
Level dc:type.qualificationlevel
Doctoral Thesis
Grantor dc:publisher.institution
Queen's University Belfast
Year dc:date.issued
2020

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Nesbitt, David
Advisors dc:contributor.advisor
  • Ji, Ming
  • McLaughlin, Niall

Subjects

dc:subject × 4

Rights

Language dc:language
eng

Identifiers

dc:identifier.*
Identifier
oai:pure.qub.ac.uk/portal:studenttheses/ce499bf6-6a19-4e07-aca3-395fffa6f1b8
OAI identifier oai:identifier
oai:pure.qub.ac.uk/portal:studenttheses/ce499bf6-6a19-4e07-aca3-395fffa6f1b8

Chain of custody

source
Harvested from
Queen's University Belfast
Base URL
pureadmin.qub.ac.uk/ws/oai
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Nesbitt, David. Speech enhancement for real-time applications. Doctoral Thesis thesis, Queen's University Belfast, 2020. https://pure.qub.ac.uk/en/studentTheses/ce499bf6-6a19-4e07-aca3-395fffa6f1b8