Robert Gordon University
Self-supervision strategies for post-ASR error correction in clinical domain.
Abstract
dc:description.abstractIn clinical settings, manual documentation of patient interactions, medical histories, and treatment plans poses significant challenges, including clinician burnout, information loss, and inefficiency. Automatic Speech Recognition (ASR) systems offer a promising solution by transcribing spoken language into written text, enabling clinicians to dictate notes in real-time. However, using ASR systems in medical environments remains challenging due to the need for high accuracy in a critical domain filled with specialised terminology. Despite advancements in ASR applications, there is a gap in the literature addressing the correction of transcription errors in scenarios with limited domain-specific data. Current models heavily rely on extensive training datasets, which may not always be available in specialised fields such as healthcare. This research identifies two key challenges in post-ASR error correction: the vocabulary gap, where pre-trained language models struggle with domain-specific terms, and the objective gap, where general-purpose models are not optimised for specific tasks such as error correction. To address these challenges, this study employs transfer learning to fine-tune pre-trained large language models for post-ASR error correction, even in low-data environments. By leveraging both internal and open-source ASR datasets, the main types of errors occurring in these datasets were identified. The study also introduces self-supervision strategies and an edit distance operations-based noising algorithm to optimise error correction processes based on error distributions. The contributions of this research include the creation of a domain-specific ASR dataset alongside a comprehensive error analysis, resulting in transcription-reference pairs that are publicly available for future work. Next, a fine-tuning approach using pre-trained large language models was developed to address the objective gap, revealing that tasks similar to mask-filling outperform other sequence-to-sequence tasks. The self-supervision strategies underscored the importance of selecting the appropriate task and bridging the vocabulary gap by employing medical mask-filling techniques. As a result, the specialised mask-filling objective for the medical domain outperformed the best-performing commercial ASR system by 10.27%. Finally, understanding that different types of errors affect the model's performance, the study looked at predefined noise patterns and analysed error distributions to see how various error levels and types influenced the results. This led to the development of an edit distance operations-based noise strategy, where synthetic noise turned out to be more effective at correcting errors than the real-world noise typically found in datasets.
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Nanayakkara, Gayani
- Advisor dc:contributor.advisor
-
- N. Wiratunga, D. Corsar and K. Martin
Subjects
dc:subject × 10Rights
- Language dc:language
- en
Identifiers
dc:identifier.*- Identifier
-
oai:rgu-repository.worktribe.com:3020496
https://doi.org/10.48526/rgu-wt-3020496 - Author Identifier
- 0000-0002-0017-6589
- OAI identifier oai:identifier
- oai:rgu-repository.worktribe.com:3020496