Back to results

University of Cambridge

The Use of Language Models in End-to-End Spoken Language Processing

Abstract

dc:description.abstract

By leveraging general deep learning techniques, the end-to-end (E2E) trainable modelling approach has achieved great success in spoken language processing including automatic speech recognition (ASR) and speech translation (ST). E2E spoken language processing systems have many advantages compared to cascaded systems, including global optimisation, a simplified pipeline, and a compact structure. However, E2E training further reduces interpretability and often lacks a modular structure, such as an explicit language model (LM) component, which could enable more efficient use of easily-collected data, such as text-only data. LMs are designed to capture dependencies between modelling units, making them particularly beneficial for spoken language processing tasks like ASR and ST. To address these limitations, this thesis investigates the use of LMs in E2E spoken language processing. To modularise E2E ASR systems, a decoupled structure is proposed, which modularises the internal LM in both attention-based encoder-decoder and neural transducer frameworks while maintaining E2E gradient back-propagation. By directly replacing the internal LM at decoding, efficient domain adaptation is achieved, and the biases introduced by the training data distribution are mitigated. The proposed decoupled structure is applied to both ASR and ST tasks. Experiments show that applying the decoupled structure to the AED model yields a 15.1% relative word error rate (WER) reduction on the cross-domain TED-LIUM 2 dataset compared with the best baseline, while substantially improving intra-domain performance on LibriSpeech. When applied to the online neural transducer, it achieves a 12.2% relative WER reduction over the best baseline while maintaining comparable intra-domain performance on LibriSpeech. This thesis first explores modularising the internal LM of E2E ASR and ST models and then further extends large language models (LLMs) so that they directly handle spoken language processing tasks. Given that the decoupled structure does not improve intra-domain streaming ASR performance, the label-synchronous neural transducer (LS-Transducer) is proposed. This new framework serves as an alternative to the traditional frame-synchronous neural transducer, naturally supporting low-latency processing and explicitly integrating a modular internal LM component. Experiments on LibriSpeech showed that LS-Transducer outperformed standard neural transducers, achieving 12.9% and 24.6% relative WER reductions in intra- and cross-domain settings, respectively. The LS-Transducer can also be extended to simultaneous speech translation, which is even more challenging compared to streaming ASR as it requires both streaming and re-ordering capabilities. By accumulating frame-level weights from left to right to decide when to emit tokens and using the attention mechanism to extract label-level representations, the LS-Transducer is adapted to naturally possess these two properties. Experiments on the Fisher-CallHome Spanish and MuST-C English-German speech translation tasks showed that the adapted LS-Transducer gives a better quality-latency trade-off than existing popular methods. For example, it gave a 3.1 point BLEU increase at a similar latency and a 1.4 s reduction in average lagging latency with similar BLEU scores. In addition to modularising the internal LM of the E2E speech system, this thesis also extends pre-trained large language models (LLMs) to directly handle speech input. However, since standard LLMs are pre-trained on text data, adapting them to handle speech input is not straightforward, and such speech-enabled LLMs are often limited to speech tasks they have been trained on and can not maintain good performance for unseen tasks. Wav2Prompt is proposed in order to take the first step towards extending LLMs to a wide range of zero-shot speech tasks including speech translation, understanding, question answering, etc, using only ASR training data. Wav2Prompt unlocks the zero-shot capability of the speech-enabled LLMs by enforcing the consistency between speech prompts and corresponding LLM token embeddings. Experimental results showed that for these zero-shot tasks, Wav2Prompt performs similarly to an ASR-LLM cascade and outperforms recent work. With limited task-specific paired data, end-to-end fine-tuning of the Wav2Prompt-LLM combination yields greatly improved results, e.g., a 5 point BLEU increase over the ASR-LLM cascade in English-French ST. While LLMs have been successfully extended to handle the speech modality, streaming remains challenging because speech is pre-pended as a prompt for the entire generation process. Hence, SimulS2S-LLM is also proposed to support streaming and direct speech generation, i.e. simultaneous speech-to-speech translation based on LLMs. In order to avoid constraining speech LLMs to specific streaming tasks, SimulS2S-LLM trains the speech LLM offline and employs a test-time policy to guide simultaneous inference. SimulS2S-LLM alleviates the mismatch between training and inference by extracting boundary-aware speech prompts, which allows it to be better matched with text input data. SimulS2S-LLM achieves speech-to-speech translation by predicting discrete output speech tokens and then synthesising output speech using a pre-trained vocoder. An incremental beam search is designed to expand the search space of speech token prediction without increasing latency. Experiments on the CVSS speech data showed that SimulS2S-LLM offers a better translation quality-latency trade-off than existing methods that use the same training data, such as improving ASR-BLEU scores by 3 points at similar latency.

Degree

thesis:*
Name dc:type.qualificationname
Doctor of Philosophy (PhD)
Level dc:type.qualificationlevel
Doctoral
Grantor dc:publisher.institution
University of Cambridge
Year dc:date.issued
2025

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Deng, Keqi
Advisor dc:contributor.advisor
  • Woodland, Philip

Subjects

dc:subject × 4

Rights

dc:rights
Language dc:language
eng

Identifiers

dc:identifier.*
DOI dc:identifier.doi
https://doi.org/10.17863/CAM.124014
OAI identifier oai:identifier
oai:www.repository.cam.ac.uk:1810/393856

Chain of custody

source
Harvested from
Cambridge University
Base URL
api.repository.cam.ac.uk/server/oai/request
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Deng, Keqi. The Use of Language Models in End-to-End Spoken Language Processing. Doctoral thesis, University of Cambridge, 2025. https://doi.org/10.17863/CAM.124014