Reykjavík University
Analyzing Icelandic conversation using State-of-the-Art ASR models
Abstract
dc:description.abstractThis thesis investigates the performance and fine-tuning of transformer-based Automatic Speech Recognition (ASR) systems applied to Icelandic conversational speech, with a primary focus on OpenAI’s Whisper model and a secondary fo- cus on Meta’s Wav2Vec 2.0. The study uses the Spjallrómur dataset, a 21-hour unscripted dialogue dataset between two speakers for benchmarking and model training. The project examines both reduced parameter number Low-Rank Adap- tation ( LoRA) and full parameter fine-tuning strategies. In this thesis, only full- parameter fine-tuning was implemented and tested. Both the base Whisper-Small and Whisper-Large models are used, starting from a model pre-trained on 967 hours of read Icelandic speech, which served as the basis for further fine-tuning. Fine-tuning from the 967h pre-trained Whisper-Large model reduced the WER from 29.84% to 22.69% and the CER from 20.48% to 13.10%. This indicates that further fine-tuning with low-quality data yields a notable improvement. Whisper- small improved considerably, reducing its Word Error Rate (WER) from 145.71% to WER of 48.30% (with CER of 29.37%). These findings indicate that even a small and misaligned dataset can yield substantial improvement on both OpenAI’s - Small and Whisper-Large. Keywords/Efnisorð: Whisper, Wav2Vec 2.0, K2, Icelandic, ASR, fine-tuning, LoRA, transformer models, conversational speech
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Páll Rúnarsson 1982-
- Contributors dc:contributor
-
- Háskólinn í Reykjavík
Subjects
dc:subject × 9Rights
- Language dc:language.iso
- en
Identifiers
dc:identifier.*- Handle dc:identifier.uri
- https://hdl.handle.net/1946/50888
- OAI identifier oai:identifier
- oai:skemman.is:1946/50888