Robert Gordon University
Beyond pre-training: continual learning and hallucinations in transformer-based language models.
Abstract
dc:description.abstractPre-trained transformer models have become the norm for various language modelling tasks from document similarity analysis and text classification to natural language generation. Despite the impressive performance on benchmark datasets, adopting pretrained models for real-world applications often requires additional stages of model fine-tuning and guard-rails setup. This thesis identifies and addresses two key research gaps in these additional stages that currently limits the application of pre-trained language models. First, the thesis addresses practical scenarios in text classification where new data domains and tasks are defined incrementally over a period of time, requiring models to continually learn on newly acquired data. The research investigates existing approaches to continual learning and identifies a gap in how impact to previous domains/tasks is determined without access to the data. Based on this, a novel method is developed that takes into account the relative importance of model parameters across old and new tasks for regularising the model training more effectively. The proposed method is evaluated on various continual learning scenarios, demonstrating improvements against state-of-the-art baselines. Next, the thesis addresses safety in natural language generation scenarios where models tend to hallucinate factual information. A study of recent approaches to detecting and mitigating hallucinations in large language models (LLMs) is conducted, identifying (1) a gap in how different LLM layers are used during detection and (2) the need for fine-grained detection between hallucinated and non-hallucinated responses to the same query. A novel hallucination detection method is presented, which trains an attention-based supervised classifier on the LLM residual stream in order to automatically identify the most relevant layers. The proposed method is evaluated on various factual recall scenarios, demonstrating improved detection compared to state-of-the-art baselines on both greedy and sampled responses. By addressing current challenges in continual learning and hallucination detection, this thesis aims to reduce the gap between pre-training and adoption of language models.
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Suresh, Malavika
- Advisor dc:contributor.advisor
-
- I. Nkisi-Orji, N. Wiratunga and R. Aljundi
Subjects
dc:subject × 6Rights
- Language dc:language
- en
Identifiers
dc:identifier.*- Identifier
-
oai:rgu-repository.worktribe.com:3476013
https://doi.org/10.48526/rgu-wt-3476013 - OAI identifier oai:identifier
- oai:rgu-repository.worktribe.com:3476013