{"id":{"repo_id":"rgu","oai_identifier":"oai:rgu-repository.worktribe.com:3476013"},"canonical_url":"https://search.dev.ndltd.org/etd/rgu/oai:rgu-repository.worktribe.com:3476013","repository":{"repo_id":"rgu","name":"Robert Gordon University","base_url":"https://rgu-repository.worktribe.com/oaiprovider"},"display":{"title":"Beyond pre-training: continual learning and hallucinations in transformer-based language models.","abstract":"Pre-trained transformer models have become the norm for various language modelling tasks from document similarity analysis and text classification to natural language generation. Despite the impressive performance on benchmark datasets, adopting pretrained models for real-world applications often requires additional stages of model fine-tuning and guard-rails setup. This thesis identifies and addresses two key research gaps in these additional stages that currently limits the application of pre-trained language models. First, the thesis addresses practical scenarios in text classification where new data domains and tasks are defined incrementally over a period of time, requiring models to continually learn on newly acquired data. The research investigates existing approaches to continual learning and identifies a gap in how impact to previous domains/tasks is determined without access to the data. Based on this, a novel method is developed that takes into account the relative importance of model parameters across old and new tasks for regularising the model training more effectively. The proposed method is evaluated on various continual learning scenarios, demonstrating improvements against state-of-the-art baselines. Next, the thesis addresses safety in natural language generation scenarios where models tend to hallucinate factual information. A study of recent approaches to detecting and mitigating hallucinations in large language models (LLMs) is conducted, identifying (1) a gap in how different LLM layers are used during detection and (2) the need for fine-grained detection between hallucinated and non-hallucinated responses to the same query. A novel hallucination detection method is presented, which trains an attention-based supervised classifier on the LLM residual stream in order to automatically identify the most relevant layers. The proposed method is evaluated on various factual recall scenarios, demonstrating improved detection compared to state-of-the-art baselines on both greedy and sampled responses. By addressing current challenges in continual learning and hallucination detection, this thesis aims to reduce the gap between pre-training and adoption of language models.","abstract_html":"Pre-trained transformer models have become the norm for various language modelling tasks from document similarity analysis and text classification to natural language generation. Despite the impressive performance on benchmark datasets, adopting pretrained models for real-world applications often requires additional stages of model fine-tuning and guard-rails setup. This thesis identifies and addresses two key research gaps in these additional stages that currently limits the application of pre-trained language models. First, the thesis addresses practical scenarios in text classification where new data domains and tasks are defined incrementally over a period of time, requiring models to continually learn on newly acquired data. The research investigates existing approaches to continual learning and identifies a gap in how impact to previous domains/tasks is determined without access to the data. Based on this, a novel method is developed that takes into account the relative importance of model parameters across old and new tasks for regularising the model training more effectively. The proposed method is evaluated on various continual learning scenarios, demonstrating improvements against state-of-the-art baselines. Next, the thesis addresses safety in natural language generation scenarios where models tend to hallucinate factual information. A study of recent approaches to detecting and mitigating hallucinations in large language models (LLMs) is conducted, identifying (1) a gap in how different LLM layers are used during detection and (2) the need for fine-grained detection between hallucinated and non-hallucinated responses to the same query. A novel hallucination detection method is presented, which trains an attention-based supervised classifier on the LLM residual stream in order to automatically identify the most relevant layers. The proposed method is evaluated on various factual recall scenarios, demonstrating improved detection compared to state-of-the-art baselines on both greedy and sampled responses. By addressing current challenges in continual learning and hallucination detection, this thesis aims to reduce the gap between pre-training and adoption of language models.","abstract_has_math":false,"creators":["Suresh, Malavika"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["I. Nkisi-Orji, N. Wiratunga and R. Aljundi"],"committee_chairs":[],"committee_members":[],"year":2026,"date_issued":"2026","date_published":"2026","updated_at":"2026-07-24T04:10:14Z","subjects":["Activation probing","Continual learning","Hallucination detection","Large language models","Parameter importance based regularisation","Transformers"],"languages":["en"],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["oai:rgu-repository.worktribe.com:3476013","https://doi.org/10.48526/rgu-wt-3476013"],"render_values":[{"text":"oai:rgu-repository.worktribe.com:3476013","href":null,"code":true},{"text":"https://doi.org/10.48526/rgu-wt-3476013","href":"https://doi.org/10.48526/rgu-wt-3476013","code":true}]}]},"links":{"outbound_url":"https://rgu-repository.worktribe.com/3476013/1/SURESH%202026%20Beyond%20pre-training","outbound_label":"Repository record","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["I. Nkisi-Orji, N. Wiratunga and R. Aljundi"]},{"key":"dc:contributor.sponsor","label":"Sponsor","values":["No Funder Acknowledged (Outputs)"]},{"key":"dc:creator","label":"Author","values":["Suresh, Malavika"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2026-03-31"]},{"key":"dc:date.issued","label":"Date","values":["2026"]},{"key":"dc:relation.isreferencedby","label":"Dc Relation Isreferencedby","values":["https://rgu-repository.worktribe.com/output/3476013"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Activation probing","Continual learning","Hallucination detection","Large language models","Parameter importance based regularisation","Transformers"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["oai:rgu-repository.worktribe.com:3476013","https://doi.org/10.48526/rgu-wt-3476013"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://rgu-repository.worktribe.com/3476013/1/SURESH%202026%20Beyond%20pre-training"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Pre-trained transformer models have become the norm for various language modelling tasks from document similarity analysis and text classification to natural language generation. Despite the impressive performance on benchmark datasets, adopting pretrained models for real-world applications often requires additional stages of model fine-tuning and guard-rails setup. This thesis identifies and addresses two key research gaps in these additional stages that currently limits the application of pre-trained language models. First, the thesis addresses practical scenarios in text classification where new data domains and tasks are defined incrementally over a period of time, requiring models to continually learn on newly acquired data. The research investigates existing approaches to continual learning and identifies a gap in how impact to previous domains/tasks is determined without access to the data. Based on this, a novel method is developed that takes into account the relative importance of model parameters across old and new tasks for regularising the model training more effectively. The proposed method is evaluated on various continual learning scenarios, demonstrating improvements against state-of-the-art baselines. Next, the thesis addresses safety in natural language generation scenarios where models tend to hallucinate factual information. A study of recent approaches to detecting and mitigating hallucinations in large language models (LLMs) is conducted, identifying (1) a gap in how different LLM layers are used during detection and (2) the need for fine-grained detection between hallucinated and non-hallucinated responses to the same query. A novel hallucination detection method is presented, which trains an attention-based supervised classifier on the LLM residual stream in order to automatically identify the most relevant layers. The proposed method is evaluated on various factual recall scenarios, demonstrating improved detection compared to state-of-the-art baselines on both greedy and sampled responses. By addressing current challenges in continual learning and hallucination detection, this thesis aims to reduce the gap between pre-training and adoption of language models."]},{"key":"dc:title","label":"Title","values":["Beyond pre-training: continual learning and hallucinations in transformer-based language models."]}]}],"canonical_facts":{"dc:contributor.advisor":["I. Nkisi-Orji, N. Wiratunga and R. Aljundi"],"dc:contributor.sponsor":["No Funder Acknowledged (Outputs)"],"dc:creator":["Suresh, Malavika"],"dc:date":["2026-03-31"],"dc:date.issued":["2026"],"dc:description.abstract":["Pre-trained transformer models have become the norm for various language modelling tasks from document similarity analysis and text classification to natural language generation. Despite the impressive performance on benchmark datasets, adopting pretrained models for real-world applications often requires additional stages of model fine-tuning and guard-rails setup. This thesis identifies and addresses two key research gaps in these additional stages that currently limits the application of pre-trained language models. First, the thesis addresses practical scenarios in text classification where new data domains and tasks are defined incrementally over a period of time, requiring models to continually learn on newly acquired data. The research investigates existing approaches to continual learning and identifies a gap in how impact to previous domains/tasks is determined without access to the data. Based on this, a novel method is developed that takes into account the relative importance of model parameters across old and new tasks for regularising the model training more effectively. The proposed method is evaluated on various continual learning scenarios, demonstrating improvements against state-of-the-art baselines. Next, the thesis addresses safety in natural language generation scenarios where models tend to hallucinate factual information. A study of recent approaches to detecting and mitigating hallucinations in large language models (LLMs) is conducted, identifying (1) a gap in how different LLM layers are used during detection and (2) the need for fine-grained detection between hallucinated and non-hallucinated responses to the same query. A novel hallucination detection method is presented, which trains an attention-based supervised classifier on the LLM residual stream in order to automatically identify the most relevant layers. The proposed method is evaluated on various factual recall scenarios, demonstrating improved detection compared to state-of-the-art baselines on both greedy and sampled responses. By addressing current challenges in continual learning and hallucination detection, this thesis aims to reduce the gap between pre-training and adoption of language models."],"dc:identifier":["oai:rgu-repository.worktribe.com:3476013","https://doi.org/10.48526/rgu-wt-3476013"],"dc:identifier.uri":["https://rgu-repository.worktribe.com/3476013/1/SURESH%202026%20Beyond%20pre-training"],"dc:language":["en"],"dc:relation.isreferencedby":["https://rgu-repository.worktribe.com/output/3476013"],"dc:subject":["Activation probing","Continual learning","Hallucination detection","Large language models","Parameter importance based regularisation","Transformers"],"dc:title":["Beyond pre-training: continual learning and hallucinations in transformer-based language models."],"dc:type":["Thesis"]},"updated_at":"2026-07-24T04:10:14Z"}