Back to results

University of Cambridge

Efficient and Composable Adaptation for Cross-Lingual Transfer

Abstract

dc:description.abstract

Parameter-efficient fine-tuning (PEFT) has emerged as a important technique for moderating the growing cost of fine-tuning state-of-the-art pre-trained language models. The modular properties of some PEFT techniques, such as reusability, composability and resistance to overfitting, lend them to applications in cross-lingual transfer, where data from "high-resource" source language(s) is employed to improve performance of NLP systems for (typically lower-resource) target languages. For instance, two independently trained PEFT modules, one for a specific task and one for a specific target language, can be composed in zero-shot fashion to perform transfer to that task-language combination. The main goal of this thesis is to advance our knowledge of, and techniques for, PEFT and cross-lingual transfer, both independently and especially where the former is used as a tool for the latter. I pay particular attention to three aspects of this problem: (i) how best to make use of the various sources of data available for cross-lingual transfer; (ii) how to devise PEFT methods which deliver the best cross-lingual transfer performance; and (iii) how to improve the efficiency of PEFT both in and beyond the context of cross-lingual transfer. The main contributions of this thesis are as follows: we first propose MAD-G, a hypernetwork method which learns to generate bottleneck adapters (a form of PEFT module) for specific languages from vectors of their typological features (such as word order and language family). In addition to the significant efficiency benefit over training independent adapters for single languages, this opens the possibility of generating modules for the many languages for which typological features are available but which lack any significant text data. We then propose sparse fine-tuning (SFT) as an alternative PEFT method for cross-lingual transfer to address the inexpressivity of bottleneck adapters. We present Lottery-Ticket Sparse Fine-Tuning (LT-SFT), a new SFT method, and show that SFTs trained with this method can be composed to obtain better cross-lingual transfer performance than with adapters. Next, we propose "BiStillation," a distillation-based method for creating more efficient models for specific cross-lingual transfer pairs with minimal loss of performance compared to the base multilingual model, and which outperforms multilingually-distilled models. Whereas much work in cross-lingual transfer confines itself to considering one or two data sources, we then consider a more general class of scenarios, where a mixture of labelled, unlabelled and parallel data is available for the source-target pair(s) of interest. We show how these data sources can be incorporated into a unified training setup in a complementary manner using a range of cross-lingual transfer techniques, yielding large performance improvements for low-resource languages in our evaluation. Finally, we address the time- and memory-inefficiency of previously proposed SFT methods, including LT-SFT, by proposing SpIEL, which preserves memory efficiency by maintaining a fixed-size set of trainable parameters which is periodically updated by replacing those which have changed the least from their pre-trained values with others which appear most likely to reduce the loss in the future. We show that SpIEL performs favourably compared to state-of-the-art PEFT methods for instruction tuning of large language models and with similar efficiency.

Degree

thesis:*
Name dc:type.qualificationname
Doctor of Philosophy (PhD)
Level dc:type.qualificationlevel
Doctoral
Grantor dc:publisher.institution
University of Cambridge
Year dc:date.issued
2024

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Ansell, Alan
Advisors dc:contributor.advisor
  • Korhonen, Anna
  • Vulic, Ivan
  • Ponti, Edoardo

Subjects

dc:subject × 3

Rights

dc:rights
Language dc:language
eng

Identifiers

dc:identifier.*
DOI dc:identifier.doi
https://doi.org/10.17863/CAM.118424
OAI identifier oai:identifier
oai:www.repository.cam.ac.uk:1810/384413

Chain of custody

source
Harvested from
Cambridge University
Base URL
api.repository.cam.ac.uk/server/oai/request
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Ansell, Alan. Efficient and Composable Adaptation for Cross-Lingual Transfer. Doctoral thesis, University of Cambridge, 2024. https://doi.org/10.17863/CAM.118424