University of Cambridge
Efficient and Composable Adaptation for Cross-Lingual Transfer
Abstract
dc:description.abstractParameter-efficient fine-tuning (PEFT) has emerged as a important technique for moderating the growing cost of fine-tuning state-of-the-art pre-trained language models. The modular properties of some PEFT techniques, such as reusability, composability and resistance to overfitting, lend them to applications in cross-lingual transfer, where data from "high-resource" source language(s) is employed to improve performance of NLP systems for (typically lower-resource) target languages. For instance, two independently trained PEFT modules, one for a specific task and one for a specific target language, can be composed in zero-shot fashion to perform transfer to that task-language combination. The main goal of this thesis is to advance our knowledge of, and techniques for, PEFT and cross-lingual transfer, both independently and especially where the former is used as a tool for the latter. I pay particular attention to three aspects of this problem: (i) how best to make use of the various sources of data available for cross-lingual transfer; (ii) how to devise PEFT methods which deliver the best cross-lingual transfer performance; and (iii) how to improve the efficiency of PEFT both in and beyond the context of cross-lingual transfer. The main contributions of this thesis are as follows: we first propose MAD-G, a hypernetwork method which learns to generate bottleneck adapters (a form of PEFT module) for specific languages from vectors of their typological features (such as word order and language family). In addition to the significant efficiency benefit over training independent adapters for single languages, this opens the possibility of generating modules for the many languages for which typological features are available but which lack any significant text data. We then propose sparse fine-tuning (SFT) as an alternative PEFT method for cross-lingual transfer to address the inexpressivity of bottleneck adapters. We present Lottery-Ticket Sparse Fine-Tuning (LT-SFT), a new SFT method, and show that SFTs trained with this method can be composed to obtain better cross-lingual transfer performance than with adapters. Next, we propose "BiStillation," a distillation-based method for creating more efficient models for specific cross-lingual transfer pairs with minimal loss of performance compared to the base multilingual model, and which outperforms multilingually-distilled models. Whereas much work in cross-lingual transfer confines itself to considering one or two data sources, we then consider a more general class of scenarios, where a mixture of labelled, unlabelled and parallel data is available for the source-target pair(s) of interest. We show how these data sources can be incorporated into a unified training setup in a complementary manner using a range of cross-lingual transfer techniques, yielding large performance improvements for low-resource languages in our evaluation. Finally, we address the time- and memory-inefficiency of previously proposed SFT methods, including LT-SFT, by proposing SpIEL, which preserves memory efficiency by maintaining a fixed-size set of trainable parameters which is periodically updated by replacing those which have changed the least from their pre-trained values with others which appear most likely to reduce the loss in the future. We show that SpIEL performs favourably compared to state-of-the-art PEFT methods for instruction tuning of large language models and with similar efficiency.
Degree
thesis:*- Name dc:type.qualificationname
- Doctor of Philosophy (PhD)
- Level dc:type.qualificationlevel
- Doctoral
- Grantor dc:publisher.institution
- University of Cambridge
- Year dc:date.issued
- 2024
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Ansell, Alan
- Advisors dc:contributor.advisor
-
- Korhonen, Anna
- Vulic, Ivan
- Ponti, Edoardo
Subjects
dc:subject × 3Rights
dc:rights- Licence
- Language dc:language
- eng
Identifiers
dc:identifier.*- DOI dc:identifier.doi
- https://doi.org/10.17863/CAM.118424
- OAI identifier oai:identifier
- oai:www.repository.cam.ac.uk:1810/384413