University of Cambridge
Inductive Biases and Generalisation in Models for Natural Language Processing
Abstract
dc:description.abstractComputers that can interact in a naturalistic way using language have long been a staple of science-fiction. Researchers in computing and artificial intelligence have made considerable efforts to realise this aim but, due to the complexity of natural language and the subtleties and nuances involved in its use, it has often been held up as an idealised, potentially unreachable goal (Turing, 1950). However, recent breakthroughs made by large language models (LLMs), such as GPT-4 (OpenAI, 2023), have made this once seemingly unassailable target feel closer than ever before. As impressive as they are, these LLMs only achieve such remarkable performance through training on enormous amounts of data, consuming large quantities of resources in the process. Human children learn to competently and fluently use language with a much smaller volume of input than current state-of-the-art computational models (Hart and Risley, 1992; Gilkerson et al., 2017). This suggests that models are not making efficient use of the data that they are being trained on, and are failing to infer correct generalisation strategies from smaller quantities of data. This leads us to ask whether imbuing model architectures with carefully considered inductive biases might lead to improved data-efficiency and performance, by nudging them towards solutions of particular forms and enabling them to generalise from small amounts of data in a more human-like way. This thesis describes work which focuses on the inductive biases and generalisation capabilities of models used for Natural Language Processing (NLP). It begins by describing a highly-controlled method for investigating the unintentional inductive biases that may arise through architectural choices and whether they may lead to certain types of natural languages being modelled more successfully than others. These experiments conclude that the inductive bias of transformers lends itself to languages with certain word orders over others, in contrast with LSTMs whose performance is largely agnostic to word order. It also demonstrates that there is no correspondence between the inductive bias of transformers and theorised human cognitive biases. Subsequently, the thesis discusses how an inductive bias towards compositional solutions can be beneficial for language-based tasks and demonstrates how this can be achieved through a group-equivariant architecture. The architecture used builds on existing work, but makes changes that result in both theoretical and empirical advantages, outperforming previous similar models on all splits of the SCAN benchmark. Finally, the thesis provides a systematic evaluation of how architectures handle tasks containing long-range dependencies, and an analysis of which elements of long-range dependencies pose difficulties. It finds that the most critical factors across architectures are the quantity of information that must be processed as part of a dependency, and the complexity of the processing required. It finds that the gap between a token and those on which it depends has little effect on performance, which brings into question the focus of some recent work on long-range dependencies. Taken together, this work highlights the importance of careful, controlled evaluations of commonly-used architectures in order to understand how their inductive biases affect their training process and ultimate performance, as well as motivating the continued pursuit of architectures with improved inductive biases to achieve the desired generalisation strategies on difficult language-based tasks. Without improved data-efficiency and robust generalisation capabilities, it is difficult to truly begin to compare the linguistic and reasoning capabilities of LLMs with those of humans.
Degree
thesis:*- Name dc:type.qualificationname
- Doctor of Philosophy (PhD)
- Level dc:type.qualificationlevel
- Doctoral
- Grantor dc:publisher.institution
- University of Cambridge
- Year dc:date.issued
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- White, Jennifer
- Advisors dc:contributor.advisor
-
- Teufel, Simone
- Cotterell, Ryan
Subjects
dc:subject × 2Rights
dc:rightsIdentifiers
dc:identifier.*- DOI dc:identifier.doi
- https://doi.org/10.17863/CAM.123310
- OAI identifier oai:identifier
- oai:www.repository.cam.ac.uk:1810/392665