University of Cambridge
From data to model behaviour: A causal and empirical analysis of how data shapes the behaviour of language models
Abstract
dc:description.abstractLarge language models have achieved remarkable capabilities, yet their behaviour remains difficult to predict and understand. A central challenge is that these models are shaped not only by their architectures or training algorithms, but fundamentally by their training data---what examples they see, how those examples are represented, when they encounter them, and how much influence individual examples (or batches of examples) exert. Despite the critical role of training data, a systematic understanding of how data-centric decisions propagate through the training pipeline to affect final model behaviour has remained limited. This thesis addresses this gap by examining four stages of the data pipeline: data sourcing and curation, data representation, model training, and model evaluation. The thesis makes four primary contributions, each addressing a distinct stage. First, I introduce AnchorAL, a novel Active Learning approach that addresses the computational and learning challenges of data sourcing and curation. The method maintains constant computational costs whilst discovering more minority class examples and (often) achieving higher performance than previous comparable approaches. Second, I develop a causal framework using the Regression Discontinuity Design to measure tokenisation bias---how vocabulary choices causally affect probability assignments. The analysis reveals that character spans represented as single tokens receive up to 17 times more probability than when split, with effects persisting even in larger models. Third, through PolyPythia---a suite of 45 new training runs spanning 5 model sizes---I characterise training stability across randomness factors. The analysis identifies remarkably consistent learning phases across seeds and scales, whilst revealing that smaller models can predict larger models' learning dynamics. Fourth, I develop a causal framework using the Difference-in-Differences methodology to measure memorisation as the causal effect of exposure to training examples. This yields memorisation profiles that track how individual batches are memorised throughout training, revealing that memorisation is stronger and more persistent in larger models, with patterns stable enough to enable cross-scale prediction. Throughout these contributions, several unifying themes emerge. Causal inference techniques provide principled methods for measuring data-centric effects that would be computationally prohibitive to estimate through direct experimentation. Scaling analyses reveal both continuities and discontinuities, informing what can be inferred from smaller models and which phenomena reflect fundamental properties rather than artefacts of scale. This thesis advances the broader research agenda of data-centric AI, demonstrating that whilst model architectures and training algorithms matter, the data itself remains fundamental. The methods developed---from strategic data selection, through causal frameworks for tokenisation bias and memorisation, to training stability analysis---provide both conceptual tools for reasoning about data-to-model relationships and practical techniques for measuring them. As language models continue to grow in scale and capability, understanding their relationship with training data becomes ever more critical for building trustworthy, effective systems.
Degree
thesis:*- Name dc:type.qualificationname
- Doctor of Philosophy (PhD)
- Level dc:type.qualificationlevel
- Doctoral
- Grantor dc:publisher.institution
- University of Cambridge
- Year dc:date.issued
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Lesci, Pietro
- Advisor dc:contributor.advisor
-
- Vlachos, Andreas
Subjects
dc:subject × 5Rights
dc:rightsIdentifiers
dc:identifier.*- DOI dc:identifier.doi
- https://doi.org/10.17863/CAM.131697
- OAI identifier oai:identifier
- oai:www.repository.cam.ac.uk:1810/405504