Back to results

University of Cambridge

Grounded Language Learning with Foundation Models

Abstract

dc:description.abstract

The field of Natural Language Processing (NLP) has experienced remarkable advancement in recent years, especially with the development of foundation models trained using language modelling objectives. These models have achieved a level of proficiency comparable to human competence across an array of applications. Through unsupervised pretraining on web data, foundation models gain a wealth of world knowledge and exhibit high competence in both probing and downstream tasks. This thesis begins with a probing study where I demonstrate that language models of sufficient scales pretrained on unstructured text can develop some grounded understanding such as accurately capturing perceptual qualities eg. colours of objects. This suggests that large-scale text compression may offer a pathway to learning grounded knowledge. Yet this raises a deeper question: is text-only training truly sufficient for achieving general intelligence? I argue that existing text foundation models still lack comprehensive grounding capabilities to effectively interact with the world, which is intrinsically multimodal and exhibits complexity beyond what is sufficiently described in language. This thesis investigate these grounding limitations from language models with challenging evaluations and introduces novel methods to address these shortcomings and enhance the grounding capabilities of pretrained foundation models. To this end, the thesis focuses on two key sources of information essential for grounding language models: (1) structured knowledge from human-curated tools and (2) perceptual information from vision. For grounding language models with structured knowledge, I designed a contrastive learning framework capable of injecting structured knowledge into language models through continual pretraining. The efficacy of this method is showcased by the state-of-the-art (SOTA) models I built in specialised domains where robust reasoning on fine-grained relational knowledge is essential. To this end, I put forward two challenging benchmarks: one on spatial reasoning which is a fundamental cornerstone for human cognition, and another on examining vision-language models' understandings of frequent concepts across languages and cultures. I also provided extensive analysis and insights on vision-language training strategies. In particularly, I proposed effective vision-language pretraining methods that achieved SOTA performance on visual language reasoning tasks by distilling skills from code executors. In conclusion, this thesis has explored the strengths of foundational language models, investigated their weaknesses as intelligent agents, and developed strategies for their refinement through grounding. These advancements pave the way for more effective multimodal interaction by enabling the integration of structured knowledge and superior visual reasoning. This progress represents a critical step towards the next generation of NLP technologies, built on top of large-scale pretraining, and capable of richer and more nuanced interactions. Finally, I discuss that scalable grounding can be the next crucial step towards building general intelligence.

Degree

thesis:*
Name dc:type.qualificationname
Doctor of Philosophy (PhD)
Level dc:type.qualificationlevel
Doctoral
Grantor dc:publisher.institution
University of Cambridge
Year dc:date.issued
2025

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Liu, Fangyu
Advisor dc:contributor.advisor
  • Collier, Nigel

Subjects

dc:subject × 4

Rights

dc:rights
Language dc:language
eng

Identifiers

dc:identifier.*
Author Identifier
0000-0001-7038-3623
OAI identifier oai:identifier
oai:www.repository.cam.ac.uk:1810/392048

Chain of custody

source
Harvested from
Cambridge University
Base URL
api.repository.cam.ac.uk/server/oai/request
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Liu, Fangyu. Grounded Language Learning with Foundation Models. Doctoral thesis, University of Cambridge, 2025. https://doi.org/10.17863/CAM.122924