{"id":{"repo_id":"cambridge","oai_identifier":"oai:www.repository.cam.ac.uk:1810/392048"},"canonical_url":"https://search.dev.ndltd.org/etd/cambridge/oai:www.repository.cam.ac.uk:1810/392048","repository":{"repo_id":"cambridge","name":"Cambridge University","base_url":"https://api.repository.cam.ac.uk/server/oai/request"},"display":{"title":"Grounded Language Learning with Foundation Models","abstract":"The field of Natural Language Processing (NLP) has experienced remarkable advancement in recent years, especially with the development of foundation models trained using language modelling objectives. These models have achieved a level of proficiency comparable to human competence across an array of applications. Through unsupervised pretraining on web data, foundation models gain a wealth of world knowledge and exhibit high competence in both probing and downstream tasks. This thesis begins with a probing study where I demonstrate that language models of sufficient scales pretrained on unstructured text can develop some grounded understanding such as accurately capturing perceptual qualities eg. colours of objects. This suggests that large-scale text compression may offer a pathway to learning grounded knowledge. Yet this raises a deeper question: is text-only training truly sufficient for achieving general intelligence? I argue that existing text foundation models still lack comprehensive grounding capabilities to effectively interact with the world, which is intrinsically multimodal and exhibits complexity beyond what is sufficiently described in language. This thesis investigate these grounding limitations from language models with challenging evaluations and introduces novel methods to address these shortcomings and enhance the grounding capabilities of pretrained foundation models. To this end, the thesis focuses on two key sources of information essential for grounding language models: (1) structured knowledge from human-curated tools and (2) perceptual information from vision. For grounding language models with structured knowledge, I designed a contrastive learning framework capable of injecting structured knowledge into language models through continual pretraining. The efficacy of this method is showcased by the state-of-the-art (SOTA) models I built in specialised domains where robust reasoning on fine-grained relational knowledge is essential. To this end, I put forward two challenging benchmarks: one on spatial reasoning which is a fundamental cornerstone for human cognition, and another on examining vision-language models' understandings of frequent concepts across languages and cultures. I also provided extensive analysis and insights on vision-language training strategies. In particularly, I proposed effective vision-language pretraining methods that achieved SOTA performance on visual language reasoning tasks by distilling skills from code executors. In conclusion, this thesis has explored the strengths of foundational language models, investigated their weaknesses as intelligent agents, and developed strategies for their refinement through grounding. These advancements pave the way for more effective multimodal interaction by enabling the integration of structured knowledge and superior visual reasoning. This progress represents a critical step towards the next generation of NLP technologies, built on top of large-scale pretraining, and capable of richer and more nuanced interactions. Finally, I discuss that scalable grounding can be the next crucial step towards building general intelligence.","abstract_html":"The field of Natural Language Processing (NLP) has experienced remarkable advancement in recent years, especially with the development of foundation models trained using language modelling objectives. These models have achieved a level of proficiency comparable to human competence across an array of applications. Through unsupervised pretraining on web data, foundation models gain a wealth of world knowledge and exhibit high competence in both probing and downstream tasks. This thesis begins with a probing study where I demonstrate that language models of sufficient scales pretrained on unstructured text can develop some grounded understanding such as accurately capturing perceptual qualities eg. colours of objects. This suggests that large-scale text compression may offer a pathway to learning grounded knowledge. Yet this raises a deeper question: is text-only training truly sufficient for achieving general intelligence? I argue that existing text foundation models still lack comprehensive grounding capabilities to effectively interact with the world, which is intrinsically multimodal and exhibits complexity beyond what is sufficiently described in language. This thesis investigate these grounding limitations from language models with challenging evaluations and introduces novel methods to address these shortcomings and enhance the grounding capabilities of pretrained foundation models. To this end, the thesis focuses on two key sources of information essential for grounding language models: (1) structured knowledge from human-curated tools and (2) perceptual information from vision. For grounding language models with structured knowledge, I designed a contrastive learning framework capable of injecting structured knowledge into language models through continual pretraining. The efficacy of this method is showcased by the state-of-the-art (SOTA) models I built in specialised domains where robust reasoning on fine-grained relational knowledge is essential. To this end, I put forward two challenging benchmarks: one on spatial reasoning which is a fundamental cornerstone for human cognition, and another on examining vision-language models&#x27; understandings of frequent concepts across languages and cultures. I also provided extensive analysis and insights on vision-language training strategies. In particularly, I proposed effective vision-language pretraining methods that achieved SOTA performance on visual language reasoning tasks by distilling skills from code executors. In conclusion, this thesis has explored the strengths of foundational language models, investigated their weaknesses as intelligent agents, and developed strategies for their refinement through grounding. These advancements pave the way for more effective multimodal interaction by enabling the integration of structured knowledge and superior visual reasoning. This progress represents a critical step towards the next generation of NLP technologies, built on top of large-scale pretraining, and capable of richer and more nuanced interactions. Finally, I discuss that scalable grounding can be the next crucial step towards building general intelligence.","abstract_has_math":false,"creators":["Liu, Fangyu"],"institution":"University of Cambridge","degree_name":"Doctor of Philosophy (PhD)","degree_level":"Doctoral","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Collier, Nigel"],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-07-21","date_published":"2025-07-21","updated_at":"2026-07-22T22:24:10Z","subjects":["Computational Linguistics","Deep Learning","Language Models","NLP"],"languages":["eng"],"rights":[],"rights_urls":["https://www.repository.cam.ac.uk/bitstreams/00253154-201b-4f7d-bac2-4110c232c128/download","https://creativecommons.org/licenses/by/4.0/"],"identifier_entries":[{"key":"dc:creator.authoridentifier","label":"Author Identifier","values":["0000000170383623"],"render_values":[{"text":"0000-0001-7038-3623","href":"https://orcid.org/0000-0001-7038-3623","code":true}]}]},"links":{"outbound_url":"https://doi.org/10.17863/CAM.122924","outbound_label":"DOI","outbound_source":"dc:identifier.doi"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Collier, Nigel"]},{"key":"dc:contributor.sponsor","label":"Sponsor","values":["Grace and Thomas C. H. Chan Scholarship - Cambridge Trust"]},{"key":"dc:creator","label":"Author","values":["Liu, Fangyu"]},{"key":"dc:creator.authoridentifier","label":"Author Identifier","values":["0000000170383623"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2025-07-21"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["University of Cambridge"]},{"key":"dc:relation.isreferencedby.uri","label":"Dc Relation Isreferencedby URI","values":["https://www.repository.cam.ac.uk/handle/1810/392048"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["Doctoral"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["Doctor of Philosophy (PhD)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computational Linguistics","Deep Learning","Language Models","NLP"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["https://www.repository.cam.ac.uk/bitstreams/00253154-201b-4f7d-bac2-4110c232c128/download","https://creativecommons.org/licenses/by/4.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.17863/CAM.122924"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://www.repository.cam.ac.uk/bitstreams/bc8bce5b-4715-4d03-8444-bb850d73ea9e/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["The field of Natural Language Processing (NLP) has experienced remarkable advancement in recent years, especially with the development of foundation models trained using language modelling objectives. These models have achieved a level of proficiency comparable to human competence across an array of applications. Through unsupervised pretraining on web data, foundation models gain a wealth of world knowledge and exhibit high competence in both probing and downstream tasks. This thesis begins with a probing study where I demonstrate that language models of sufficient scales pretrained on unstructured text can develop some grounded understanding such as accurately capturing perceptual qualities eg. colours of objects. This suggests that large-scale text compression may offer a pathway to learning grounded knowledge. Yet this raises a deeper question: is text-only training truly sufficient for achieving general intelligence? I argue that existing text foundation models still lack comprehensive grounding capabilities to effectively interact with the world, which is intrinsically multimodal and exhibits complexity beyond what is sufficiently described in language. This thesis investigate these grounding limitations from language models with challenging evaluations and introduces novel methods to address these shortcomings and enhance the grounding capabilities of pretrained foundation models. To this end, the thesis focuses on two key sources of information essential for grounding language models: (1) structured knowledge from human-curated tools and (2) perceptual information from vision. For grounding language models with structured knowledge, I designed a contrastive learning framework capable of injecting structured knowledge into language models through continual pretraining. The efficacy of this method is showcased by the state-of-the-art (SOTA) models I built in specialised domains where robust reasoning on fine-grained relational knowledge is essential. To this end, I put forward two challenging benchmarks: one on spatial reasoning which is a fundamental cornerstone for human cognition, and another on examining vision-language models' understandings of frequent concepts across languages and cultures. I also provided extensive analysis and insights on vision-language training strategies. In particularly, I proposed effective vision-language pretraining methods that achieved SOTA performance on visual language reasoning tasks by distilling skills from code executors. In conclusion, this thesis has explored the strengths of foundational language models, investigated their weaknesses as intelligent agents, and developed strategies for their refinement through grounding. These advancements pave the way for more effective multimodal interaction by enabling the integration of structured knowledge and superior visual reasoning. This progress represents a critical step towards the next generation of NLP technologies, built on top of large-scale pretraining, and capable of richer and more nuanced interactions. Finally, I discuss that scalable grounding can be the next crucial step towards building general intelligence."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["5301e3a8e659aaba17fccf22304c2a35","87eda9de84448d1f82354d60eee3eb5f"]},{"key":"dc:title","label":"Title","values":["Grounded Language Learning with Foundation Models"]}]}],"canonical_facts":{"dc:contributor.advisor":["Collier, Nigel"],"dc:contributor.sponsor":["Grace and Thomas C. H. Chan Scholarship - Cambridge Trust"],"dc:creator":["Liu, Fangyu"],"dc:creator.authoridentifier":["0000000170383623"],"dc:date.issued":["2025-07-21"],"dc:description.abstract":["The field of Natural Language Processing (NLP) has experienced remarkable advancement in recent years, especially with the development of foundation models trained using language modelling objectives. These models have achieved a level of proficiency comparable to human competence across an array of applications. Through unsupervised pretraining on web data, foundation models gain a wealth of world knowledge and exhibit high competence in both probing and downstream tasks. This thesis begins with a probing study where I demonstrate that language models of sufficient scales pretrained on unstructured text can develop some grounded understanding such as accurately capturing perceptual qualities eg. colours of objects. This suggests that large-scale text compression may offer a pathway to learning grounded knowledge. Yet this raises a deeper question: is text-only training truly sufficient for achieving general intelligence? I argue that existing text foundation models still lack comprehensive grounding capabilities to effectively interact with the world, which is intrinsically multimodal and exhibits complexity beyond what is sufficiently described in language. This thesis investigate these grounding limitations from language models with challenging evaluations and introduces novel methods to address these shortcomings and enhance the grounding capabilities of pretrained foundation models. To this end, the thesis focuses on two key sources of information essential for grounding language models: (1) structured knowledge from human-curated tools and (2) perceptual information from vision. For grounding language models with structured knowledge, I designed a contrastive learning framework capable of injecting structured knowledge into language models through continual pretraining. The efficacy of this method is showcased by the state-of-the-art (SOTA) models I built in specialised domains where robust reasoning on fine-grained relational knowledge is essential. To this end, I put forward two challenging benchmarks: one on spatial reasoning which is a fundamental cornerstone for human cognition, and another on examining vision-language models' understandings of frequent concepts across languages and cultures. I also provided extensive analysis and insights on vision-language training strategies. In particularly, I proposed effective vision-language pretraining methods that achieved SOTA performance on visual language reasoning tasks by distilling skills from code executors. In conclusion, this thesis has explored the strengths of foundational language models, investigated their weaknesses as intelligent agents, and developed strategies for their refinement through grounding. These advancements pave the way for more effective multimodal interaction by enabling the integration of structured knowledge and superior visual reasoning. This progress represents a critical step towards the next generation of NLP technologies, built on top of large-scale pretraining, and capable of richer and more nuanced interactions. Finally, I discuss that scalable grounding can be the next crucial step towards building general intelligence."],"dc:format.checksum.md5":["5301e3a8e659aaba17fccf22304c2a35","87eda9de84448d1f82354d60eee3eb5f"],"dc:identifier.doi":["https://doi.org/10.17863/CAM.122924"],"dc:identifier.uri":["https://www.repository.cam.ac.uk/bitstreams/bc8bce5b-4715-4d03-8444-bb850d73ea9e/download"],"dc:language":["eng"],"dc:publisher.institution":["University of Cambridge"],"dc:relation.isreferencedby.uri":["https://www.repository.cam.ac.uk/handle/1810/392048"],"dc:rights":["https://www.repository.cam.ac.uk/bitstreams/00253154-201b-4f7d-bac2-4110c232c128/download","https://creativecommons.org/licenses/by/4.0/"],"dc:subject":["Computational Linguistics","Deep Learning","Language Models","NLP"],"dc:title":["Grounded Language Learning with Foundation Models"],"dc:type":["Thesis"],"dc:type.qualificationlevel":["Doctoral"],"dc:type.qualificationname":["Doctor of Philosophy (PhD)"]},"updated_at":"2026-07-22T22:24:10Z"}