University of Cambridge
Multimodal learning and language models for enhanced knowledge representations
Abstract
dc:description.abstractMachine learning with modern neural architectures, particularly large-scale language models, has enabled impressive progress across diverse problem domains such as natural language processing, graph representation learning, and tabular data analysis. However, these successes have predominantly relied on abundant, single-modal datasets, often overlooking the inherently multimodal and structurally complex nature of real-world data. Real-world data typically combines multiple modalities—text, graph structures, and tabular formats—each encoding distinct but complementary insights. Effectively integrating these heterogeneous sources remains one of the most significant open challenges in artificial intelligence research. In this thesis, I directly address this challenge by proposing three pioneering contributions that leverage multimodal data through sophisticated neural architectures and advanced representation learning techniques. First, I introduce HyperBERT, a novel neural architecture that integrates hypergraph neural networks with pretrained language models, achieving superior node classification performance on text-attributed hypergraphs. The second contribution tackles dense retrieval in environments with limited labeled data by developing a novel weakly supervised semantic distillation framework. Leveraging the rich semantic understanding embedded in large language models along with negative sampling strategies, this framework significantly improves retrieval accuracy and generalizability, particularly for tasks such as claim verification and evidence retrieval. The third contribution introduces TabMDA, a Transformer-based manifold data augmentation method for tabular data, leveraging pretrained Transformer encoders in combination with an advanced in-context subsetting technique. This approach generates synthetic samples that faithfully preserve data dependencies, enhancing downstream classification performance. Collectively, these interconnected contributions significantly advance the frontier of multimodal representation learning. Through the explicit modeling and integration of structural, textual, and tabular information, this thesis paves the way towards building more sophisticated, adaptable, and intelligent computational systems capable of nuanced understanding and interaction with complex real-world information structures.
Degree
thesis:*- Name dc:type.qualificationname
- Doctor of Philosophy (PhD)
- Level dc:type.qualificationlevel
- Doctoral
- Grantor dc:publisher.institution
- University of Cambridge
- Year dc:date.issued
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Rodriguez Bazaga, Adrian
- Advisors dc:contributor.advisor
-
- Micklem, Gos
- Lio, Pietro
Subjects
dc:subject × 13Rights
dc:rightsIdentifiers
dc:identifier.*- DOI dc:identifier.doi
- https://doi.org/10.17863/CAM.125995
- OAI identifier oai:identifier
- oai:www.repository.cam.ac.uk:1810/396718