{"id":{"repo_id":"cambridge","oai_identifier":"oai:www.repository.cam.ac.uk:1810/396718"},"canonical_url":"https://search.dev.ndltd.org/etd/cambridge/oai:www.repository.cam.ac.uk:1810/396718","repository":{"repo_id":"cambridge","name":"Cambridge University","base_url":"https://api.repository.cam.ac.uk/server/oai/request"},"display":{"title":"Multimodal learning and language models for enhanced knowledge representations","abstract":"Machine learning with modern neural architectures, particularly large-scale language models, has enabled impressive progress across diverse problem domains such as natural language processing, graph representation learning, and tabular data analysis. However, these successes have predominantly relied on abundant, single-modal datasets, often overlooking the inherently multimodal and structurally complex nature of real-world data. Real-world data typically combines multiple modalities—text, graph structures, and tabular formats—each encoding distinct but complementary insights. Effectively integrating these heterogeneous sources remains one of the most significant open challenges in artificial intelligence research. In this thesis, I directly address this challenge by proposing three pioneering contributions that leverage multimodal data through sophisticated neural architectures and advanced representation learning techniques. First, I introduce HyperBERT, a novel neural architecture that integrates hypergraph neural networks with pretrained language models, achieving superior node classification performance on text-attributed hypergraphs. The second contribution tackles dense retrieval in environments with limited labeled data by developing a novel weakly supervised semantic distillation framework. Leveraging the rich semantic understanding embedded in large language models along with negative sampling strategies, this framework significantly improves retrieval accuracy and generalizability, particularly for tasks such as claim verification and evidence retrieval. The third contribution introduces TabMDA, a Transformer-based manifold data augmentation method for tabular data, leveraging pretrained Transformer encoders in combination with an advanced in-context subsetting technique. This approach generates synthetic samples that faithfully preserve data dependencies, enhancing downstream classification performance. Collectively, these interconnected contributions significantly advance the frontier of multimodal representation learning. Through the explicit modeling and integration of structural, textual, and tabular information, this thesis paves the way towards building more sophisticated, adaptable, and intelligent computational systems capable of nuanced understanding and interaction with complex real-world information structures.","abstract_html":"Machine learning with modern neural architectures, particularly large-scale language models, has enabled impressive progress across diverse problem domains such as natural language processing, graph representation learning, and tabular data analysis. However, these successes have predominantly relied on abundant, single-modal datasets, often overlooking the inherently multimodal and structurally complex nature of real-world data. Real-world data typically combines multiple modalities—text, graph structures, and tabular formats—each encoding distinct but complementary insights. Effectively integrating these heterogeneous sources remains one of the most significant open challenges in artificial intelligence research. In this thesis, I directly address this challenge by proposing three pioneering contributions that leverage multimodal data through sophisticated neural architectures and advanced representation learning techniques. First, I introduce HyperBERT, a novel neural architecture that integrates hypergraph neural networks with pretrained language models, achieving superior node classification performance on text-attributed hypergraphs. The second contribution tackles dense retrieval in environments with limited labeled data by developing a novel weakly supervised semantic distillation framework. Leveraging the rich semantic understanding embedded in large language models along with negative sampling strategies, this framework significantly improves retrieval accuracy and generalizability, particularly for tasks such as claim verification and evidence retrieval. The third contribution introduces TabMDA, a Transformer-based manifold data augmentation method for tabular data, leveraging pretrained Transformer encoders in combination with an advanced in-context subsetting technique. This approach generates synthetic samples that faithfully preserve data dependencies, enhancing downstream classification performance. Collectively, these interconnected contributions significantly advance the frontier of multimodal representation learning. Through the explicit modeling and integration of structural, textual, and tabular information, this thesis paves the way towards building more sophisticated, adaptable, and intelligent computational systems capable of nuanced understanding and interaction with complex real-world information structures.","abstract_has_math":false,"creators":["Rodriguez Bazaga, Adrian"],"institution":"University of Cambridge","degree_name":"Doctor of Philosophy (PhD)","degree_level":"Doctoral","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Micklem, Gos","Lio, Pietro"],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-04-11","date_published":"2025-04-11","updated_at":"2026-07-22T22:23:59Z","subjects":["machine learning","deep learning","language models","LLMs","graph neural networks","GNNs","multimodal learning","node classification","hypergraphs","dense retrieval","data augmentation","contrastive learning","pretraining"],"languages":["eng"],"rights":[],"rights_urls":["https://www.repository.cam.ac.uk/bitstreams/5a438c2c-912c-49a6-b87b-3b8475819040/download","http://purl.org/NET/rdflicense/allrightsreserved"],"identifier_entries":[]},"links":{"outbound_url":"https://doi.org/10.17863/CAM.125995","outbound_label":"DOI","outbound_source":"dc:identifier.doi"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Micklem, Gos","Lio, Pietro"]},{"key":"dc:creator","label":"Author","values":["Rodriguez Bazaga, Adrian"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2025-04-11"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["University of Cambridge"]},{"key":"dc:relation.isreferencedby.uri","label":"Dc Relation Isreferencedby URI","values":["https://www.repository.cam.ac.uk/handle/1810/396718"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["Doctoral"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["Doctor of Philosophy (PhD)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["machine learning","deep learning","language models","LLMs","graph neural networks","GNNs","multimodal learning","node classification","hypergraphs","dense retrieval","data augmentation","contrastive learning","pretraining"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["https://www.repository.cam.ac.uk/bitstreams/5a438c2c-912c-49a6-b87b-3b8475819040/download","http://purl.org/NET/rdflicense/allrightsreserved"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.17863/CAM.125995"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://www.repository.cam.ac.uk/bitstreams/364013b4-2964-4784-b6b7-2fda8d726151/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Machine learning with modern neural architectures, particularly large-scale language models, has enabled impressive progress across diverse problem domains such as natural language processing, graph representation learning, and tabular data analysis. However, these successes have predominantly relied on abundant, single-modal datasets, often overlooking the inherently multimodal and structurally complex nature of real-world data. Real-world data typically combines multiple modalities—text, graph structures, and tabular formats—each encoding distinct but complementary insights. Effectively integrating these heterogeneous sources remains one of the most significant open challenges in artificial intelligence research. In this thesis, I directly address this challenge by proposing three pioneering contributions that leverage multimodal data through sophisticated neural architectures and advanced representation learning techniques. First, I introduce HyperBERT, a novel neural architecture that integrates hypergraph neural networks with pretrained language models, achieving superior node classification performance on text-attributed hypergraphs. The second contribution tackles dense retrieval in environments with limited labeled data by developing a novel weakly supervised semantic distillation framework. Leveraging the rich semantic understanding embedded in large language models along with negative sampling strategies, this framework significantly improves retrieval accuracy and generalizability, particularly for tasks such as claim verification and evidence retrieval. The third contribution introduces TabMDA, a Transformer-based manifold data augmentation method for tabular data, leveraging pretrained Transformer encoders in combination with an advanced in-context subsetting technique. This approach generates synthetic samples that faithfully preserve data dependencies, enhancing downstream classification performance. Collectively, these interconnected contributions significantly advance the frontier of multimodal representation learning. Through the explicit modeling and integration of structural, textual, and tabular information, this thesis paves the way towards building more sophisticated, adaptable, and intelligent computational systems capable of nuanced understanding and interaction with complex real-world information structures."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["882fa6f774139abd9f016aabefeaba28","87eda9de84448d1f82354d60eee3eb5f"]},{"key":"dc:title","label":"Title","values":["Multimodal learning and language models for enhanced knowledge representations"]}]}],"canonical_facts":{"dc:contributor.advisor":["Micklem, Gos","Lio, Pietro"],"dc:creator":["Rodriguez Bazaga, Adrian"],"dc:date.issued":["2025-04-11"],"dc:description.abstract":["Machine learning with modern neural architectures, particularly large-scale language models, has enabled impressive progress across diverse problem domains such as natural language processing, graph representation learning, and tabular data analysis. However, these successes have predominantly relied on abundant, single-modal datasets, often overlooking the inherently multimodal and structurally complex nature of real-world data. Real-world data typically combines multiple modalities—text, graph structures, and tabular formats—each encoding distinct but complementary insights. Effectively integrating these heterogeneous sources remains one of the most significant open challenges in artificial intelligence research. In this thesis, I directly address this challenge by proposing three pioneering contributions that leverage multimodal data through sophisticated neural architectures and advanced representation learning techniques. First, I introduce HyperBERT, a novel neural architecture that integrates hypergraph neural networks with pretrained language models, achieving superior node classification performance on text-attributed hypergraphs. The second contribution tackles dense retrieval in environments with limited labeled data by developing a novel weakly supervised semantic distillation framework. Leveraging the rich semantic understanding embedded in large language models along with negative sampling strategies, this framework significantly improves retrieval accuracy and generalizability, particularly for tasks such as claim verification and evidence retrieval. The third contribution introduces TabMDA, a Transformer-based manifold data augmentation method for tabular data, leveraging pretrained Transformer encoders in combination with an advanced in-context subsetting technique. This approach generates synthetic samples that faithfully preserve data dependencies, enhancing downstream classification performance. Collectively, these interconnected contributions significantly advance the frontier of multimodal representation learning. Through the explicit modeling and integration of structural, textual, and tabular information, this thesis paves the way towards building more sophisticated, adaptable, and intelligent computational systems capable of nuanced understanding and interaction with complex real-world information structures."],"dc:format.checksum.md5":["882fa6f774139abd9f016aabefeaba28","87eda9de84448d1f82354d60eee3eb5f"],"dc:identifier.doi":["https://doi.org/10.17863/CAM.125995"],"dc:identifier.uri":["https://www.repository.cam.ac.uk/bitstreams/364013b4-2964-4784-b6b7-2fda8d726151/download"],"dc:language":["eng"],"dc:publisher.institution":["University of Cambridge"],"dc:relation.isreferencedby.uri":["https://www.repository.cam.ac.uk/handle/1810/396718"],"dc:rights":["https://www.repository.cam.ac.uk/bitstreams/5a438c2c-912c-49a6-b87b-3b8475819040/download","http://purl.org/NET/rdflicense/allrightsreserved"],"dc:subject":["machine learning","deep learning","language models","LLMs","graph neural networks","GNNs","multimodal learning","node classification","hypergraphs","dense retrieval","data augmentation","contrastive learning","pretraining"],"dc:title":["Multimodal learning and language models for enhanced knowledge representations"],"dc:type":["Thesis"],"dc:type.qualificationlevel":["Doctoral"],"dc:type.qualificationname":["Doctor of Philosophy (PhD)"]},"updated_at":"2026-07-22T22:23:59Z"}