University of Cambridge
Principled Methods for Advancing Generative Machine Learning
Abstract
dc:description.abstractWith the aim of generating novel data beyond collected samples, generative machine learning (ML) has become one of the most actively explored research areas in recent years, witnessing a series of groundbreaking advances. In computer vision (CV), diffusion models mark a major milestone — they can synthesize photo-realistic and diverse images, outperforming previous state-of-the-art approaches such as generative adversarial networks (GANs). In computer graphics, Gaussian splatting achieves high-fidelity novel view synthesis, with significantly higher efficiency than methods like neural radiance fields (NeRFs). Despite their remarkable success on standard benchmarks (e.g., text-to-image generation), current generative ML models often perform poorly in more challenging scenarios that violate their implicit mathematical assumptions or demand the incorporation of physical constraints. In some cases, these scenarios even lie beyond the models’ scopes of applicability. This thesis identifies several such settings, aiming to understand why existing models fail and how to overcome these limitations in a principled manner. For instance, diffusion models are highly sensitive to noise, and their performance relies heavily on the availability of clean data. Consequently, their generation quality on non-image domains (e.g., medical time series) remains significantly lower than on natural images. To address this issue, we analyse its underlying causes and propose a new class of stochastic differential equations (SDEs) — risk-sensitive SDEs — as the backbone of continuous-time diffusion models. This approach is principled and accommodates various noise assumptions. Another example explored in this thesis concerns neural operators in scientific ML. While these data-driven models efficiently solve partial differential equations (PDEs), they often violate fundamental conservation laws (e.g., mass conservation) that are essential for describing real-world physical systems. To resolve this, we introduce a collection of local correction operators with learnable interpolation coefficients. Theoretically, we establish that this interpolation ensures consistency with key conservation laws while producing PDE solutions that are, in principle, more accurate than those from previous methods in terms of squared error. Beyond diffusion models and neural operators, we also investigate Gaussian splatting as an important generative model for 3D reconstruction. By analyzing the geometric structure of the surfaces formed by Gaussian primitives, we derive structural properties that can serve as constraints for optimizing Gaussian splatting. This geometry-aware approach substantially mitigates performance degradation in low-resource settings. Finally, the thesis concludes by summarizing our studies and discussing potential future directions, particularly in the era of large language models (LLMs). The rich priors that LLMs capture from massive internet-scale data may help overcome some inherent limitations of the settings we studied.
Degree
thesis:*- Name dc:type.qualificationname
- Doctor of Philosophy (PhD)
- Level dc:type.qualificationlevel
- Doctoral
- Grantor dc:publisher.institution
- University of Cambridge
- Year dc:date.issued
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Li, Yangming
- Advisor dc:contributor.advisor
-
- Schönlieb, Carola-Bibiane
Subjects
dc:subject × 1Rights
dc:rightsIdentifiers
dc:identifier.*- DOI dc:identifier.doi
- https://doi.org/10.17863/CAM.130189
- OAI identifier oai:identifier
- oai:www.repository.cam.ac.uk:1810/403117