University of Cambridge
Generative representations of 2D and 3D visual content: semantics, geometry, and appearance
Abstract
dc:description.abstractGenerative modelling has transformed how visual data is represented, synthesized, and manipulated across computer vision and graphics. From images and 3D shapes to material appearance, generative models offer the potential to create content that is diverse, controllable, and physically realistic. Yet, major challenges remain: ensuring semantic alignment with human intent, incorporating geometric and structural priors, and enforcing physical plausibility in generated outputs. This thesis investigates generative representations of 2D and 3D visual content, with a focus on semantics, geometry, and appearance — three interconnected aspects that together underpin realistic and controllable generation. It is organized into three parts. The first part addresses text-guided image manipulation, analysing the limitations of CLIP-based approaches and introducing CLIP-PAE, a projection-augmentation embedding that enables more accurate, disentangled, and controllable editing. The second part focuses on 3D shape generation, proposing FrePolad, a frequency-rectified latent diffusion framework for efficient and flexible synthesis, and the Quartet of Diffusions, a structure-aware model that explicitly encodes part composition and symmetry for controllable shape generation. The third part explores neural appearance modelling, presenting M 3 ashy, a multimodal BRDF synthesis framework based on hyperdiffusion, and PBNBRDF, a physically based neural BRDF representation that enforces Helmholtz reciprocity and energy conservation for realistic material reconstruction and generation. Together, these contributions advance the state of generative modelling by combining expressiveness, interpretability, and physical consistency across different visual modalities. More broadly, this thesis moves towards the longer-term vision of generative systems that are not only visually convincing, but also semantically aligned, structurally coherent, and physically grounded, laying the foundation for immersive, editable, and richly structured digital worlds.
Degree
thesis:*- Name dc:type.qualificationname
- Doctor of Philosophy (PhD)
- Level dc:type.qualificationlevel
- Doctoral
- Grantor dc:publisher.institution
- University of Cambridge
- Year dc:date.issued
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Zhou, Chenliang
- Advisor dc:contributor.advisor
-
- Oztireli, Cengiz
Subjects
dc:subject × 4Rights
dc:rightsIdentifiers
dc:identifier.*- DOI dc:identifier.doi
- https://doi.org/10.17863/CAM.126935
- OAI identifier oai:identifier
- oai:www.repository.cam.ac.uk:1810/397968