Back to results

University of Cambridge

Generative representations of 2D and 3D visual content: semantics, geometry, and appearance

Abstract

dc:description.abstract

Generative modelling has transformed how visual data is represented, synthesized, and manipulated across computer vision and graphics. From images and 3D shapes to material appearance, generative models offer the potential to create content that is diverse, controllable, and physically realistic. Yet, major challenges remain: ensuring semantic alignment with human intent, incorporating geometric and structural priors, and enforcing physical plausibility in generated outputs. This thesis investigates generative representations of 2D and 3D visual content, with a focus on semantics, geometry, and appearance — three interconnected aspects that together underpin realistic and controllable generation. It is organized into three parts. The first part addresses text-guided image manipulation, analysing the limitations of CLIP-based approaches and introducing CLIP-PAE, a projection-augmentation embedding that enables more accurate, disentangled, and controllable editing. The second part focuses on 3D shape generation, proposing FrePolad, a frequency-rectified latent diffusion framework for efficient and flexible synthesis, and the Quartet of Diffusions, a structure-aware model that explicitly encodes part composition and symmetry for controllable shape generation. The third part explores neural appearance modelling, presenting M 3 ashy, a multimodal BRDF synthesis framework based on hyperdiffusion, and PBNBRDF, a physically based neural BRDF representation that enforces Helmholtz reciprocity and energy conservation for realistic material reconstruction and generation. Together, these contributions advance the state of generative modelling by combining expressiveness, interpretability, and physical consistency across different visual modalities. More broadly, this thesis moves towards the longer-term vision of generative systems that are not only visually convincing, but also semantically aligned, structurally coherent, and physically grounded, laying the foundation for immersive, editable, and richly structured digital worlds.

Degree

thesis:*
Name dc:type.qualificationname
Doctor of Philosophy (PhD)
Level dc:type.qualificationlevel
Doctoral
Grantor dc:publisher.institution
University of Cambridge
Year dc:date.issued
2025

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Zhou, Chenliang
Advisor dc:contributor.advisor
  • Oztireli, Cengiz

Subjects

dc:subject × 4

Rights

dc:rights
Language dc:language
eng

Identifiers

dc:identifier.*
DOI dc:identifier.doi
https://doi.org/10.17863/CAM.126935
OAI identifier oai:identifier
oai:www.repository.cam.ac.uk:1810/397968

Chain of custody

source
Harvested from
Cambridge University
Base URL
api.repository.cam.ac.uk/server/oai/request
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Zhou, Chenliang. Generative representations of 2D and 3D visual content: semantics, geometry, and appearance. Doctoral thesis, University of Cambridge, 2025. https://doi.org/10.17863/CAM.126935