{"id":{"repo_id":"cambridge","oai_identifier":"oai:www.repository.cam.ac.uk:1810/397968"},"canonical_url":"https://search.dev.ndltd.org/etd/cambridge/oai:www.repository.cam.ac.uk:1810/397968","repository":{"repo_id":"cambridge","name":"Cambridge University","base_url":"https://api.repository.cam.ac.uk/server/oai/request"},"display":{"title":"Generative representations of 2D and 3D visual content: semantics, geometry, and appearance","abstract":"Generative modelling has transformed how visual data is represented, synthesized, and manipulated across computer vision and graphics. From images and 3D shapes to material appearance, generative models offer the potential to create content that is diverse, controllable, and physically realistic. Yet, major challenges remain: ensuring semantic alignment with human intent, incorporating geometric and structural priors, and enforcing physical plausibility in generated outputs. This thesis investigates generative representations of 2D and 3D visual content, with a focus on semantics, geometry, and appearance — three interconnected aspects that together underpin realistic and controllable generation. It is organized into three parts. The first part addresses text-guided image manipulation, analysing the limitations of CLIP-based approaches and introducing CLIP-PAE, a projection-augmentation embedding that enables more accurate, disentangled, and controllable editing. The second part focuses on 3D shape generation, proposing FrePolad, a frequency-rectified latent diffusion framework for eﬀicient and flexible synthesis, and the Quartet of Diffusions, a structure-aware model that explicitly encodes part composition and symmetry for controllable shape generation. The third part explores neural appearance modelling, presenting M 3 ashy, a multimodal BRDF synthesis framework based on hyperdiffusion, and PBNBRDF, a physically based neural BRDF representation that enforces Helmholtz reciprocity and energy conservation for realistic material reconstruction and generation. Together, these contributions advance the state of generative modelling by combining expressiveness, interpretability, and physical consistency across different visual modalities. More broadly, this thesis moves towards the longer-term vision of generative systems that are not only visually convincing, but also semantically aligned, structurally coherent, and physically grounded, laying the foundation for immersive, editable, and richly structured digital worlds.","abstract_html":"Generative modelling has transformed how visual data is represented, synthesized, and manipulated across computer vision and graphics. From images and 3D shapes to material appearance, generative models offer the potential to create content that is diverse, controllable, and physically realistic. Yet, major challenges remain: ensuring semantic alignment with human intent, incorporating geometric and structural priors, and enforcing physical plausibility in generated outputs. This thesis investigates generative representations of 2D and 3D visual content, with a focus on semantics, geometry, and appearance — three interconnected aspects that together underpin realistic and controllable generation. It is organized into three parts. The first part addresses text-guided image manipulation, analysing the limitations of CLIP-based approaches and introducing CLIP-PAE, a projection-augmentation embedding that enables more accurate, disentangled, and controllable editing. The second part focuses on 3D shape generation, proposing FrePolad, a frequency-rectified latent diffusion framework for eﬀicient and flexible synthesis, and the Quartet of Diffusions, a structure-aware model that explicitly encodes part composition and symmetry for controllable shape generation. The third part explores neural appearance modelling, presenting M 3 ashy, a multimodal BRDF synthesis framework based on hyperdiffusion, and PBNBRDF, a physically based neural BRDF representation that enforces Helmholtz reciprocity and energy conservation for realistic material reconstruction and generation. Together, these contributions advance the state of generative modelling by combining expressiveness, interpretability, and physical consistency across different visual modalities. More broadly, this thesis moves towards the longer-term vision of generative systems that are not only visually convincing, but also semantically aligned, structurally coherent, and physically grounded, laying the foundation for immersive, editable, and richly structured digital worlds.","abstract_has_math":false,"creators":["Zhou, Chenliang"],"institution":"University of Cambridge","degree_name":"Doctor of Philosophy (PhD)","degree_level":"Doctoral","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Oztireli, Cengiz"],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-09-01","date_published":"2025-09-01","updated_at":"2026-07-22T22:24:07Z","subjects":["machine learning","artificial intelligence","computer vision","computer graphics"],"languages":["eng"],"rights":[],"rights_urls":["https://www.repository.cam.ac.uk/bitstreams/10768240-819e-4ac2-aec1-db973bdb98b0/download","https://creativecommons.org/licenses/by-nc-sa/4.0/"],"identifier_entries":[]},"links":{"outbound_url":"https://doi.org/10.17863/CAM.126935","outbound_label":"DOI","outbound_source":"dc:identifier.doi"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Oztireli, Cengiz"]},{"key":"dc:creator","label":"Author","values":["Zhou, Chenliang"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2025-09-01"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["University of Cambridge"]},{"key":"dc:relation.isreferencedby.uri","label":"Dc Relation Isreferencedby URI","values":["https://www.repository.cam.ac.uk/handle/1810/397968"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["Doctoral"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["Doctor of Philosophy (PhD)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["machine learning","artificial intelligence","computer vision","computer graphics"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["https://www.repository.cam.ac.uk/bitstreams/10768240-819e-4ac2-aec1-db973bdb98b0/download","https://creativecommons.org/licenses/by-nc-sa/4.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.17863/CAM.126935"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://www.repository.cam.ac.uk/bitstreams/d23c1003-a61a-47c0-a074-06a75830d0ea/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Generative modelling has transformed how visual data is represented, synthesized, and manipulated across computer vision and graphics. From images and 3D shapes to material appearance, generative models offer the potential to create content that is diverse, controllable, and physically realistic. Yet, major challenges remain: ensuring semantic alignment with human intent, incorporating geometric and structural priors, and enforcing physical plausibility in generated outputs. This thesis investigates generative representations of 2D and 3D visual content, with a focus on semantics, geometry, and appearance — three interconnected aspects that together underpin realistic and controllable generation. It is organized into three parts. The first part addresses text-guided image manipulation, analysing the limitations of CLIP-based approaches and introducing CLIP-PAE, a projection-augmentation embedding that enables more accurate, disentangled, and controllable editing. The second part focuses on 3D shape generation, proposing FrePolad, a frequency-rectified latent diffusion framework for eﬀicient and flexible synthesis, and the Quartet of Diffusions, a structure-aware model that explicitly encodes part composition and symmetry for controllable shape generation. The third part explores neural appearance modelling, presenting M 3 ashy, a multimodal BRDF synthesis framework based on hyperdiffusion, and PBNBRDF, a physically based neural BRDF representation that enforces Helmholtz reciprocity and energy conservation for realistic material reconstruction and generation. Together, these contributions advance the state of generative modelling by combining expressiveness, interpretability, and physical consistency across different visual modalities. More broadly, this thesis moves towards the longer-term vision of generative systems that are not only visually convincing, but also semantically aligned, structurally coherent, and physically grounded, laying the foundation for immersive, editable, and richly structured digital worlds."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["4c61f4e813181fc18598e0ddbd7dd2c6","87eda9de84448d1f82354d60eee3eb5f"]},{"key":"dc:title","label":"Title","values":["Generative representations of 2D and 3D visual content: semantics, geometry, and appearance"]}]}],"canonical_facts":{"dc:contributor.advisor":["Oztireli, Cengiz"],"dc:creator":["Zhou, Chenliang"],"dc:date.issued":["2025-09-01"],"dc:description.abstract":["Generative modelling has transformed how visual data is represented, synthesized, and manipulated across computer vision and graphics. From images and 3D shapes to material appearance, generative models offer the potential to create content that is diverse, controllable, and physically realistic. Yet, major challenges remain: ensuring semantic alignment with human intent, incorporating geometric and structural priors, and enforcing physical plausibility in generated outputs. This thesis investigates generative representations of 2D and 3D visual content, with a focus on semantics, geometry, and appearance — three interconnected aspects that together underpin realistic and controllable generation. It is organized into three parts. The first part addresses text-guided image manipulation, analysing the limitations of CLIP-based approaches and introducing CLIP-PAE, a projection-augmentation embedding that enables more accurate, disentangled, and controllable editing. The second part focuses on 3D shape generation, proposing FrePolad, a frequency-rectified latent diffusion framework for eﬀicient and flexible synthesis, and the Quartet of Diffusions, a structure-aware model that explicitly encodes part composition and symmetry for controllable shape generation. The third part explores neural appearance modelling, presenting M 3 ashy, a multimodal BRDF synthesis framework based on hyperdiffusion, and PBNBRDF, a physically based neural BRDF representation that enforces Helmholtz reciprocity and energy conservation for realistic material reconstruction and generation. Together, these contributions advance the state of generative modelling by combining expressiveness, interpretability, and physical consistency across different visual modalities. More broadly, this thesis moves towards the longer-term vision of generative systems that are not only visually convincing, but also semantically aligned, structurally coherent, and physically grounded, laying the foundation for immersive, editable, and richly structured digital worlds."],"dc:format.checksum.md5":["4c61f4e813181fc18598e0ddbd7dd2c6","87eda9de84448d1f82354d60eee3eb5f"],"dc:identifier.doi":["https://doi.org/10.17863/CAM.126935"],"dc:identifier.uri":["https://www.repository.cam.ac.uk/bitstreams/d23c1003-a61a-47c0-a074-06a75830d0ea/download"],"dc:language":["eng"],"dc:publisher.institution":["University of Cambridge"],"dc:relation.isreferencedby.uri":["https://www.repository.cam.ac.uk/handle/1810/397968"],"dc:rights":["https://www.repository.cam.ac.uk/bitstreams/10768240-819e-4ac2-aec1-db973bdb98b0/download","https://creativecommons.org/licenses/by-nc-sa/4.0/"],"dc:subject":["machine learning","artificial intelligence","computer vision","computer graphics"],"dc:title":["Generative representations of 2D and 3D visual content: semantics, geometry, and appearance"],"dc:type":["Thesis"],"dc:type.qualificationlevel":["Doctoral"],"dc:type.qualificationname":["Doctor of Philosophy (PhD)"]},"updated_at":"2026-07-22T22:24:07Z"}