Back to results

University of Cambridge

Generalizable and Efficient Novel View Synthesis with Perceptual Foundations

Abstract

dc:description.abstract

Recent advancements in novel view synthesis have demonstrated remarkable potential. Although these methods are capable of generating visually compelling novel views, their reliance on per-scene retraining severely limits their scalability and practical applications. The central goal of this dissertation is to transform novel view synthesis into a practical technology by developing generalizable models that are robust, efficient, dynamic-aware, and perceptually aligned. To achieve this objective, we introduce a series of integrated contributions that enhance generalizable novel view synthesis models by improving reconstruction robustness, boosting rendering efficiency, and incorporating support for dynamic scenes, while also providing a systematic analysis of the perceptual soundness of current techniques. These advancements aim to make novel view synthesis practical, enabling users to effortlessly capture, reconstruct, and interact with scenes in an efficient and realistic manner. To improve the robustness and realistic rendering quality of generalizable novel view synthesis, our first contribution enhances generalizable Neural Radiance Field (NeRF) architectures through the integration of a Mixture-of-View-Experts (MoE) paradigm. Our proposed model, GNT-MOVE, builds upon recent generalizable NeRF transformer by incorporating expert modules and geometry-aware consistency losses to balance overall model capacity with instance-specific specialization. This design boosts cross-scene generalization of the novel view synthesis model by a great margin. Aiming to address the need for real-time rendering of generalizable models, our second contribution, Evolutive Primitive Organization (EPO), directly predicts 3D Gaussian parameters in a feed-forward manner. By introducing learnable mechanisms for both growing and splitting of 3D Gaussians, EPO maintains end-to-end differentiability, ensuring gradient flow from the final training objective. This formulation not only supports predicting 3D-GS in a generalizable feed-forward manner, but also alleviates the local minimum issue in per-scene optimization setting. Recognizing the importance of handling dynamic content in real-world applications, we propose BTimer, the first motion-aware, generalizable feed-forward model for real-time reconstruction and novel view synthesis of dynamic scenes. BTimer reconstructs a full scene at a target timestamp by aggregating multi-frame context into a 3D-GS representation, achieving state-of-the-art performance with reconstruction times under 150 ms per frame. Finally, building on existing technical advancements, we consider whether these improvements truly enhance the overall user experience and for that, we present a comprehensive study on the perceptual quality of neural view synthesis methods. By building datasets of both controlled and in-the-wild scenes with corresponding reference videos, we systematically evaluate temporal artifacts and subtle distortions that conventional static image metrics may overlook. Our perceptual analysis yields insights and recommendations for improved dataset and metric selection in evaluating novel view synthesis techniques. Together, these contributions push the boundaries of generalizable novel view synthesis, offering a robust, efficient, and perceptually sound foundation for real-time applications in computer vision and computer graphics.

Degree

thesis:*
Name dc:type.qualificationname
Doctor of Philosophy (PhD)
Level dc:type.qualificationlevel
Doctoral
Grantor dc:publisher.institution
University of Cambridge
Year dc:date.issued
2025

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Liang, Hanxue
Advisors dc:contributor.advisor
  • Oztireli, A Cengiz
  • Mantiuk, Rafal

Subjects

dc:subject × 3

Rights

dc:rights
Language dc:language
eng

Identifiers

dc:identifier.*
DOI dc:identifier.doi
https://doi.org/10.17863/CAM.122970
OAI identifier oai:identifier
oai:www.repository.cam.ac.uk:1810/392180

Chain of custody

source
Harvested from
Cambridge University
Base URL
api.repository.cam.ac.uk/server/oai/request
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Liang, Hanxue. Generalizable and Efficient Novel View Synthesis with Perceptual Foundations. Doctoral thesis, University of Cambridge, 2025. https://doi.org/10.17863/CAM.122970