University of Cambridge
Generalizable and Efficient Novel View Synthesis with Perceptual Foundations
Abstract
dc:description.abstractRecent advancements in novel view synthesis have demonstrated remarkable potential. Although these methods are capable of generating visually compelling novel views, their reliance on per-scene retraining severely limits their scalability and practical applications. The central goal of this dissertation is to transform novel view synthesis into a practical technology by developing generalizable models that are robust, efficient, dynamic-aware, and perceptually aligned. To achieve this objective, we introduce a series of integrated contributions that enhance generalizable novel view synthesis models by improving reconstruction robustness, boosting rendering efficiency, and incorporating support for dynamic scenes, while also providing a systematic analysis of the perceptual soundness of current techniques. These advancements aim to make novel view synthesis practical, enabling users to effortlessly capture, reconstruct, and interact with scenes in an efficient and realistic manner. To improve the robustness and realistic rendering quality of generalizable novel view synthesis, our first contribution enhances generalizable Neural Radiance Field (NeRF) architectures through the integration of a Mixture-of-View-Experts (MoE) paradigm. Our proposed model, GNT-MOVE, builds upon recent generalizable NeRF transformer by incorporating expert modules and geometry-aware consistency losses to balance overall model capacity with instance-specific specialization. This design boosts cross-scene generalization of the novel view synthesis model by a great margin. Aiming to address the need for real-time rendering of generalizable models, our second contribution, Evolutive Primitive Organization (EPO), directly predicts 3D Gaussian parameters in a feed-forward manner. By introducing learnable mechanisms for both growing and splitting of 3D Gaussians, EPO maintains end-to-end differentiability, ensuring gradient flow from the final training objective. This formulation not only supports predicting 3D-GS in a generalizable feed-forward manner, but also alleviates the local minimum issue in per-scene optimization setting. Recognizing the importance of handling dynamic content in real-world applications, we propose BTimer, the first motion-aware, generalizable feed-forward model for real-time reconstruction and novel view synthesis of dynamic scenes. BTimer reconstructs a full scene at a target timestamp by aggregating multi-frame context into a 3D-GS representation, achieving state-of-the-art performance with reconstruction times under 150 ms per frame. Finally, building on existing technical advancements, we consider whether these improvements truly enhance the overall user experience and for that, we present a comprehensive study on the perceptual quality of neural view synthesis methods. By building datasets of both controlled and in-the-wild scenes with corresponding reference videos, we systematically evaluate temporal artifacts and subtle distortions that conventional static image metrics may overlook. Our perceptual analysis yields insights and recommendations for improved dataset and metric selection in evaluating novel view synthesis techniques. Together, these contributions push the boundaries of generalizable novel view synthesis, offering a robust, efficient, and perceptually sound foundation for real-time applications in computer vision and computer graphics.
Degree
thesis:*- Name dc:type.qualificationname
- Doctor of Philosophy (PhD)
- Level dc:type.qualificationlevel
- Doctoral
- Grantor dc:publisher.institution
- University of Cambridge
- Year dc:date.issued
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Liang, Hanxue
- Advisors dc:contributor.advisor
-
- Oztireli, A Cengiz
- Mantiuk, Rafal
Subjects
dc:subject × 3Rights
dc:rightsIdentifiers
dc:identifier.*- DOI dc:identifier.doi
- https://doi.org/10.17863/CAM.122970
- OAI identifier oai:identifier
- oai:www.repository.cam.ac.uk:1810/392180