Back to search

Cal Poly

Adapting Single-View View Synthesis with Multiplane Images for 3D Video Chat

Abstract

dc:description.abstract

<p>Activities like one-on-one video chatting and video conferencing with multiple participants are more prevalent than ever today as we continue to tackle the pandemic. Bringing a 3D feel to video chat has always been a hot topic in Vision and Graphics communities. In this thesis, we have employed novel view synthesis in attempting to turn one-on-one video chatting into 3D. We have tuned the learning pipeline of Tucker and Snavely's <em>single-view</em> view synthesis <a href="https://single-view-mpi.github.io/" target="_blank">paper</a> — by retraining it on <em>MannequinChallenge</em> <a href="https://mannequin-depth.github.io/" target="_blank">dataset</a> — to better predict a layered representation of the scene viewed by either video chat participant at any given time. This intermediate representation of the local light field — called a <em>Multiplane Image</em> (MPI) — may then be used to rerender the scene at an arbitrary viewpoint which, in our case, would match with the head pose of the watcher in the opposite, concurrent video frame. We discuss that our pipeline, when implemented in real-time, would allow both video chat participants to unravel occluded scene content and "peer into" each other's dynamic video scenes to a certain extent. It would enable full parallax up to the baselines of small head rotations and/or translations. It would be similar to a VR headset's ability to determine the position and orientation of the wearer's head in 3D space and render any scene in alignment with this estimated head pose. We have attempted to improve the performance of the retrained model by extending MannequinChallenge with the much larger <em>RealEstate10K</em> <a href="https://tinghuiz.github.io/projects/mpi/" target="_blank">dataset</a>. We present a quantitative and qualitative comparison of the model variants and describe our impactful dataset curation process, among other aspects.</p>

Degree

thesis:*
Name thesis:degree_name
MS in Computer Science
Discipline thesis:degree_discipline
Computer Science
Year dc:date.available
2021

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Uppuluri, Anurag Venkata
Contributors dc:contributor
  • Jonathan Ventura
  • Computer Science
  • College of Engineering

Subjects

dc:subject × 7

Identifiers

dc:identifier.*
OAI identifier oai:identifier
oai:digitalcommons.calpoly.edu:theses-3987

Chain of custody

source
Harvested from
Cal Poly
Base URL
digitalcommons.calpoly.edu/do/oai/
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Uppuluri, Anurag Venkata. Adapting Single-View View Synthesis with Multiplane Images for 3D Video Chat. 2021. https://digitalcommons.calpoly.edu/theses/2572