{"id":{"repo_id":"calpoly","oai_identifier":"oai:digitalcommons.calpoly.edu:theses-3987"},"canonical_url":"https://search.dev.ndltd.org/etd/calpoly/oai:digitalcommons.calpoly.edu:theses-3987","repository":{"repo_id":"calpoly","name":"Cal Poly","base_url":"https://digitalcommons.calpoly.edu/do/oai/"},"display":{"title":"Adapting Single-View View Synthesis with Multiplane Images for 3D Video Chat","abstract":"<p>Activities like one-on-one video chatting and video conferencing with multiple participants are more prevalent than ever today as we continue to tackle the pandemic. Bringing a 3D feel to video chat has always been a hot topic in Vision and Graphics communities. In this thesis, we have employed novel view synthesis in attempting to turn one-on-one video chatting into 3D. We have tuned the learning pipeline of Tucker and Snavely's <em>single-view</em> view synthesis <a href=\"https://single-view-mpi.github.io/\" target=\"_blank\">paper</a> — by retraining it on <em>MannequinChallenge</em> <a href=\"https://mannequin-depth.github.io/\" target=\"_blank\">dataset</a> — to better predict a layered representation of the scene viewed by either video chat participant at any given time. This intermediate representation of the local light field — called a <em>Multiplane Image</em> (MPI) — may then be used to rerender the scene at an arbitrary viewpoint which, in our case, would match with the head pose of the watcher in the opposite, concurrent video frame. We discuss that our pipeline, when implemented in real-time, would allow both video chat participants to unravel occluded scene content and \"peer into\" each other's dynamic video scenes to a certain extent. It would enable full parallax up to the baselines of small head rotations and/or translations. It would be similar to a VR headset's ability to determine the position and orientation of the wearer's head in 3D space and render any scene in alignment with this estimated head pose. We have attempted to improve the performance of the retrained model by extending MannequinChallenge with the much larger <em>RealEstate10K</em> <a href=\"https://tinghuiz.github.io/projects/mpi/\" target=\"_blank\">dataset</a>. We present a quantitative and qualitative comparison of the model variants and describe our impactful dataset curation process, among other aspects.</p>","abstract_html":"&lt;p&gt;Activities like one-on-one video chatting and video conferencing with multiple participants are more prevalent than ever today as we continue to tackle the pandemic. Bringing a 3D feel to video chat has always been a hot topic in Vision and Graphics communities. In this thesis, we have employed novel view synthesis in attempting to turn one-on-one video chatting into 3D. We have tuned the learning pipeline of Tucker and Snavely&#x27;s &lt;em&gt;single-view&lt;/em&gt; view synthesis &lt;a href=&quot;https://single-view-mpi.github.io/&quot; target=&quot;_blank&quot;&gt;paper&lt;/a&gt; — by retraining it on &lt;em&gt;MannequinChallenge&lt;/em&gt; &lt;a href=&quot;https://mannequin-depth.github.io/&quot; target=&quot;_blank&quot;&gt;dataset&lt;/a&gt; — to better predict a layered representation of the scene viewed by either video chat participant at any given time. This intermediate representation of the local light field — called a &lt;em&gt;Multiplane Image&lt;/em&gt; (MPI) — may then be used to rerender the scene at an arbitrary viewpoint which, in our case, would match with the head pose of the watcher in the opposite, concurrent video frame. We discuss that our pipeline, when implemented in real-time, would allow both video chat participants to unravel occluded scene content and &quot;peer into&quot; each other&#x27;s dynamic video scenes to a certain extent. It would enable full parallax up to the baselines of small head rotations and/or translations. It would be similar to a VR headset&#x27;s ability to determine the position and orientation of the wearer&#x27;s head in 3D space and render any scene in alignment with this estimated head pose. We have attempted to improve the performance of the retrained model by extending MannequinChallenge with the much larger &lt;em&gt;RealEstate10K&lt;/em&gt; &lt;a href=&quot;https://tinghuiz.github.io/projects/mpi/&quot; target=&quot;_blank&quot;&gt;dataset&lt;/a&gt;. We present a quantitative and qualitative comparison of the model variants and describe our impactful dataset curation process, among other aspects.&lt;/p&gt;","abstract_has_math":false,"creators":["Uppuluri, Anurag Venkata"],"institution":null,"degree_name":"MS in Computer Science","degree_level":null,"degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Jonathan Ventura","Computer Science","College of Engineering"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2021,"date_issued":"2021-12-01T08:00:00Z","date_published":"2021-12-01T08:00:00Z","updated_at":"2026-07-24T01:32:37Z","subjects":["Virtual Reality","Neural Networks","Deep Learning","Image-Based Rendering","Computational Photography","3D Video Conferencing","Artificial Intelligence and Robotics"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["10.15368/theses.2021.165"],"render_values":[{"text":"10.15368/theses.2021.165","href":"https://doi.org/10.15368/theses.2021.165","code":true}]}]},"links":{"outbound_url":"https://digitalcommons.calpoly.edu/theses/2572","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Jonathan Ventura","Computer Science","College of Engineering"]},{"key":"dc:creator","label":"Author","values":["Uppuluri, Anurag Venkata"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2021-12-10T08:00:00Z"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_name","label":"Degree Name","values":["MS in Computer Science"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Virtual Reality","Neural Networks","Deep Learning","Image-Based Rendering","Computational Photography","3D Video Conferencing","Artificial Intelligence and Robotics"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://digitalcommons.calpoly.edu/theses/2572","10.15368/theses.2021.165"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>Activities like one-on-one video chatting and video conferencing with multiple participants are more prevalent than ever today as we continue to tackle the pandemic. Bringing a 3D feel to video chat has always been a hot topic in Vision and Graphics communities. In this thesis, we have employed novel view synthesis in attempting to turn one-on-one video chatting into 3D. We have tuned the learning pipeline of Tucker and Snavely's <em>single-view</em> view synthesis <a href=\"https://single-view-mpi.github.io/\" target=\"_blank\">paper</a> — by retraining it on <em>MannequinChallenge</em> <a href=\"https://mannequin-depth.github.io/\" target=\"_blank\">dataset</a> — to better predict a layered representation of the scene viewed by either video chat participant at any given time. This intermediate representation of the local light field — called a <em>Multiplane Image</em> (MPI) — may then be used to rerender the scene at an arbitrary viewpoint which, in our case, would match with the head pose of the watcher in the opposite, concurrent video frame. We discuss that our pipeline, when implemented in real-time, would allow both video chat participants to unravel occluded scene content and \"peer into\" each other's dynamic video scenes to a certain extent. It would enable full parallax up to the baselines of small head rotations and/or translations. It would be similar to a VR headset's ability to determine the position and orientation of the wearer's head in 3D space and render any scene in alignment with this estimated head pose. We have attempted to improve the performance of the retrained model by extending MannequinChallenge with the much larger <em>RealEstate10K</em> <a href=\"https://tinghuiz.github.io/projects/mpi/\" target=\"_blank\">dataset</a>. We present a quantitative and qualitative comparison of the model variants and describe our impactful dataset curation process, among other aspects.</p>"]},{"key":"dc:title","label":"Title","values":["Adapting Single-View View Synthesis with Multiplane Images for 3D Video Chat"]}]}],"canonical_facts":{"dc:contributor":["Jonathan Ventura","Computer Science","College of Engineering"],"dc:creator":["Uppuluri, Anurag Venkata"],"dc:date.available":["2021-12-10T08:00:00Z"],"dc:description.abstract":["<p>Activities like one-on-one video chatting and video conferencing with multiple participants are more prevalent than ever today as we continue to tackle the pandemic. Bringing a 3D feel to video chat has always been a hot topic in Vision and Graphics communities. In this thesis, we have employed novel view synthesis in attempting to turn one-on-one video chatting into 3D. We have tuned the learning pipeline of Tucker and Snavely's <em>single-view</em> view synthesis <a href=\"https://single-view-mpi.github.io/\" target=\"_blank\">paper</a> — by retraining it on <em>MannequinChallenge</em> <a href=\"https://mannequin-depth.github.io/\" target=\"_blank\">dataset</a> — to better predict a layered representation of the scene viewed by either video chat participant at any given time. This intermediate representation of the local light field — called a <em>Multiplane Image</em> (MPI) — may then be used to rerender the scene at an arbitrary viewpoint which, in our case, would match with the head pose of the watcher in the opposite, concurrent video frame. We discuss that our pipeline, when implemented in real-time, would allow both video chat participants to unravel occluded scene content and \"peer into\" each other's dynamic video scenes to a certain extent. It would enable full parallax up to the baselines of small head rotations and/or translations. It would be similar to a VR headset's ability to determine the position and orientation of the wearer's head in 3D space and render any scene in alignment with this estimated head pose. We have attempted to improve the performance of the retrained model by extending MannequinChallenge with the much larger <em>RealEstate10K</em> <a href=\"https://tinghuiz.github.io/projects/mpi/\" target=\"_blank\">dataset</a>. We present a quantitative and qualitative comparison of the model variants and describe our impactful dataset curation process, among other aspects.</p>"],"dc:identifier":["https://digitalcommons.calpoly.edu/theses/2572","10.15368/theses.2021.165"],"dc:subject":["Virtual Reality","Neural Networks","Deep Learning","Image-Based Rendering","Computational Photography","3D Video Conferencing","Artificial Intelligence and Robotics"],"dc:title":["Adapting Single-View View Synthesis with Multiplane Images for 3D Video Chat"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_name":["MS in Computer Science"]},"updated_at":"2026-07-24T01:32:37Z"}