{"id":{"repo_id":"wustl","oai_identifier":"oai:openscholarship.wustl.edu:eng_etds-1743"},"canonical_url":"https://search.dev.ndltd.org/etd/wustl/oai:openscholarship.wustl.edu:eng_etds-1743","repository":{"repo_id":"wustl","name":"Washington University in St. Louis","base_url":"https://openscholarship.wustl.edu/do/oai/"},"display":{"title":"NeVR: Learning Continuous Neural Video Representation with Local Feature Codes for Video Interpolation","abstract":"<p>Video frame interpolation aims to synthesis a non-exists intermediate frame guided by two successive frames. Recently, some work shows excellent results in learning continuous representation of temporally-varying 3D objects with neural field (NF), which could be used for interpolating the original video. However, these methods require several videos from different viewing angles, the information of camera poses, learning for each specific scene, and achieving sub-optimal results for video frame interpolation. To this end, we propose a new learning neural field representation-based model, Neural Video Representation (NeVR) to learn a continuous representation of videos for high-quality video interpolation. Unlike the traditional video interpolation algorithm, which directly synthesis the whole intermediate frame, our model aims to map the temporal-spatial coordinates of the queried pixels to the corresponding pixel value of the interpolated frame. Additionally, NeVR takes a latent feature code associated with queried pixels as input to enhance the image quality. That feature code contains the information of local implicit features and bilateral motion of the input frames and is obtained by a jointly trained encoder. Our experiments show that the proposed algorithm outperforms the state-of-the-art methods in video frame interpolation on several benchmark datasets.</p>","abstract_html":"&lt;p&gt;Video frame interpolation aims to synthesis a non-exists intermediate frame guided by two successive frames. Recently, some work shows excellent results in learning continuous representation of temporally-varying 3D objects with neural field (NF), which could be used for interpolating the original video. However, these methods require several videos from different viewing angles, the information of camera poses, learning for each specific scene, and achieving sub-optimal results for video frame interpolation. To this end, we propose a new learning neural field representation-based model, Neural Video Representation (NeVR) to learn a continuous representation of videos for high-quality video interpolation. Unlike the traditional video interpolation algorithm, which directly synthesis the whole intermediate frame, our model aims to map the temporal-spatial coordinates of the queried pixels to the corresponding pixel value of the interpolated frame. Additionally, NeVR takes a latent feature code associated with queried pixels as input to enhance the image quality. That feature code contains the information of local implicit features and bilateral motion of the input frames and is obtained by a jointly trained encoder. Our experiments show that the proposed algorithm outperforms the state-of-the-art methods in video frame interpolation on several benchmark datasets.&lt;/p&gt;","abstract_has_math":false,"creators":["Shangguan, Wentao"],"institution":null,"degree_name":"Master of Science (MS)","degree_level":"Dissertation","degree_discipline":"Electrical & Systems Engineering","degree_department":null,"school":null,"contributors":["Ulugbek Kamilov","Joseph A. O’Sullivan, Umberto Villa"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2021,"date_issued":"2021-12-01T08:00:00Z","date_published":"2021-12-01T08:00:00Z","updated_at":"2026-07-24T06:13:05Z","subjects":["Video Interpolation","Neural Field Representation","Video Enhancement","Engineering","Signal Processing"],"languages":["English (en)"],"rights":["I have not registered my thesis with the U.S. Copyright Office, and do not intend to."],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["https://openscholarship.wustl.edu/eng_etds/681"],"render_values":[{"text":"https://openscholarship.wustl.edu/eng_etds/681","href":"https://openscholarship.wustl.edu/eng_etds/681","code":true}]}]},"links":{"outbound_url":"https://doi.org/10.7936/r51t-mq19","outbound_label":"DOI","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Ulugbek Kamilov","Joseph A. O’Sullivan, Umberto Villa"]},{"key":"dc:creator","label":"Author","values":["Shangguan, Wentao"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2023-12-29T08:00:00Z"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Systems Engineering","McKelvey School of Engineering"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science (MS)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Video Interpolation","Neural Field Representation","Video Enhancement","Engineering","Signal Processing"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["English (en)"]},{"key":"dc:rights","label":"Dc Rights","values":["I have not registered my thesis with the U.S. Copyright Office, and do not intend to."]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://doi.org/10.7936/r51t-mq19","https://openscholarship.wustl.edu/eng_etds/681"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>Video frame interpolation aims to synthesis a non-exists intermediate frame guided by two successive frames. Recently, some work shows excellent results in learning continuous representation of temporally-varying 3D objects with neural field (NF), which could be used for interpolating the original video. However, these methods require several videos from different viewing angles, the information of camera poses, learning for each specific scene, and achieving sub-optimal results for video frame interpolation. To this end, we propose a new learning neural field representation-based model, Neural Video Representation (NeVR) to learn a continuous representation of videos for high-quality video interpolation. Unlike the traditional video interpolation algorithm, which directly synthesis the whole intermediate frame, our model aims to map the temporal-spatial coordinates of the queried pixels to the corresponding pixel value of the interpolated frame. Additionally, NeVR takes a latent feature code associated with queried pixels as input to enhance the image quality. That feature code contains the information of local implicit features and bilateral motion of the input frames and is obtained by a jointly trained encoder. Our experiments show that the proposed algorithm outperforms the state-of-the-art methods in video frame interpolation on several benchmark datasets.</p>"]},{"key":"dc:title","label":"Title","values":["NeVR: Learning Continuous Neural Video Representation with Local Feature Codes for Video Interpolation"]}]}],"canonical_facts":{"dc:contributor":["Ulugbek Kamilov","Joseph A. O’Sullivan, Umberto Villa"],"dc:creator":["Shangguan, Wentao"],"dc:date.available":["2023-12-29T08:00:00Z"],"dc:description.abstract":["<p>Video frame interpolation aims to synthesis a non-exists intermediate frame guided by two successive frames. Recently, some work shows excellent results in learning continuous representation of temporally-varying 3D objects with neural field (NF), which could be used for interpolating the original video. However, these methods require several videos from different viewing angles, the information of camera poses, learning for each specific scene, and achieving sub-optimal results for video frame interpolation. To this end, we propose a new learning neural field representation-based model, Neural Video Representation (NeVR) to learn a continuous representation of videos for high-quality video interpolation. Unlike the traditional video interpolation algorithm, which directly synthesis the whole intermediate frame, our model aims to map the temporal-spatial coordinates of the queried pixels to the corresponding pixel value of the interpolated frame. Additionally, NeVR takes a latent feature code associated with queried pixels as input to enhance the image quality. That feature code contains the information of local implicit features and bilateral motion of the input frames and is obtained by a jointly trained encoder. Our experiments show that the proposed algorithm outperforms the state-of-the-art methods in video frame interpolation on several benchmark datasets.</p>"],"dc:identifier":["https://doi.org/10.7936/r51t-mq19","https://openscholarship.wustl.edu/eng_etds/681"],"dc:language":["English (en)"],"dc:rights":["I have not registered my thesis with the U.S. Copyright Office, and do not intend to."],"dc:subject":["Video Interpolation","Neural Field Representation","Video Enhancement","Engineering","Signal Processing"],"dc:title":["NeVR: Learning Continuous Neural Video Representation with Local Feature Codes for Video Interpolation"],"thesis:degree_discipline":["Electrical & Systems Engineering","McKelvey School of Engineering"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Master of Science (MS)"]},"updated_at":"2026-07-24T06:13:05Z"}