{"id":{"repo_id":"vt","oai_identifier":"oai:vtechworks.lib.vt.edu:10919/79702"},"canonical_url":"https://search.dev.ndltd.org/etd/vt/oai:vtechworks.lib.vt.edu:10919/79702","repository":{"repo_id":"vt","name":"Virginia Tech","base_url":"https://vtechworks.lib.vt.edu/oai/request"},"display":{"title":"Unsupervised Learning of Spatiotemporal Features by Video Completion","abstract":"In this work, we present an unsupervised representation learning approach for learning rich spatiotemporal features from videos without the supervision from semantic labels. We propose to learn the spatiotemporal features by training a 3D convolutional neural network (CNN) using video completion as a surrogate task. Using a large collection of unlabeled videos, we train the CNN to predict the missing pixels of a spatiotemporal hole given the remaining parts of the video through minimizing per-pixel reconstruction loss. To achieve good reconstruction results using color videos, the CNN needs to have a certain level of understanding of the scene dynamics and predict plausible, temporally coherent contents. We further explore to jointly reconstruct both color frames and flow fields. By exploiting the statistical temporal structure of images, we show that the learned representations capture meaningful spatiotemporal structures from raw videos. We validate the effectiveness of our approach for CNN pre-training on action recognition and action similarity labeling problems. Our quantitative results demonstrate that our method compares favorably against learning without external data and existing unsupervised learning approaches.","abstract_html":"In this work, we present an unsupervised representation learning approach for learning rich spatiotemporal features from videos without the supervision from semantic labels. We propose to learn the spatiotemporal features by training a 3D convolutional neural network (CNN) using video completion as a surrogate task. Using a large collection of unlabeled videos, we train the CNN to predict the missing pixels of a spatiotemporal hole given the remaining parts of the video through minimizing per-pixel reconstruction loss. To achieve good reconstruction results using color videos, the CNN needs to have a certain level of understanding of the scene dynamics and predict plausible, temporally coherent contents. We further explore to jointly reconstruct both color frames and flow fields. By exploiting the statistical temporal structure of images, we show that the learned representations capture meaningful spatiotemporal structures from raw videos. We validate the effectiveness of our approach for CNN pre-training on action recognition and action similarity labeling problems. Our quantitative results demonstrate that our method compares favorably against learning without external data and existing unsupervised learning approaches.","abstract_has_math":false,"creators":["Nallabolu, Adithya Reddy"],"institution":"Virginia Tech","degree_name":"Master of Science","degree_level":"masters","degree_discipline":"Computer Engineering","degree_department":"Electrical and Computer Engineering","school":null,"contributors":[],"advisors":[],"committee_chairs":["Kochersberger, Kevin B.","Huang, Jia-Bin"],"committee_members":["Dhillon, Harpreet Singh"],"year":2017,"date_issued":"2017-10-18","date_published":"2017-10-18","updated_at":"2026-07-22T22:20:27Z","subjects":["Representation Learning","Supervised","Unsupervised"],"languages":[],"rights":["In Copyright"],"rights_urls":["http://rightsstatements.org/vocab/InC/1.0/"],"identifier_entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:12668"],"render_values":[{"text":"vt_gsexam:12668","href":null,"code":true}]}]},"links":{"outbound_url":"http://hdl.handle.net/10919/79702","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.committeechair","label":"Committee Chair","values":["Kochersberger, Kevin B.","Huang, Jia-Bin"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Dhillon, Harpreet Singh"]},{"key":"dc:contributor.department","label":"Department","values":["Electrical and Computer Engineering"]},{"key":"dc:creator","label":"Author","values":["Nallabolu, Adithya Reddy"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2017-10-19T08:00:43Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2017-10-19T08:00:43Z"]},{"key":"dc:date.issued","label":"Date","values":["2017-10-18"]},{"key":"dc:publisher","label":"Institution","values":["Virginia Tech"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Engineering"]},{"key":"thesis:degree_level","label":"Degree Level","values":["masters"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Virginia Polytechnic Institute and State University"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Representation Learning","Supervised","Unsupervised"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["In Copyright"]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://rightsstatements.org/vocab/InC/1.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:12668"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["http://hdl.handle.net/10919/79702"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["In this work, we present an unsupervised representation learning approach for learning rich spatiotemporal features from videos without the supervision from semantic labels. We propose to learn the spatiotemporal features by training a 3D convolutional neural network (CNN) using video completion as a surrogate task. Using a large collection of unlabeled videos, we train the CNN to predict the missing pixels of a spatiotemporal hole given the remaining parts of the video through minimizing per-pixel reconstruction loss. To achieve good reconstruction results using color videos, the CNN needs to have a certain level of understanding of the scene dynamics and predict plausible, temporally coherent contents. We further explore to jointly reconstruct both color frames and flow fields. By exploiting the statistical temporal structure of images, we show that the learned representations capture meaningful spatiotemporal structures from raw videos. We validate the effectiveness of our approach for CNN pre-training on action recognition and action similarity labeling problems. Our quantitative results demonstrate that our method compares favorably against learning without external data and existing unsupervised learning approaches."]},{"key":"dc:description.abstractgeneral","label":"General Abstract","values":["The current supervised representation learning methods leverage large datasets of millions of labeled examples to learn semantically meaningful visual representations. Thousands of boring human hours are spent on manually labeling these datasets. But, do we need semantically labeled images to learn good visual representation? Humans learn visual representations using little or no semantic supervision but the existing approaches are mostly supervised. In this work, we propose an unsupervised visual representation learning algorithm to learn useful spatiotemporal features by formulating a video completion problem. To predict the missing pixels of the video, the model needs to have a high-level semantic understanding and motion patterns of people and objects. We demonstrate that video completion task effectively learns semantically meaningful spatiotemporal features from raw natural videos without semantic labels. The learned representation provide a good network weight initialization for applications with few training examples. We show significant performance gain over training the model from scratch and demonstrate improved performance in action recognition and action similarity labeling tasks when compared with competitive unsupervised learning algorithms."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["Master of Science"]},{"key":"dc:format.medium","label":"Dc Format Medium","values":["ETD"]},{"key":"dc:title","label":"Title","values":["Unsupervised Learning of Spatiotemporal Features by Video Completion"]}]}],"canonical_facts":{"dc:contributor.committeechair":["Kochersberger, Kevin B.","Huang, Jia-Bin"],"dc:contributor.committeemember":["Dhillon, Harpreet Singh"],"dc:contributor.department":["Electrical and Computer Engineering"],"dc:creator":["Nallabolu, Adithya Reddy"],"dc:date.accessioned":["2017-10-19T08:00:43Z"],"dc:date.available":["2017-10-19T08:00:43Z"],"dc:date.issued":["2017-10-18"],"dc:description.abstract":["In this work, we present an unsupervised representation learning approach for learning rich spatiotemporal features from videos without the supervision from semantic labels. We propose to learn the spatiotemporal features by training a 3D convolutional neural network (CNN) using video completion as a surrogate task. Using a large collection of unlabeled videos, we train the CNN to predict the missing pixels of a spatiotemporal hole given the remaining parts of the video through minimizing per-pixel reconstruction loss. To achieve good reconstruction results using color videos, the CNN needs to have a certain level of understanding of the scene dynamics and predict plausible, temporally coherent contents. We further explore to jointly reconstruct both color frames and flow fields. By exploiting the statistical temporal structure of images, we show that the learned representations capture meaningful spatiotemporal structures from raw videos. We validate the effectiveness of our approach for CNN pre-training on action recognition and action similarity labeling problems. Our quantitative results demonstrate that our method compares favorably against learning without external data and existing unsupervised learning approaches."],"dc:description.abstractgeneral":["The current supervised representation learning methods leverage large datasets of millions of labeled examples to learn semantically meaningful visual representations. Thousands of boring human hours are spent on manually labeling these datasets. But, do we need semantically labeled images to learn good visual representation? Humans learn visual representations using little or no semantic supervision but the existing approaches are mostly supervised. In this work, we propose an unsupervised visual representation learning algorithm to learn useful spatiotemporal features by formulating a video completion problem. To predict the missing pixels of the video, the model needs to have a high-level semantic understanding and motion patterns of people and objects. We demonstrate that video completion task effectively learns semantically meaningful spatiotemporal features from raw natural videos without semantic labels. The learned representation provide a good network weight initialization for applications with few training examples. We show significant performance gain over training the model from scratch and demonstrate improved performance in action recognition and action similarity labeling tasks when compared with competitive unsupervised learning algorithms."],"dc:description.degree":["Master of Science"],"dc:format.medium":["ETD"],"dc:identifier.other":["vt_gsexam:12668"],"dc:identifier.uri":["http://hdl.handle.net/10919/79702"],"dc:publisher":["Virginia Tech"],"dc:rights":["In Copyright"],"dc:rights.uri":["http://rightsstatements.org/vocab/InC/1.0/"],"dc:subject":["Representation Learning","Supervised","Unsupervised"],"dc:title":["Unsupervised Learning of Spatiotemporal Features by Video Completion"],"dc:type":["Thesis"],"thesis:degree_discipline":["Computer Engineering"],"thesis:degree_level":["masters"],"thesis:degree_name":["Master of Science"],"thesis:institution_name":["Virginia Polytechnic Institute and State University"]},"updated_at":"2026-07-22T22:20:27Z"}