{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/105728"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/105728","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Self-supervised learning of spatiotemporal features from video colorization","abstract":"The student, Zubin Pahuja, accepted the attached license on 2019-07-19 at 13:10.","abstract_html":"The student, Zubin Pahuja, accepted the attached license on 2019-07-19 at 13:10.","abstract_has_math":false,"creators":["Pahuja, Zubin"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Forsyth, David A"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019-11-26T20:35:18Z","date_published":"2019-11-26T20:35:18Z","updated_at":"2026-07-22T22:24:45Z","subjects":["self-supervised learning","colorization","tracking","video"],"languages":["en"],"rights":["Copyright 2019 Zubin Pahuja"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/105728","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Forsyth, David A"]},{"key":"dc:creator","label":"Author","values":["Pahuja, Zubin"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2019-11-26T20:35:18Z","2019-07-19","2019-08"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["self-supervised learning","colorization","tracking","video"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2019 Zubin Pahuja"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/105728"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The student, Zubin Pahuja, accepted the attached license on 2019-07-19 at 13:10.","The student, Zubin Pahuja, submitted this Thesis for approval on 2019-07-19 at 13:18.","This Thesis was approved for publication on 2019-07-19 at 13:31.","DSpace SAF Submission Ingestion Package generated from Vireo submission #14395 on 2019-11-26 at 12:54:33","Made available in DSpace on 2019-11-26T20:35:18Z (GMT). No. of bitstreams: 2 PAHUJA-THESIS-2019.pdf: 4431266 bytes, checksum: b4d45542ac3dfdfd431c9a33498000b8 (MD5) LICENSE.txt: 4209 bytes, checksum: b126ca84b61ae45174be281af9b0ffc7 (MD5) Previous issue date: 2019-07-19","We pose video colorization as a self-supervised learning problem for visual tracking. We use large amounts of freely available unlabeled video from YouTube to learn colorization without explicit supervision. However, instead of predicting the color directly from the gray-scale frame, we constrain the model to solve this task by learning to copy colors from a reference frame. By equipping the model with a pointing mechanism into a reference frame, we learn an explicit spatiotemporal feature representation that can be used as a generic tracker for new tracking tasks without additional training or fine-tuning. Our self-supervised model can propagate any annotation from the first frame as a reference to the rest of the video. Experimental results suggest that the learned feature representations can be effectively transferred to video tracking and object segmentation tasks. We perform extensive quantitative and qualitative evaluations on the DAVIS-2017 video object segmentation dataset and demonstrate significant improvements over the baseline. Although the model is trained without any ground-truth labels, our method learns to track well enough to outperform the latest methods based on optical flow. Since annotating videos is expensive and tracking has many applications in robotics and graphics, we believe learning to track with self-supervision can have a large impact. More broadly, we show that the features learned from a task for which cheap training data is readily available can be used to learn a task which would otherwise require an expensive, large-scale dataset with minimal supervision. Thus, we hope our results encourage a broader exploration in the promising field of self-supervised learning.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2019-11-26 without embargo terms"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Self-supervised learning of spatiotemporal features from video colorization"]}]}],"canonical_facts":{"dc:contributor":["Forsyth, David A"],"dc:creator":["Pahuja, Zubin"],"dc:date":["2019-11-26T20:35:18Z","2019-07-19","2019-08"],"dc:description":["The student, Zubin Pahuja, accepted the attached license on 2019-07-19 at 13:10.","The student, Zubin Pahuja, submitted this Thesis for approval on 2019-07-19 at 13:18.","This Thesis was approved for publication on 2019-07-19 at 13:31.","DSpace SAF Submission Ingestion Package generated from Vireo submission #14395 on 2019-11-26 at 12:54:33","Made available in DSpace on 2019-11-26T20:35:18Z (GMT). No. of bitstreams: 2 PAHUJA-THESIS-2019.pdf: 4431266 bytes, checksum: b4d45542ac3dfdfd431c9a33498000b8 (MD5) LICENSE.txt: 4209 bytes, checksum: b126ca84b61ae45174be281af9b0ffc7 (MD5) Previous issue date: 2019-07-19","We pose video colorization as a self-supervised learning problem for visual tracking. We use large amounts of freely available unlabeled video from YouTube to learn colorization without explicit supervision. However, instead of predicting the color directly from the gray-scale frame, we constrain the model to solve this task by learning to copy colors from a reference frame. By equipping the model with a pointing mechanism into a reference frame, we learn an explicit spatiotemporal feature representation that can be used as a generic tracker for new tracking tasks without additional training or fine-tuning. Our self-supervised model can propagate any annotation from the first frame as a reference to the rest of the video. Experimental results suggest that the learned feature representations can be effectively transferred to video tracking and object segmentation tasks. We perform extensive quantitative and qualitative evaluations on the DAVIS-2017 video object segmentation dataset and demonstrate significant improvements over the baseline. Although the model is trained without any ground-truth labels, our method learns to track well enough to outperform the latest methods based on optical flow. Since annotating videos is expensive and tracking has many applications in robotics and graphics, we believe learning to track with self-supervision can have a large impact. More broadly, we show that the features learned from a task for which cheap training data is readily available can be used to learn a task which would otherwise require an expensive, large-scale dataset with minimal supervision. Thus, we hope our results encourage a broader exploration in the promising field of self-supervised learning.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2019-11-26 without embargo terms"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/105728"],"dc:language":["en"],"dc:rights":["Copyright 2019 Zubin Pahuja"],"dc:subject":["self-supervised learning","colorization","tracking","video"],"dc:title":["Self-supervised learning of spatiotemporal features from video colorization"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:45Z"}