{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/121988"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/121988","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Learning video representations with limited supervision","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-03-01 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2024-03-01 without embargo terms","abstract_has_math":false,"creators":["McKee, Daniel Benjamin"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Lazebnik, Svetlana","Forsyth, David","Hoiem, Derek","Tighe, Joseph"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2023,"date_issued":"2023-12","date_published":"2023-12","updated_at":"2026-07-22T22:25:00Z","subjects":["Computer Vision","Deep Learning","Video","Self-supervised Learning","Weakly Supervised Learning","Language Models","Multi-modal Models","Object Tracking"],"languages":["en","eng"],"rights":["Copyright 2023 Daniel McKee"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/121988","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Lazebnik, Svetlana","Forsyth, David","Hoiem, Derek","Tighe, Joseph"]},{"key":"dc:creator","label":"Author","values":["McKee, Daniel Benjamin"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2023-12","2023-11-22"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer Vision","Deep Learning","Video","Self-supervised Learning","Weakly Supervised Learning","Language Models","Multi-modal Models","Object Tracking"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2023 Daniel McKee"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/121988"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-03-01 without embargo terms","The student, Daniel McKee, accepted the attached license on 2023-11-19 at 23:21.","The student, Daniel McKee, submitted this Dissertation for approval on 2023-11-19 at 23:32.","This Dissertation was approved for publication on 2023-11-22 at 09:10.","DSpace SAF Submission Ingestion Package generated from Vireo submission #19925 on 2024-03-01 at 13:14:31","With the rapid growth of deep computer vision models, demand for large quantities of annotated data has risen higher than ever. Obtaining visual annotations, especially dense annotations requiring fine-grained localization of objects, is a costly and intensive process. Dense video tasks like tracking or video object segmentation provide an even greater annotation challenge due to the steep cost increase associated with labeling many individual frames. As a result, datasets for these tasks often lack the scale and diversity of samples in annotated image datasets. To combat such limitations, we investigate how we can take advantage of unlabeled videos, image annotations, and transfer of large-scale pretrained models to achieve effective performance on dense video tasks. First, we study representations for dense label propagation tasks in video, focusing on self-supervised approaches to learning temporal correspondence and comparing how image-trained models might be adapted for these tasks. Second, we investigate how to train a multi-object tracking model in the absence of tracking annotations. In place of fully supervised annotations, we demonstrate how to learn from unlabeled videos and videos that are hallucinated from annotated images using data augmentation techniques. Lastly, we explore a multi-modal problem setting where we wish to automatically recommend an audio soundtrack for an input video and text description of desired music. In this setting, we explore adapting large scale models like CLIP for joint modeling of video, text, and audio. We also investigate mechanisms for generating text pseudo-label descriptions for training using recent large language models."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Learning video representations with limited supervision"]}]}],"canonical_facts":{"dc:contributor":["Lazebnik, Svetlana","Forsyth, David","Hoiem, Derek","Tighe, Joseph"],"dc:creator":["McKee, Daniel Benjamin"],"dc:date":["2023-12","2023-11-22"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2024-03-01 without embargo terms","The student, Daniel McKee, accepted the attached license on 2023-11-19 at 23:21.","The student, Daniel McKee, submitted this Dissertation for approval on 2023-11-19 at 23:32.","This Dissertation was approved for publication on 2023-11-22 at 09:10.","DSpace SAF Submission Ingestion Package generated from Vireo submission #19925 on 2024-03-01 at 13:14:31","With the rapid growth of deep computer vision models, demand for large quantities of annotated data has risen higher than ever. Obtaining visual annotations, especially dense annotations requiring fine-grained localization of objects, is a costly and intensive process. Dense video tasks like tracking or video object segmentation provide an even greater annotation challenge due to the steep cost increase associated with labeling many individual frames. As a result, datasets for these tasks often lack the scale and diversity of samples in annotated image datasets. To combat such limitations, we investigate how we can take advantage of unlabeled videos, image annotations, and transfer of large-scale pretrained models to achieve effective performance on dense video tasks. First, we study representations for dense label propagation tasks in video, focusing on self-supervised approaches to learning temporal correspondence and comparing how image-trained models might be adapted for these tasks. Second, we investigate how to train a multi-object tracking model in the absence of tracking annotations. In place of fully supervised annotations, we demonstrate how to learn from unlabeled videos and videos that are hallucinated from annotated images using data augmentation techniques. Lastly, we explore a multi-modal problem setting where we wish to automatically recommend an audio soundtrack for an input video and text description of desired music. In this setting, we explore adapting large scale models like CLIP for joint modeling of video, text, and audio. We also investigate mechanisms for generating text pseudo-label descriptions for training using recent large language models."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/121988"],"dc:language":["en","eng"],"dc:rights":["Copyright 2023 Daniel McKee"],"dc:subject":["Computer Vision","Deep Learning","Video","Self-supervised Learning","Weakly Supervised Learning","Language Models","Multi-modal Models","Object Tracking"],"dc:title":["Learning video representations with limited supervision"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:00Z"}