Back to search

University of Illinois at Urbana-Champaign

Learning video representations with limited supervision

Abstract

dc:description

With the rapid growth of deep computer vision models, demand for large quantities of annotated data has risen higher than ever. Obtaining visual annotations, especially dense annotations requiring fine-grained localization of objects, is a costly and intensive process. Dense video tasks like tracking or video object segmentation provide an even greater annotation challenge due to the steep cost increase associated with labeling many individual frames. As a result, datasets for these tasks often lack the scale and diversity of samples in annotated image datasets. To combat such limitations, we investigate how we can take advantage of unlabeled videos, image annotations, and transfer of large-scale pretrained models to achieve effective performance on dense video tasks. First, we study representations for dense label propagation tasks in video, focusing on self-supervised approaches to learning temporal correspondence and comparing how image-trained models might be adapted for these tasks. Second, we investigate how to train a multi-object tracking model in the absence of tracking annotations. In place of fully supervised annotations, we demonstrate how to learn from unlabeled videos and videos that are hallucinated from annotated images using data augmentation techniques. Lastly, we explore a multi-modal problem setting where we wish to automatically recommend an audio soundtrack for an input video and text description of desired music. In this setting, we explore adapting large scale models like CLIP for joint modeling of video, text, and audio. We also investigate mechanisms for generating text pseudo-label descriptions for training using recent large language models.

Degree

thesis:*
Name thesis:degree_name
Ph.D.
Level thesis:degree_level
Dissertation
Discipline thesis:degree_discipline
Computer Science
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2023

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • McKee, Daniel Benjamin
Contributors dc:contributor
  • Lazebnik, Svetlana
  • Forsyth, David
  • Hoiem, Derek
  • Tighe, Joseph

Subjects

dc:subject × 8

Rights

dc:rights
Statement dc:rights
  • Copyright 2023 Daniel McKee
Language dc:language
en, eng

Identifiers

dc:identifier.*
Handle dc:identifier
https://hdl.handle.net/2142/121988

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

McKee, Daniel Benjamin. Learning video representations with limited supervision. Dissertation thesis, University of Illinois at Urbana-Champaign, 2023. https://hdl.handle.net/2142/121988