Back to results

University of Illinois at Urbana-Champaign

Video pretrained transformer with an ensemble of experts

Abstract

dc:description

I present VideoSemble, a novel multimodal decoder-only model capable of comprehending and generating outputs across diverse modalities, including text, audio, images, and scene graphs. By incorporating state-of-the-art pretrained encoders for each modality, the model demonstrates a deep understanding of the underlying context and relationships present in the data. To train the model, the extensive YouTube-1B corpus is leveraged, consisting of 20 million YouTube videos that provide rich, multimodal context. This pretraining objective focuses on autoregressive output generation and contrastive learning, utilizing the non-output modali- ties as context. This approach encourages the model to form meaningful connections between various modalities and develop a comprehensive understanding of the data. Following pretraining, the model is finetuned and evaluated on two benchmark datasets: the TV Question dataset, designed to assess multimodal question-answering capabilities, and the Kinetics-600 dataset, which measures action recognition and understanding in videos. The proposed model demonstrates inconsistent performances in both tasks, showcasing its potential ability to effectively synthesize information from multiple modalities and generate coherent, context-aware textual outputs, while also providing reservations about pretraining and finetuning methodologies utilized. The findings presented in this thesis contribute to the growing body of research in multi- modal understanding and generation, providing a robust and versatile framework for future exploration in the field. By combining state-of-the-art encoders with a decoder-only archi- tecture, VideoSemble offer new insights into the potential for deep learning models to grasp the complex interplay between modalities and generate meaningful outputs across a diverse range of contexts.

Degree

thesis:*
Name thesis:degree_name
M.S.
Level thesis:degree_level
Thesis
Discipline thesis:degree_discipline
Computer Science
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2023

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Christl, Daniel
Contributors dc:contributor
  • Ji, Heng

Subjects

dc:subject × 2

Rights

dc:rights
Statement dc:rights
  • Copyright 2023 Daniel Christl
Language dc:language
en, eng

Identifiers

dc:identifier.*
Handle dc:identifier
https://hdl.handle.net/2142/121245

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Christl, Daniel. Video pretrained transformer with an ensemble of experts. Thesis thesis, University of Illinois at Urbana-Champaign, 2023. https://hdl.handle.net/2142/121245