University of Illinois at Urbana-Champaign
Video pretrained transformer with an ensemble of experts
Abstract
dc:descriptionI present VideoSemble, a novel multimodal decoder-only model capable of comprehending and generating outputs across diverse modalities, including text, audio, images, and scene graphs. By incorporating state-of-the-art pretrained encoders for each modality, the model demonstrates a deep understanding of the underlying context and relationships present in the data. To train the model, the extensive YouTube-1B corpus is leveraged, consisting of 20 million YouTube videos that provide rich, multimodal context. This pretraining objective focuses on autoregressive output generation and contrastive learning, utilizing the non-output modali- ties as context. This approach encourages the model to form meaningful connections between various modalities and develop a comprehensive understanding of the data. Following pretraining, the model is finetuned and evaluated on two benchmark datasets: the TV Question dataset, designed to assess multimodal question-answering capabilities, and the Kinetics-600 dataset, which measures action recognition and understanding in videos. The proposed model demonstrates inconsistent performances in both tasks, showcasing its potential ability to effectively synthesize information from multiple modalities and generate coherent, context-aware textual outputs, while also providing reservations about pretraining and finetuning methodologies utilized. The findings presented in this thesis contribute to the growing body of research in multi- modal understanding and generation, providing a robust and versatile framework for future exploration in the field. By combining state-of-the-art encoders with a decoder-only archi- tecture, VideoSemble offer new insights into the potential for deep learning models to grasp the complex interplay between modalities and generate meaningful outputs across a diverse range of contexts.
Degree
thesis:*- Name thesis:degree_name
- M.S.
- Level thesis:degree_level
- Thesis
- Discipline thesis:degree_discipline
- Computer Science
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2023
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Christl, Daniel
- Contributors dc:contributor
-
- Ji, Heng
Subjects
dc:subject × 2Rights
dc:rights- Statement dc:rights
-
- Copyright 2023 Daniel Christl
- Language dc:language
- en, eng
Identifiers
dc:identifier.*- Handle dc:identifier
- https://hdl.handle.net/2142/121245