{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/121245"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/121245","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Video pretrained transformer with an ensemble of experts","abstract":"Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2025-08-01","abstract_html":"Submission published under a 24 month embargo labeled &#x27;Closed Access&#x27;, the embargo will last until 2025-08-01","abstract_has_math":false,"creators":["Christl, Daniel"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Ji, Heng"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2023,"date_issued":"2023-08","date_published":"2023-08","updated_at":"2026-07-22T22:24:57Z","subjects":["Transfer Learning","Multimodal Learning"],"languages":["en","eng"],"rights":["Copyright 2023 Daniel Christl"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/121245","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Ji, Heng"]},{"key":"dc:creator","label":"Author","values":["Christl, Daniel"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2023-08","2023-07-19"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Transfer Learning","Multimodal Learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2023 Daniel Christl"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/121245"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2025-08-01","The student, Daniel Christl, accepted the attached license on 2023-07-10 at 15:08.","The student, Daniel Christl, submitted this Thesis for approval on 2023-07-10 at 15:13.","This Thesis was approved for publication on 2023-07-19 at 08:30.","DSpace SAF Submission Ingestion Package generated from Vireo submission #19603 on 2023-12-04 at 17:33:04","I present VideoSemble, a novel multimodal decoder-only model capable of comprehending and generating outputs across diverse modalities, including text, audio, images, and scene graphs. By incorporating state-of-the-art pretrained encoders for each modality, the model demonstrates a deep understanding of the underlying context and relationships present in the data. To train the model, the extensive YouTube-1B corpus is leveraged, consisting of 20 million YouTube videos that provide rich, multimodal context. This pretraining objective focuses on autoregressive output generation and contrastive learning, utilizing the non-output modali- ties as context. This approach encourages the model to form meaningful connections between various modalities and develop a comprehensive understanding of the data. Following pretraining, the model is finetuned and evaluated on two benchmark datasets: the TV Question dataset, designed to assess multimodal question-answering capabilities, and the Kinetics-600 dataset, which measures action recognition and understanding in videos. The proposed model demonstrates inconsistent performances in both tasks, showcasing its potential ability to effectively synthesize information from multiple modalities and generate coherent, context-aware textual outputs, while also providing reservations about pretraining and finetuning methodologies utilized. The findings presented in this thesis contribute to the growing body of research in multi- modal understanding and generation, providing a robust and versatile framework for future exploration in the field. By combining state-of-the-art encoders with a decoder-only archi- tecture, VideoSemble offer new insights into the potential for deep learning models to grasp the complex interplay between modalities and generate meaningful outputs across a diverse range of contexts."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Video pretrained transformer with an ensemble of experts"]}]}],"canonical_facts":{"dc:contributor":["Ji, Heng"],"dc:creator":["Christl, Daniel"],"dc:date":["2023-08","2023-07-19"],"dc:description":["Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2025-08-01","The student, Daniel Christl, accepted the attached license on 2023-07-10 at 15:08.","The student, Daniel Christl, submitted this Thesis for approval on 2023-07-10 at 15:13.","This Thesis was approved for publication on 2023-07-19 at 08:30.","DSpace SAF Submission Ingestion Package generated from Vireo submission #19603 on 2023-12-04 at 17:33:04","I present VideoSemble, a novel multimodal decoder-only model capable of comprehending and generating outputs across diverse modalities, including text, audio, images, and scene graphs. By incorporating state-of-the-art pretrained encoders for each modality, the model demonstrates a deep understanding of the underlying context and relationships present in the data. To train the model, the extensive YouTube-1B corpus is leveraged, consisting of 20 million YouTube videos that provide rich, multimodal context. This pretraining objective focuses on autoregressive output generation and contrastive learning, utilizing the non-output modali- ties as context. This approach encourages the model to form meaningful connections between various modalities and develop a comprehensive understanding of the data. Following pretraining, the model is finetuned and evaluated on two benchmark datasets: the TV Question dataset, designed to assess multimodal question-answering capabilities, and the Kinetics-600 dataset, which measures action recognition and understanding in videos. The proposed model demonstrates inconsistent performances in both tasks, showcasing its potential ability to effectively synthesize information from multiple modalities and generate coherent, context-aware textual outputs, while also providing reservations about pretraining and finetuning methodologies utilized. The findings presented in this thesis contribute to the growing body of research in multi- modal understanding and generation, providing a robust and versatile framework for future exploration in the field. By combining state-of-the-art encoders with a decoder-only archi- tecture, VideoSemble offer new insights into the potential for deep learning models to grasp the complex interplay between modalities and generate meaningful outputs across a diverse range of contexts."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/121245"],"dc:language":["en","eng"],"dc:rights":["Copyright 2023 Daniel Christl"],"dc:subject":["Transfer Learning","Multimodal Learning"],"dc:title":["Video pretrained transformer with an ensemble of experts"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:57Z"}