Back to results

Massachusetts Institute of Technology

Text-Free Audio Captions of Short Videos from Latent Space Representation

Abstract

dc:description.abstract

In this thesis, we re-implement previous work exploring image to speech captioning. We expand upon the work to implement video to speech captioning. Specifically, we implement a text-free image to speech captioning pipeline that integrates four distinct machine learning models. We alter the models to process video data rather than image data and analyze the resulting speech captions. We conduct experiments on the Wav2Vec2 and HuBERT Automatic Speech Recognition models, and identify which works best with synthesized speech.

Degree

thesis:*
Name thesis:degree_name
Master
Department dc:contributor.department
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
Grantor dc:publisher
Massachusetts Institute of Technology
Year dc:date.issued
2022

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Agarwal, Anisha
Advisor dc:contributor.advisor
  • Oliva, Aude

Rights

dc:rights
Statement dc:rights
  • In Copyright - Educational Use Permitted
  • Copyright MIT

Identifiers

dc:identifier.*
Handle dc:identifier.uri
https://hdl.handle.net/1721.1/144873
OAI identifier oai:identifier
oai:dspace.mit.edu:1721.1/144873

Chain of custody

source
Harvested from
MIT
Base URL
dspace.mit.edu/oai/request
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
related terms
citation

Agarwal, Anisha. Text-Free Audio Captions of Short Videos from Latent Space Representation. Massachusetts Institute of Technology, 2022. https://hdl.handle.net/1721.1/144873