Back to results

Massachusetts Institute of Technology

Large scale video action understanding

Abstract

dc:description.abstract

The goal of the project is to build a large scale video dataset called Moments, and train existing/novel models for action recognition. To aid automation of video collection and annotation selection, I trained Convolutional Neural Network models to estimate the likelihood of a desired action appearing in video clips. Selecting clips, which are highly probable to contain the wanted action, for annotation leads to a more efficient process overall with higher yield. Once a sizable dataset had been amassed, I investigated new multi-modal models that make use of different (spatial, temporal, auditory) signals in the video. I also conducted preliminary experiments into several promising directions that Moments opens up, including multi-label training. Lastly, I trained baseline models on Moments to calibrate the performance of existing techniques. Post-training, I diagnosed the shortcomings of the models and visualized videos that were found to be particularly difficult. I discovered that the difficulty largely arises due to the great variety in quality/perspective/subjects found in Moments videos. This highlights the challenging nature of the dataset and its value to the research community.

Degree

thesis:*
Department dc:contributor.department
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science.
Grantor dc:publisher
Massachusetts Institute of Technology
Year dc:date.issued
2017

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Yan, Tom, M. Eng. Massachusetts Institute of Technology
Advisor dc:contributor.advisor
  • Aude Oliva.

Subjects

dc:subject × 1

Rights

dc:rights
Statement dc:rights
  • MIT theses are protected by copyright. They may be viewed, downloaded, or printed from this source but further reproduction or distribution in any format is prohibited without written permission.
Language dc:language.iso
eng

Identifiers

dc:identifier.*
Handle dc:identifier.uri
http://hdl.handle.net/1721.1/119541
OAI identifier oai:identifier
oai:dspace.mit.edu:1721.1/119541

Chain of custody

source
Harvested from
MIT
Base URL
dspace.mit.edu/oai/request
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Yan, Tom, M. Eng. Massachusetts Institute of Technology. Large scale video action understanding. Massachusetts Institute of Technology, 2017. http://hdl.handle.net/1721.1/119541