Oxford Brookes University
Online Spatiotemporal Action Detection and Prediction via Causal Representations
Abstract
dc:descriptionIn this thesis, we focus on video action understanding problems from an online and real-time processing point of view. We start with the conversion of the traditional oine spatiotemporal action detection pipeline into an online spatiotemporal action tube detection system. An action tube is a set of bounding connected over time, which bounds an action instance in space and time. Next, we explore the future prediction capabilities of such detection methods by extending the an existing action tube into the future by regression. Later, we seek to establish that online/causal representations can achieve similar performance to that of oine three dimensional (3D) convolutional neural networks (CNNs) on various tasks, including action recognition, temporal action segmentation and early prediction. To this end, we propose various action tube detection approaches from either single or multiple frames. We start by introducing supervised action proposals for frame-level action detection and solving two energy optimisation formulations to detect the spatial and temporal boundaries of action tubes. Further, we propose an incremental tube construction algorithm to handle the online action detection problem. There, the real-time capabilities are made possible by introducing real-time frame-level action detection and real-time optical ow in the action detection pipeline for eciency. Next, we extend our frame-level approach to multiple frames with the help of a novel proposal to for predicting exible action 'micro-tubes' from a pair of frames. We extend the micro-tube prediction network in order to regress the future of each micro-tube, which is then fed to our proposed future action tube prediction framework. We convert 3D CNNs to causal 3D CNNs by replacing every 3D convolution with recurrent convolution, and by making use of sophisticated initialisation to handle the problems of recurrent modules. We show that our action tube detectors perform better than previous stateof- the-art methods, while exhibiting online and real-time capabilities. We evaluate each action tube detector and predictor on publicly available benchmarks to show the comparison with other state-of-the-art approaches. We also show that our exible micro-tube proposals not only improve action detection performance but can also handle sparse annotations. Finally, we demonstrate the causal capabilities of our causal 3D CNN.
Degree
thesis:*- Grantor dc:publisher
- Oxford Brookes University
- Year dc:date
- 2019
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Singh, Gurkirt
- Contributors dc:contributor
-
- Cuzzolin, Fabio
- Crook, Nigel
Rights
dc:rights- Statement dc:rights
-
- All rights reserved
- Language dc:language
- en
Identifiers
dc:identifier.*- DOI dc:identifier
- https://doi.org/10.24384/c7sk-6p02
- OAI identifier oai:identifier
- tle:62d12f94-2d53-469d-a917-9a46797eb393:d6bd9758-527a-46cd-bfe2-c433766e8fca:1