National University of Singapore
A multi-resolution multi-source and multi-modal (M3) transductive framework for concept detection in news video
Abstract
dc:description.abstractWe study the problem of detecting concepts in news video. Most existing algorithms for news video concept detection are based on single-resolution (shot), single source (training data), and multi-modal fusion methods under a supervised inductive inference framework. In this thesis, we present a novel multi-resolution, multi-source and multi-modal transductive learning framework. As different modal features only work well in different temporal resolutions and different resolutions exhibit different types of semantics, we perform a multi-resolution analysis at the shot, multimedia discourse and story levels to capture the semantics. Our multi-source inference model makes use of the knowledge not only from training data but also from other online information resources. We perform transductive inference to better capture the distributions of data from both the test and specific training cases to train the classifiers. We test our framework in the TRECVID 2004 dataset. Experimental results demonstrate that our approach is effective.
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- WANG GANG