Back to results

West Virginia University

Human Interaction Recognition with Audio and Visual Cues

Abstract

dc:description.abstract

The automated recognition of human activities from video is a fundamental problem with applications in several areas, ranging from video surveillance, and robotics, to smart healthcare, and multimedia indexing and retrieval, just to mention a few. However, the pervasive diffusion of cameras capable of recording audio also makes available to those applications a complementary modality. Despite the sizable progress made in the area of modeling and recognizing group activities, and actions performed by people in isolation from video, the availability of audio cues has rarely being leveraged. This is even more so in the area of modeling and recognizing binary interactions between humans, where also the use of video has been limited.;This thesis introduces a modeling framework for binary human interactions based on audio and visual cues. The main idea is to describe an interaction with a spatio-temporal trajectory modeling the visual motion cues, and a temporal trajectory modeling the audio cues. This poses the problem of how to fuse temporal trajectories from multiple modalities for the purpose of recognition. We propose a solution whereby trajectories are modeled as the output of kernel state space models. Then, we developed kernel-based methods for the audio-visual fusion that act at the feature level, as well as at the kernel level, by exploiting multiple kernel learning techniques. The approaches have been extensively tested and evaluated with a dataset made of videos obtained from TV shows and Hollywood movies, containing five different interactions. The results show the promise of this approach by producing a significant improvement of the recognition rate when audio cues are exploited, clearly setting the state-of-the-art in this particular application.

Degree

thesis:*
Name thesis:degree_name
MS
Level thesis:degree_level
Thesis
Discipline thesis:degree_discipline
Lane Department of Computer Science and Electrical Engineering
Year dc:date.available
2014

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Almohsen, Ranya
Contributors dc:contributor
  • Gianfranco Doretto
  • Hani Ammar
  • Natalia Schmie

Subjects

dc:subject × 1

Identifiers

dc:identifier.*
OAI identifier oai:identifier
oai:researchrepository.wvu.edu:etd-1536

Chain of custody

source
Harvested from
West Virginia University
Base URL
researchrepository.wvu.edu/do/oai/
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Almohsen, Ranya. Human Interaction Recognition with Audio and Visual Cues. Thesis thesis, 2014. https://doi.org/10.33915/etd.533