University of Ontario Institute of Technology
Hybrid architecture for human action recognition using skeleton data
Abstract
dc:description.abstractIn this work, we propose a deep learning architecture, incorporating a Graph Convolutional Network (GCN) backbone combined with a partitioning transformer, that achieves results comparable to the state-of-the-art methods in skeleton based multi-person, multiview human action recognition. By leveraging attention-based GCN, the model captures context-dependent intrinsic topology while enhancing discriminative information. Furthermore, utilizing transformers, we harness their ability to aggregate long-range temporal information, allowing us to learn complex actions by attending to both short-term and long-term temporal windows. This is achieved through our partitioning strategy, which efficiently captures the relationships between neighboring and distant joints, enabling a comprehensive understanding of human movement dynamics. In this work, we also introduce a Cosine-based noise as a new data augmentation strategy for joints across time. This helps our model, Hybrid-Graformer, achieve accuracy comparable to the state-of-the-art across various skeleton-based action recognition benchmarks.
Degree
thesis:*- Name thesis:degree_name
- Master of Science (MSc)
- Discipline thesis:degree_discipline
- Computer Science
- Grantor
- University of Ontario Institute of Technology
- Year dc:date.issued
- 2024
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Nadeem, Muhammad Salik
- Advisor dc:contributor.advisor
-
- Qureshi, Faisal
Rights
- Language dc:language.iso
- en
Identifiers
dc:identifier.*- Handle dc:identifier.uri
- https://hdl.handle.net/10155/1901
- OAI identifier oai:identifier
- oai:ontariotechu.scholaris.ca:10155/1901