Texas A&M University
Embodied Interaction in Virtual Meetings: From Static Gesture Recognition to Few-Shot Gesture Adaptation
Abstract
dc:description.abstractAs remote collaboration becomes a standard mode of work, existing virtual meeting platforms often fail to capture the embodied cues of in-person interactions, particularly hand gestures that convey intent or feedback. Grounded in theories of embodiment, this dissertation explores hand gesture-based interaction as an essential channel for real-time virtual meetings, assuming that users should be able to interact in digital spaces in ways that reflect real-world behaviors. The research aims to: (1) design an intuitive gesture input method for real-time question-response interaction; (2) implement a gesture-based interface that restores spontaneity and clarity in group collaboration; and (3) evaluate a few-shot recognition model that can adapt to new, user-defined gestures with minimal examples. A gesture elicitation study identified intuitive response gestures suitable for large-group virtual settings. Based on these findings, a static gesture recognition interface was implemented using consumer augmented reality tools to detect finger-count gestures (one to five) in real time, enabling natural responses to binary and multiple-choice questions within Zoom. Recognizing variability in gesture use across individuals, cultures, and contexts, the study also investigated rapid adaptation to new gesture sets. A few-shot learning framework was implemented using the publicly available Jester dataset, which contains 27 gesture classes. Experiments with Prototypical Networks, using both convolutional neural network (CNN) and spatio-temporal graph convolutional network (ST-GCN) backbones, achieved approximately 75% accuracy in 5-way 5-shot settings. These results demonstrate the potential to expand gesture vocabularies quickly and effectively with limited training data. These contributions advance the development of adaptive, expressive, and scalable gesture-based interfaces for more human-centered virtual collaboration. By integrating intuitive gesture elicitation, real-time recognition, and few-shot learning, the work lays the foundation for virtual meeting systems that more closely mirror the fluid, multimodal communication of in-person interactions.
Degree
thesis:*- Name thesis:degree_name
- Doctor of Philosophy
- Level thesis:degree_level
- Doctoral
- Discipline thesis:degree_discipline
- Computer Science
- Grantor
- Texas A&M University
- Year dc:date.issued
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Koh, Jung In
- Advisor dc:contributor.advisor
-
- Hammond, Tracy
- Committee members dc:contributor.committeemember
-
- Goldberg, Daniel
- Ioerger, Thomas
- Chaspari, Theodora
Subjects
dc:subject × 1Rights
- Language dc:language.iso
- English
Identifiers
dc:identifier.*- Handle dc:identifier.uri
- https://hdl.handle.net/1969.1/1599940