{"id":{"repo_id":"tamu","oai_identifier":"oai:oaktrust.library.tamu.edu:1969.1/1599940"},"canonical_url":"https://search.dev.ndltd.org/etd/tamu/oai:oaktrust.library.tamu.edu:1969.1/1599940","repository":{"repo_id":"tamu","name":"Texas A&M University","base_url":"https://oaktrust.library.tamu.edu/server/oai/request"},"display":{"title":"Embodied Interaction in Virtual Meetings: From Static Gesture Recognition to Few-Shot Gesture Adaptation","abstract":"As remote collaboration becomes a standard mode of work, existing virtual meeting platforms often fail to capture the embodied cues of in-person interactions, particularly hand gestures that convey intent or feedback. Grounded in theories of embodiment, this dissertation explores hand gesture-based interaction as an essential channel for real-time virtual meetings, assuming that users should be able to interact in digital spaces in ways that reflect real-world behaviors. The research aims to: (1) design an intuitive gesture input method for real-time question-response interaction; (2) implement a gesture-based interface that restores spontaneity and clarity in group collaboration; and (3) evaluate a few-shot recognition model that can adapt to new, user-defined gestures with minimal examples. A gesture elicitation study identified intuitive response gestures suitable for large-group virtual settings. Based on these findings, a static gesture recognition interface was implemented using consumer augmented reality tools to detect finger-count gestures (one to five) in real time, enabling natural responses to binary and multiple-choice questions within Zoom. Recognizing variability in gesture use across individuals, cultures, and contexts, the study also investigated rapid adaptation to new gesture sets. A few-shot learning framework was implemented using the publicly available Jester dataset, which contains 27 gesture classes. Experiments with Prototypical Networks, using both convolutional neural network (CNN) and spatio-temporal graph convolutional network (ST-GCN) backbones, achieved approximately 75% accuracy in 5-way 5-shot settings. These results demonstrate the potential to expand gesture vocabularies quickly and effectively with limited training data. These contributions advance the development of adaptive, expressive, and scalable gesture-based interfaces for more human-centered virtual collaboration. By integrating intuitive gesture elicitation, real-time recognition, and few-shot learning, the work lays the foundation for virtual meeting systems that more closely mirror the fluid, multimodal communication of in-person interactions.","abstract_html":"As remote collaboration becomes a standard mode of work, existing virtual meeting platforms often fail to capture the embodied cues of in-person interactions, particularly hand gestures that convey intent or feedback. Grounded in theories of embodiment, this dissertation explores hand gesture-based interaction as an essential channel for real-time virtual meetings, assuming that users should be able to interact in digital spaces in ways that reflect real-world behaviors. The research aims to: (1) design an intuitive gesture input method for real-time question-response interaction; (2) implement a gesture-based interface that restores spontaneity and clarity in group collaboration; and (3) evaluate a few-shot recognition model that can adapt to new, user-defined gestures with minimal examples. A gesture elicitation study identified intuitive response gestures suitable for large-group virtual settings. Based on these findings, a static gesture recognition interface was implemented using consumer augmented reality tools to detect finger-count gestures (one to five) in real time, enabling natural responses to binary and multiple-choice questions within Zoom. Recognizing variability in gesture use across individuals, cultures, and contexts, the study also investigated rapid adaptation to new gesture sets. A few-shot learning framework was implemented using the publicly available Jester dataset, which contains 27 gesture classes. Experiments with Prototypical Networks, using both convolutional neural network (CNN) and spatio-temporal graph convolutional network (ST-GCN) backbones, achieved approximately 75% accuracy in 5-way 5-shot settings. These results demonstrate the potential to expand gesture vocabularies quickly and effectively with limited training data. These contributions advance the development of adaptive, expressive, and scalable gesture-based interfaces for more human-centered virtual collaboration. By integrating intuitive gesture elicitation, real-time recognition, and few-shot learning, the work lays the foundation for virtual meeting systems that more closely mirror the fluid, multimodal communication of in-person interactions.","abstract_has_math":false,"creators":["Koh, Jung In"],"institution":"Texas A&M University","degree_name":"Doctor of Philosophy","degree_level":"Doctoral","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":[],"advisors":["Hammond, Tracy"],"committee_chairs":[],"committee_members":["Goldberg, Daniel","Ioerger, Thomas","Chaspari, Theodora"],"year":2025,"date_issued":"2025-12","date_published":"2025-12","updated_at":"2026-08-21T16:48:44Z","subjects":["Computer Science"],"languages":["English"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/1969.1/1599940","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"source_record":{"url":"https://oaktrust.library.tamu.edu/server/oai/request?verb=GetRecord&metadataPrefix=dim&identifier=oai%3Aoaktrust.library.tamu.edu%3A1969.1%2F1599940","prefix":"dim"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Hammond, Tracy"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Goldberg, Daniel","Ioerger, Thomas","Chaspari, Theodora"]},{"key":"dc:creator","label":"Author","values":["Koh, Jung In"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2026-03-05T21:05:33Z"]},{"key":"dc:date.issued","label":"Date","values":["2025-12"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Doctoral"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctor of Philosophy"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Texas A&M University"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer Science"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["English"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/1969.1/1599940"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["As remote collaboration becomes a standard mode of work, existing virtual meeting platforms often fail to capture the embodied cues of in-person interactions, particularly hand gestures that convey intent or feedback. Grounded in theories of embodiment, this dissertation explores hand gesture-based interaction as an essential channel for real-time virtual meetings, assuming that users should be able to interact in digital spaces in ways that reflect real-world behaviors. The research aims to: (1) design an intuitive gesture input method for real-time question-response interaction; (2) implement a gesture-based interface that restores spontaneity and clarity in group collaboration; and (3) evaluate a few-shot recognition model that can adapt to new, user-defined gestures with minimal examples. A gesture elicitation study identified intuitive response gestures suitable for large-group virtual settings. Based on these findings, a static gesture recognition interface was implemented using consumer augmented reality tools to detect finger-count gestures (one to five) in real time, enabling natural responses to binary and multiple-choice questions within Zoom. Recognizing variability in gesture use across individuals, cultures, and contexts, the study also investigated rapid adaptation to new gesture sets. A few-shot learning framework was implemented using the publicly available Jester dataset, which contains 27 gesture classes. Experiments with Prototypical Networks, using both convolutional neural network (CNN) and spatio-temporal graph convolutional network (ST-GCN) backbones, achieved approximately 75% accuracy in 5-way 5-shot settings. These results demonstrate the potential to expand gesture vocabularies quickly and effectively with limited training data. These contributions advance the development of adaptive, expressive, and scalable gesture-based interfaces for more human-centered virtual collaboration. By integrating intuitive gesture elicitation, real-time recognition, and few-shot learning, the work lays the foundation for virtual meeting systems that more closely mirror the fluid, multimodal communication of in-person interactions."]},{"key":"dc:format.mimetype","label":"Dc Format Mimetype","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Embodied Interaction in Virtual Meetings: From Static Gesture Recognition to Few-Shot Gesture Adaptation"]}]}],"canonical_facts":{"dc:contributor.advisor":["Hammond, Tracy"],"dc:contributor.committeemember":["Goldberg, Daniel","Ioerger, Thomas","Chaspari, Theodora"],"dc:creator":["Koh, Jung In"],"dc:date.accessioned":["2026-03-05T21:05:33Z"],"dc:date.issued":["2025-12"],"dc:description.abstract":["As remote collaboration becomes a standard mode of work, existing virtual meeting platforms often fail to capture the embodied cues of in-person interactions, particularly hand gestures that convey intent or feedback. Grounded in theories of embodiment, this dissertation explores hand gesture-based interaction as an essential channel for real-time virtual meetings, assuming that users should be able to interact in digital spaces in ways that reflect real-world behaviors. The research aims to: (1) design an intuitive gesture input method for real-time question-response interaction; (2) implement a gesture-based interface that restores spontaneity and clarity in group collaboration; and (3) evaluate a few-shot recognition model that can adapt to new, user-defined gestures with minimal examples. A gesture elicitation study identified intuitive response gestures suitable for large-group virtual settings. Based on these findings, a static gesture recognition interface was implemented using consumer augmented reality tools to detect finger-count gestures (one to five) in real time, enabling natural responses to binary and multiple-choice questions within Zoom. Recognizing variability in gesture use across individuals, cultures, and contexts, the study also investigated rapid adaptation to new gesture sets. A few-shot learning framework was implemented using the publicly available Jester dataset, which contains 27 gesture classes. Experiments with Prototypical Networks, using both convolutional neural network (CNN) and spatio-temporal graph convolutional network (ST-GCN) backbones, achieved approximately 75% accuracy in 5-way 5-shot settings. These results demonstrate the potential to expand gesture vocabularies quickly and effectively with limited training data. These contributions advance the development of adaptive, expressive, and scalable gesture-based interfaces for more human-centered virtual collaboration. By integrating intuitive gesture elicitation, real-time recognition, and few-shot learning, the work lays the foundation for virtual meeting systems that more closely mirror the fluid, multimodal communication of in-person interactions."],"dc:format.mimetype":["application/pdf"],"dc:identifier.uri":["https://hdl.handle.net/1969.1/1599940"],"dc:language.iso":["English"],"dc:subject":["Computer Science"],"dc:title":["Embodied Interaction in Virtual Meetings: From Static Gesture Recognition to Few-Shot Gesture Adaptation"],"dc:type":["Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Doctoral"],"thesis:degree_name":["Doctor of Philosophy"],"thesis:institution_name":["Texas A&M University"]},"updated_at":"2026-08-21T16:48:44Z"}