Back to results

Virginia Tech

Bridging Multimodal Learning and Planning for Intelligent Task Assistance

Abstract

dc:description.abstract

Task-assistance systems provide adaptive, multimodal guidance for complex, step-based activities such as cooking and DIY projects. A central challenge lies in enabling these systems to interpret real-world scenarios—understanding user intent from verbal, visual, or textual cues and generating coherent, multimodal instructions enriched with relevant visual. To tackle this, modern systems leverage advanced machine learning techniques, from representation learning that processes information from diverse modalities (e.g., text, images, audio) to procedural planning which provides dynamic, context-driven guidance, enabling systems to provide precise, real-time assistance tailored to user needs. This work addresses core challenges in representation learning and multimodal planning through three key contributions. First, we introduce a modality-agnostic contrastive learning framework that optimizes negative sample selection by jointly balancing anchor similarity, influence and diversity, improving generalization across vision, language, and graph tasks. Second, we propose a tuning strategy for masked audio models that leverages unsupervised audio mixtures to enhance adaptation to downstream tasks with less labeled data, such as few-shot learning. Third, we present a zero-shot framework for generating multimodal procedural plans with explicit object-state consistency, paired with two novel evaluation metrics and an evaluation task to assess planning accuracy, cross-modal alignment and temporal coherence. These contributions are integrated into a context-aware multimodal task assistant, empirically validated through real-world user studies. Our work establishes a foundation for more robust, adaptable, and user-centric task-assistance systems, bridging critical gaps in multimodal understanding and guidance.

Degree

thesis:*
Name thesis:degree_name
Doctor of Philosophy
Level thesis:degree_level
doctoral
Discipline thesis:degree_discipline
Computer Science & Applications
Department dc:contributor.department
Computer Science and#38; Applications
Grantor dc:publisher
Virginia Tech
Year dc:date.issued
2025

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Tabassum, Afrina
Chairs dc:contributor.committeechair
  • Eldardiry, Hoda Mohamed
  • Lourentzou, Ismini
Committee members dc:contributor.committeemember
  • Jin, Ran
  • Thomas, Christopher Lee
  • Huang, Jia-Bin

Subjects

dc:subject × 2

Rights

dc:rights
Statement dc:rights
  • In Copyright
Language dc:language.iso
en

Identifiers

dc:identifier.*
Dc Identifier Other
vt_gsexam:42996
OAI identifier oai:identifier
oai:vtechworks.lib.vt.edu:10919/132474

Chain of custody

source
Harvested from
Virginia Tech
Base URL
vtechworks.lib.vt.edu/oai/request
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Tabassum, Afrina. Bridging Multimodal Learning and Planning for Intelligent Task Assistance. doctoral thesis, Virginia Tech, 2025. https://hdl.handle.net/10919/132474