{"id":{"repo_id":"vt","oai_identifier":"oai:vtechworks.lib.vt.edu:10919/119417"},"canonical_url":"https://search.dev.ndltd.org/etd/vt/oai:vtechworks.lib.vt.edu:10919/119417","repository":{"repo_id":"vt","name":"Virginia Tech","base_url":"https://vtechworks.lib.vt.edu/oai/request"},"display":{"title":"General-Purpose Task Guidance from Natural Language in Augmented Reality using Vision-Language Models","abstract":"Augmented reality task guidance systems provide assistance for procedural tasks, which require a sequence of physical actions, by rendering virtual guidance visuals within the real-world environment. An example of such a task would be to secure two wood parts together, which could display guidance visuals indicating the user to pick up a drill and drill each screw. Current AR task guidance systems are limited in that they require AR system experts for use, require CAD models of real-world objects, or only function for limited types of tasks or environments. We propose a general-purpose AR task guidance approach and proof-of-concept system to generate guidance for tasks defined by natural language. Our approach allows an operator to take pictures of relevant objects and write task instructions for an end user, which are used by the system to determine where to place guidance visuals. Then, an end user can receive and follow guidance even if objects change location or environment. Guidance includes reusable visuals that display generic actions, such as our system's 3D hand animations. Our approach utilizes current vision-language machine learning models for text and image semantic understanding and object localization. We built a proof-of-concept system using our approach and tested its accuracy and usability in a user study. We found that all operators were able to generate clear guidance for tasks in an office room, and end users were able to follow the guidance visuals to complete the expected action 85.7% of the time without knowledge of their tasks. Participants rated that our system was easy to use to generate guidance visuals they expected.","abstract_html":"Augmented reality task guidance systems provide assistance for procedural tasks, which require a sequence of physical actions, by rendering virtual guidance visuals within the real-world environment. An example of such a task would be to secure two wood parts together, which could display guidance visuals indicating the user to pick up a drill and drill each screw. Current AR task guidance systems are limited in that they require AR system experts for use, require CAD models of real-world objects, or only function for limited types of tasks or environments. We propose a general-purpose AR task guidance approach and proof-of-concept system to generate guidance for tasks defined by natural language. Our approach allows an operator to take pictures of relevant objects and write task instructions for an end user, which are used by the system to determine where to place guidance visuals. Then, an end user can receive and follow guidance even if objects change location or environment. Guidance includes reusable visuals that display generic actions, such as our system&#x27;s 3D hand animations. Our approach utilizes current vision-language machine learning models for text and image semantic understanding and object localization. We built a proof-of-concept system using our approach and tested its accuracy and usability in a user study. We found that all operators were able to generate clear guidance for tasks in an office room, and end users were able to follow the guidance visuals to complete the expected action 85.7% of the time without knowledge of their tasks. Participants rated that our system was easy to use to generate guidance visuals they expected.","abstract_has_math":false,"creators":["Stover, Daniel James"],"institution":"Virginia Tech","degree_name":"Master of Science","degree_level":"masters","degree_discipline":"Computer Engineering","degree_department":"Electrical and Computer Engineering","school":null,"contributors":[],"advisors":[],"committee_chairs":["Abbott, Amos L.","Bowman, Douglas A."],"committee_members":["Thomas, Christopher Lee","Jones, Creed Farris"],"year":2024,"date_issued":"2024-06-12","date_published":"2024-06-12","updated_at":"2026-07-22T22:20:27Z","subjects":["Augmented Reality","Machine Learning","Task Guidance"],"languages":["en"],"rights":["Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International"],"rights_urls":["http://creativecommons.org/licenses/by-nc-sa/4.0/"],"identifier_entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:40807"],"render_values":[{"text":"vt_gsexam:40807","href":null,"code":true}]}]},"links":{"outbound_url":"https://hdl.handle.net/10919/119417","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.committeechair","label":"Committee Chair","values":["Abbott, Amos L.","Bowman, Douglas A."]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Thomas, Christopher Lee","Jones, Creed Farris"]},{"key":"dc:contributor.department","label":"Department","values":["Electrical and Computer Engineering"]},{"key":"dc:creator","label":"Author","values":["Stover, Daniel James"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2024-06-13T08:01:42Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2024-06-13T08:01:42Z"]},{"key":"dc:date.issued","label":"Date","values":["2024-06-12"]},{"key":"dc:publisher","label":"Institution","values":["Virginia Tech"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Engineering"]},{"key":"thesis:degree_level","label":"Degree Level","values":["masters"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Virginia Polytechnic Institute and State University"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Augmented Reality","Machine Learning","Task Guidance"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International"]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://creativecommons.org/licenses/by-nc-sa/4.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:40807"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10919/119417"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Augmented reality task guidance systems provide assistance for procedural tasks, which require a sequence of physical actions, by rendering virtual guidance visuals within the real-world environment. An example of such a task would be to secure two wood parts together, which could display guidance visuals indicating the user to pick up a drill and drill each screw. Current AR task guidance systems are limited in that they require AR system experts for use, require CAD models of real-world objects, or only function for limited types of tasks or environments. We propose a general-purpose AR task guidance approach and proof-of-concept system to generate guidance for tasks defined by natural language. Our approach allows an operator to take pictures of relevant objects and write task instructions for an end user, which are used by the system to determine where to place guidance visuals. Then, an end user can receive and follow guidance even if objects change location or environment. Guidance includes reusable visuals that display generic actions, such as our system's 3D hand animations. Our approach utilizes current vision-language machine learning models for text and image semantic understanding and object localization. We built a proof-of-concept system using our approach and tested its accuracy and usability in a user study. We found that all operators were able to generate clear guidance for tasks in an office room, and end users were able to follow the guidance visuals to complete the expected action 85.7% of the time without knowledge of their tasks. Participants rated that our system was easy to use to generate guidance visuals they expected."]},{"key":"dc:description.abstractgeneral","label":"General Abstract","values":["Augmented Reality (AR) task guidance systems provide assistance for tasks by placing virtual guidance visuals on top of the real world through displays. An example of such a task would be to secure two wood parts together, which could display guidance visuals indicating the user to pick up a drill and drill each screw. Current AR task guidance systems are limited in that they require AR system experts for use, require detailed models of real-world objects, or only function for limited types of tasks or environments. We propose a new task guidance approach and built a system to generate guidance for tasks defined by written instructions. Our approach allows an operator to take pictures of relevant objects and write task instructions for an end user, which are used by the system to determine where to place digital visuals. Then, an end user can receive and follow guidance even if objects change location or environment. Guidance includes visuals that display generic actions, such as our system's 3D hand animations that mimic human hand actions. Our approach utilizes AI models for text and image understanding and object detection. We built a proof-of-concept system using our approach and tested its accuracy and usability in a user study. We found that all operators were able to generate clear guidance for tasks in an office room, and end users were able to follow the guidance visuals to complete the expected action 85.7% of the time without knowledge of the tasks. Participants rated that our system made it easy to write instructions and take pictures to create guidance visuals."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["Master of Science"]},{"key":"dc:format.medium","label":"Dc Format Medium","values":["ETD"]},{"key":"dc:title","label":"Title","values":["General-Purpose Task Guidance from Natural Language in Augmented Reality using Vision-Language Models"]}]}],"canonical_facts":{"dc:contributor.committeechair":["Abbott, Amos L.","Bowman, Douglas A."],"dc:contributor.committeemember":["Thomas, Christopher Lee","Jones, Creed Farris"],"dc:contributor.department":["Electrical and Computer Engineering"],"dc:creator":["Stover, Daniel James"],"dc:date.accessioned":["2024-06-13T08:01:42Z"],"dc:date.available":["2024-06-13T08:01:42Z"],"dc:date.issued":["2024-06-12"],"dc:description.abstract":["Augmented reality task guidance systems provide assistance for procedural tasks, which require a sequence of physical actions, by rendering virtual guidance visuals within the real-world environment. An example of such a task would be to secure two wood parts together, which could display guidance visuals indicating the user to pick up a drill and drill each screw. Current AR task guidance systems are limited in that they require AR system experts for use, require CAD models of real-world objects, or only function for limited types of tasks or environments. We propose a general-purpose AR task guidance approach and proof-of-concept system to generate guidance for tasks defined by natural language. Our approach allows an operator to take pictures of relevant objects and write task instructions for an end user, which are used by the system to determine where to place guidance visuals. Then, an end user can receive and follow guidance even if objects change location or environment. Guidance includes reusable visuals that display generic actions, such as our system's 3D hand animations. Our approach utilizes current vision-language machine learning models for text and image semantic understanding and object localization. We built a proof-of-concept system using our approach and tested its accuracy and usability in a user study. We found that all operators were able to generate clear guidance for tasks in an office room, and end users were able to follow the guidance visuals to complete the expected action 85.7% of the time without knowledge of their tasks. Participants rated that our system was easy to use to generate guidance visuals they expected."],"dc:description.abstractgeneral":["Augmented Reality (AR) task guidance systems provide assistance for tasks by placing virtual guidance visuals on top of the real world through displays. An example of such a task would be to secure two wood parts together, which could display guidance visuals indicating the user to pick up a drill and drill each screw. Current AR task guidance systems are limited in that they require AR system experts for use, require detailed models of real-world objects, or only function for limited types of tasks or environments. We propose a new task guidance approach and built a system to generate guidance for tasks defined by written instructions. Our approach allows an operator to take pictures of relevant objects and write task instructions for an end user, which are used by the system to determine where to place digital visuals. Then, an end user can receive and follow guidance even if objects change location or environment. Guidance includes visuals that display generic actions, such as our system's 3D hand animations that mimic human hand actions. Our approach utilizes AI models for text and image understanding and object detection. We built a proof-of-concept system using our approach and tested its accuracy and usability in a user study. We found that all operators were able to generate clear guidance for tasks in an office room, and end users were able to follow the guidance visuals to complete the expected action 85.7% of the time without knowledge of the tasks. Participants rated that our system made it easy to write instructions and take pictures to create guidance visuals."],"dc:description.degree":["Master of Science"],"dc:format.medium":["ETD"],"dc:identifier.other":["vt_gsexam:40807"],"dc:identifier.uri":["https://hdl.handle.net/10919/119417"],"dc:language.iso":["en"],"dc:publisher":["Virginia Tech"],"dc:rights":["Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International"],"dc:rights.uri":["http://creativecommons.org/licenses/by-nc-sa/4.0/"],"dc:subject":["Augmented Reality","Machine Learning","Task Guidance"],"dc:title":["General-Purpose Task Guidance from Natural Language in Augmented Reality using Vision-Language Models"],"dc:type":["Thesis"],"thesis:degree_discipline":["Computer Engineering"],"thesis:degree_level":["masters"],"thesis:degree_name":["Master of Science"],"thesis:institution_name":["Virginia Polytechnic Institute and State University"]},"updated_at":"2026-07-22T22:20:27Z"}