{"id":{"repo_id":"vt","oai_identifier":"oai:vtechworks.lib.vt.edu:10919/138118"},"canonical_url":"https://search.dev.ndltd.org/etd/vt/oai:vtechworks.lib.vt.edu:10919/138118","repository":{"repo_id":"vt","name":"Virginia Tech","base_url":"https://vtechworks.lib.vt.edu/oai/request"},"display":{"title":"Toward AI-Mediated Immersive Sensemaking with Gaze-Aware Semantic Interaction","abstract":"Motivation. Analysts who work with large text corpora must forage for evidence, con- nect disparate facts, and synthesize explanations, which imposes a heavy cognitive load. Immersive Analytics offers improving the experience with spatial memory and embodied interaction that can reduce this burden, but does not save the analyst from exhaustively browsing the corpus to find what matters. However, modern head-worn displays include eye tracking, creating an opportunity to infer an analyst's perceived interest implicitly and to provide timely, intelligible, attention-aware assistance that can essentially help in offloading some of the cognitive work. Problem. How can we model an analyst's interest from their gaze so that an AI assistant guides foraging and supports synthesis while preserving analysts' agency over the layout? Specifically, we need methods that (a) predict perceived relevance at document and term levels during multi-document investigations, and (b) expose those predictions through visual cues that help users make sense of complex evidence. Approach. We introduced a gaze-derived interest model that combines fixation duration and dwell count, adjusted for high-frequency terms, to compute GazeScore for documents and words. In parallel, we studied analysts' acceptance of different automation levels to ground design principles for a gaze-aware assistant. We operationalized the model in EyeST, an immersive analytic tool that presents two levels of visual cues to externalize the analyst's interest. Global cues provide interpretable, low-overhead signals by ranking and color-encoded evidence. Local cues reveal relationships between documents to promote discovery without clutter, while being grounded in the analyst's interest. We conducted a feasibility study that compared GazeScore to the analyst's perceived relevance. In parallel, we assessed analysts' acceptance of automated systems with a clustering task offering three levels of automation. The findings from the two studies enabled us to develop gaze-aware semantic interactions for immersive sensemaking, followed by two studies: one examining its effects on foraging, while the other tested the effects of adaptive annotation during synthesis. Results. GazeScore separated relevant from irrelevant content at the word level from the outset, enabling a real-time document relevance predictor with high precision. The clustering study showed that analysts favored assistance that preserves analyst's control over the layout and provides clear rationales. Subsequent studies revealed that global cues increased the efficiency of the analyst by helping them spend more time on relevant information while avoiding noise. Local cues encouraged individual exploration and surfaced overlooked but useful documents. Both gaze-derived cues guided analysts to essential clues vital to the sensemaking task, reduced perceived physical demand, and reduced the need for explicit externalization of the analyst's interest. Implications. The findings point toward a design pattern for AI-mediated immersive sense- making: invest in bootstrapping evidence, emphasize global high-level signals early, reveal local relationships on demand, and pair every suggestion with a clear rationale to preserve trust and agency. More broadly, this work extends semantic interaction into implicit chan- nels by showing how gaze can externalize evolving interest in real time. While our focus was on predicting perceived relevance, the approach opens pathways to incorporate other implicit signals to capture a richer picture of analysts' cognitive states. Together, these insights pave the way for more adaptive, trustworthy, and human-centered gaze-aware systems that deepen human–AI collaboration in immersive analytics.","abstract_html":"Motivation. Analysts who work with large text corpora must forage for evidence, con- nect disparate facts, and synthesize explanations, which imposes a heavy cognitive load. Immersive Analytics offers improving the experience with spatial memory and embodied interaction that can reduce this burden, but does not save the analyst from exhaustively browsing the corpus to find what matters. However, modern head-worn displays include eye tracking, creating an opportunity to infer an analyst&#x27;s perceived interest implicitly and to provide timely, intelligible, attention-aware assistance that can essentially help in offloading some of the cognitive work. Problem. How can we model an analyst&#x27;s interest from their gaze so that an AI assistant guides foraging and supports synthesis while preserving analysts&#x27; agency over the layout? Specifically, we need methods that (a) predict perceived relevance at document and term levels during multi-document investigations, and (b) expose those predictions through visual cues that help users make sense of complex evidence. Approach. We introduced a gaze-derived interest model that combines fixation duration and dwell count, adjusted for high-frequency terms, to compute GazeScore for documents and words. In parallel, we studied analysts&#x27; acceptance of different automation levels to ground design principles for a gaze-aware assistant. We operationalized the model in EyeST, an immersive analytic tool that presents two levels of visual cues to externalize the analyst&#x27;s interest. Global cues provide interpretable, low-overhead signals by ranking and color-encoded evidence. Local cues reveal relationships between documents to promote discovery without clutter, while being grounded in the analyst&#x27;s interest. We conducted a feasibility study that compared GazeScore to the analyst&#x27;s perceived relevance. In parallel, we assessed analysts&#x27; acceptance of automated systems with a clustering task offering three levels of automation. The findings from the two studies enabled us to develop gaze-aware semantic interactions for immersive sensemaking, followed by two studies: one examining its effects on foraging, while the other tested the effects of adaptive annotation during synthesis. Results. GazeScore separated relevant from irrelevant content at the word level from the outset, enabling a real-time document relevance predictor with high precision. The clustering study showed that analysts favored assistance that preserves analyst&#x27;s control over the layout and provides clear rationales. Subsequent studies revealed that global cues increased the efficiency of the analyst by helping them spend more time on relevant information while avoiding noise. Local cues encouraged individual exploration and surfaced overlooked but useful documents. Both gaze-derived cues guided analysts to essential clues vital to the sensemaking task, reduced perceived physical demand, and reduced the need for explicit externalization of the analyst&#x27;s interest. Implications. The findings point toward a design pattern for AI-mediated immersive sense- making: invest in bootstrapping evidence, emphasize global high-level signals early, reveal local relationships on demand, and pair every suggestion with a clear rationale to preserve trust and agency. More broadly, this work extends semantic interaction into implicit chan- nels by showing how gaze can externalize evolving interest in real time. While our focus was on predicting perceived relevance, the approach opens pathways to incorporate other implicit signals to capture a richer picture of analysts&#x27; cognitive states. Together, these insights pave the way for more adaptive, trustworthy, and human-centered gaze-aware systems that deepen human–AI collaboration in immersive analytics.","abstract_has_math":false,"creators":["Tahmid, Ibrahim Asadullah"],"institution":"Virginia Tech","degree_name":"Doctor of Philosophy","degree_level":"doctoral","degree_discipline":"Computer Science & Applications","degree_department":"Computer Science and Applications","school":null,"contributors":[],"advisors":[],"committee_chairs":["Bowman, Douglas Andrew","North, Christopher L."],"committee_members":["Wenskovitch, John Edward","Whitley, Kirsten","David-John, Brendan Matthew"],"year":2025,"date_issued":"2025-10-09","date_published":"2025-10-09","updated_at":"2026-07-22T22:19:29Z","subjects":["Mixed Reality","Sensemaking","Eye-Tracking","Human-Centered AI","Semantic Interaction","Recommendation Model"],"languages":["en"],"rights":["Creative Commons Attribution 4.0 International"],"rights_urls":["http://creativecommons.org/licenses/by/4.0/"],"identifier_entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:44732"],"render_values":[{"text":"vt_gsexam:44732","href":null,"code":true}]}]},"links":{"outbound_url":"https://hdl.handle.net/10919/138118","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.committeechair","label":"Committee Chair","values":["Bowman, Douglas Andrew","North, Christopher L."]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Wenskovitch, John Edward","Whitley, Kirsten","David-John, Brendan Matthew"]},{"key":"dc:contributor.department","label":"Department","values":["Computer Science and Applications"]},{"key":"dc:creator","label":"Author","values":["Tahmid, Ibrahim Asadullah"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2025-10-10T08:00:23Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2025-10-10T08:00:23Z"]},{"key":"dc:date.issued","label":"Date","values":["2025-10-09"]},{"key":"dc:publisher","label":"Institution","values":["Virginia Tech"]},{"key":"dc:type","label":"Dc Type","values":["Dissertation"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science & Applications"]},{"key":"thesis:degree_level","label":"Degree Level","values":["doctoral"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctor of Philosophy"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Virginia Polytechnic Institute and State University"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Mixed Reality","Sensemaking","Eye-Tracking","Human-Centered AI","Semantic Interaction","Recommendation Model"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Creative Commons Attribution 4.0 International"]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://creativecommons.org/licenses/by/4.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:44732"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10919/138118"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Motivation. Analysts who work with large text corpora must forage for evidence, con- nect disparate facts, and synthesize explanations, which imposes a heavy cognitive load. Immersive Analytics offers improving the experience with spatial memory and embodied interaction that can reduce this burden, but does not save the analyst from exhaustively browsing the corpus to find what matters. However, modern head-worn displays include eye tracking, creating an opportunity to infer an analyst's perceived interest implicitly and to provide timely, intelligible, attention-aware assistance that can essentially help in offloading some of the cognitive work. Problem. How can we model an analyst's interest from their gaze so that an AI assistant guides foraging and supports synthesis while preserving analysts' agency over the layout? Specifically, we need methods that (a) predict perceived relevance at document and term levels during multi-document investigations, and (b) expose those predictions through visual cues that help users make sense of complex evidence. Approach. We introduced a gaze-derived interest model that combines fixation duration and dwell count, adjusted for high-frequency terms, to compute GazeScore for documents and words. In parallel, we studied analysts' acceptance of different automation levels to ground design principles for a gaze-aware assistant. We operationalized the model in EyeST, an immersive analytic tool that presents two levels of visual cues to externalize the analyst's interest. Global cues provide interpretable, low-overhead signals by ranking and color-encoded evidence. Local cues reveal relationships between documents to promote discovery without clutter, while being grounded in the analyst's interest. We conducted a feasibility study that compared GazeScore to the analyst's perceived relevance. In parallel, we assessed analysts' acceptance of automated systems with a clustering task offering three levels of automation. The findings from the two studies enabled us to develop gaze-aware semantic interactions for immersive sensemaking, followed by two studies: one examining its effects on foraging, while the other tested the effects of adaptive annotation during synthesis. Results. GazeScore separated relevant from irrelevant content at the word level from the outset, enabling a real-time document relevance predictor with high precision. The clustering study showed that analysts favored assistance that preserves analyst's control over the layout and provides clear rationales. Subsequent studies revealed that global cues increased the efficiency of the analyst by helping them spend more time on relevant information while avoiding noise. Local cues encouraged individual exploration and surfaced overlooked but useful documents. Both gaze-derived cues guided analysts to essential clues vital to the sensemaking task, reduced perceived physical demand, and reduced the need for explicit externalization of the analyst's interest. Implications. The findings point toward a design pattern for AI-mediated immersive sense- making: invest in bootstrapping evidence, emphasize global high-level signals early, reveal local relationships on demand, and pair every suggestion with a clear rationale to preserve trust and agency. More broadly, this work extends semantic interaction into implicit chan- nels by showing how gaze can externalize evolving interest in real time. While our focus was on predicting perceived relevance, the approach opens pathways to incorporate other implicit signals to capture a richer picture of analysts' cognitive states. Together, these insights pave the way for more adaptive, trustworthy, and human-centered gaze-aware systems that deepen human–AI collaboration in immersive analytics."]},{"key":"dc:description.abstractgeneral","label":"General Abstract","values":["We live in the peak of the information age, surrounded by more data than current tools can easily handle. To manage this complexity, two disruptive technologies, extended reality (XR) and artificial intelligence (AI), offer complementary strengths. XR provides an expansive workspace where digital information can be spread out, navigated, and manipulated as naturally as physical documents, freeing analysts from the constraints of two-dimensional screens. AI, on the other hand, excels at processing vast information spaces quickly, surfacing insights that would take humans much longer to uncover. Yet, true value emerges only when AI adapts to human intent, aligning its exploration with the human's evolving needs. This dissertation argues that the synergy of XR and AI can fundamentally improve work- flows in domains where professionals must sift through large, interconnected datasets, such as intelligence analysis, investigative journalism, or academic research. The key lies in leveraging the built-in eye-tracking sensors of XR headsets. By observing where users look, what they read, and how long they dwell on different topics, we can construct an interest model that maps their perceived interest in different topics. With empirical evidence, we showed that this model can, indeed, predict evolving user interests with high precision. In parallel, we investigated how analysts perceive automation in immersive environments. Our studies revealed that users welcome intelligent assistance as long as they retain control, the automation is transparent, and they can accept, dismiss, or even undo system actions. Guided by these insights, we developed a gaze-aware immersive analytic tool that uses the interest model to generate interpretable visual cues. Evaluations with novice and professional intelligence analysts showed that these gaze-derived cues help users navigate complex datasets more efficiently, identify relevant information while avoiding distractions, and synthesize their findings with reduced need for manually typing out their intents. Together, the cues enabled a form of semantic interaction that lowered physical workload and enriched human–AI collaboration. While this work centers on predicting user perception from gaze, our approach paves the way for a broader future. Other implicit signals, such as indicators of stress, could extend the model to capture not just interest but also cognitive and emotional states. This opens the door to richer semantic interactions and more adaptive human–AI partnerships in immersive analytics. The findings presented here lay the foundation for such systems, demonstrating how gaze-aware interest models can leverage the strengths of XR and AI together to meaningfully support AI-mediated immersive sensemaking."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["Doctor of Philosophy"]},{"key":"dc:format.medium","label":"Dc Format Medium","values":["ETD"]},{"key":"dc:title","label":"Title","values":["Toward AI-Mediated Immersive Sensemaking with Gaze-Aware Semantic Interaction"]}]}],"canonical_facts":{"dc:contributor.committeechair":["Bowman, Douglas Andrew","North, Christopher L."],"dc:contributor.committeemember":["Wenskovitch, John Edward","Whitley, Kirsten","David-John, Brendan Matthew"],"dc:contributor.department":["Computer Science and Applications"],"dc:creator":["Tahmid, Ibrahim Asadullah"],"dc:date.accessioned":["2025-10-10T08:00:23Z"],"dc:date.available":["2025-10-10T08:00:23Z"],"dc:date.issued":["2025-10-09"],"dc:description.abstract":["Motivation. Analysts who work with large text corpora must forage for evidence, con- nect disparate facts, and synthesize explanations, which imposes a heavy cognitive load. Immersive Analytics offers improving the experience with spatial memory and embodied interaction that can reduce this burden, but does not save the analyst from exhaustively browsing the corpus to find what matters. However, modern head-worn displays include eye tracking, creating an opportunity to infer an analyst's perceived interest implicitly and to provide timely, intelligible, attention-aware assistance that can essentially help in offloading some of the cognitive work. Problem. How can we model an analyst's interest from their gaze so that an AI assistant guides foraging and supports synthesis while preserving analysts' agency over the layout? Specifically, we need methods that (a) predict perceived relevance at document and term levels during multi-document investigations, and (b) expose those predictions through visual cues that help users make sense of complex evidence. Approach. We introduced a gaze-derived interest model that combines fixation duration and dwell count, adjusted for high-frequency terms, to compute GazeScore for documents and words. In parallel, we studied analysts' acceptance of different automation levels to ground design principles for a gaze-aware assistant. We operationalized the model in EyeST, an immersive analytic tool that presents two levels of visual cues to externalize the analyst's interest. Global cues provide interpretable, low-overhead signals by ranking and color-encoded evidence. Local cues reveal relationships between documents to promote discovery without clutter, while being grounded in the analyst's interest. We conducted a feasibility study that compared GazeScore to the analyst's perceived relevance. In parallel, we assessed analysts' acceptance of automated systems with a clustering task offering three levels of automation. The findings from the two studies enabled us to develop gaze-aware semantic interactions for immersive sensemaking, followed by two studies: one examining its effects on foraging, while the other tested the effects of adaptive annotation during synthesis. Results. GazeScore separated relevant from irrelevant content at the word level from the outset, enabling a real-time document relevance predictor with high precision. The clustering study showed that analysts favored assistance that preserves analyst's control over the layout and provides clear rationales. Subsequent studies revealed that global cues increased the efficiency of the analyst by helping them spend more time on relevant information while avoiding noise. Local cues encouraged individual exploration and surfaced overlooked but useful documents. Both gaze-derived cues guided analysts to essential clues vital to the sensemaking task, reduced perceived physical demand, and reduced the need for explicit externalization of the analyst's interest. Implications. The findings point toward a design pattern for AI-mediated immersive sense- making: invest in bootstrapping evidence, emphasize global high-level signals early, reveal local relationships on demand, and pair every suggestion with a clear rationale to preserve trust and agency. More broadly, this work extends semantic interaction into implicit chan- nels by showing how gaze can externalize evolving interest in real time. While our focus was on predicting perceived relevance, the approach opens pathways to incorporate other implicit signals to capture a richer picture of analysts' cognitive states. Together, these insights pave the way for more adaptive, trustworthy, and human-centered gaze-aware systems that deepen human–AI collaboration in immersive analytics."],"dc:description.abstractgeneral":["We live in the peak of the information age, surrounded by more data than current tools can easily handle. To manage this complexity, two disruptive technologies, extended reality (XR) and artificial intelligence (AI), offer complementary strengths. XR provides an expansive workspace where digital information can be spread out, navigated, and manipulated as naturally as physical documents, freeing analysts from the constraints of two-dimensional screens. AI, on the other hand, excels at processing vast information spaces quickly, surfacing insights that would take humans much longer to uncover. Yet, true value emerges only when AI adapts to human intent, aligning its exploration with the human's evolving needs. This dissertation argues that the synergy of XR and AI can fundamentally improve work- flows in domains where professionals must sift through large, interconnected datasets, such as intelligence analysis, investigative journalism, or academic research. The key lies in leveraging the built-in eye-tracking sensors of XR headsets. By observing where users look, what they read, and how long they dwell on different topics, we can construct an interest model that maps their perceived interest in different topics. With empirical evidence, we showed that this model can, indeed, predict evolving user interests with high precision. In parallel, we investigated how analysts perceive automation in immersive environments. Our studies revealed that users welcome intelligent assistance as long as they retain control, the automation is transparent, and they can accept, dismiss, or even undo system actions. Guided by these insights, we developed a gaze-aware immersive analytic tool that uses the interest model to generate interpretable visual cues. Evaluations with novice and professional intelligence analysts showed that these gaze-derived cues help users navigate complex datasets more efficiently, identify relevant information while avoiding distractions, and synthesize their findings with reduced need for manually typing out their intents. Together, the cues enabled a form of semantic interaction that lowered physical workload and enriched human–AI collaboration. While this work centers on predicting user perception from gaze, our approach paves the way for a broader future. Other implicit signals, such as indicators of stress, could extend the model to capture not just interest but also cognitive and emotional states. This opens the door to richer semantic interactions and more adaptive human–AI partnerships in immersive analytics. The findings presented here lay the foundation for such systems, demonstrating how gaze-aware interest models can leverage the strengths of XR and AI together to meaningfully support AI-mediated immersive sensemaking."],"dc:description.degree":["Doctor of Philosophy"],"dc:format.medium":["ETD"],"dc:identifier.other":["vt_gsexam:44732"],"dc:identifier.uri":["https://hdl.handle.net/10919/138118"],"dc:language.iso":["en"],"dc:publisher":["Virginia Tech"],"dc:rights":["Creative Commons Attribution 4.0 International"],"dc:rights.uri":["http://creativecommons.org/licenses/by/4.0/"],"dc:subject":["Mixed Reality","Sensemaking","Eye-Tracking","Human-Centered AI","Semantic Interaction","Recommendation Model"],"dc:title":["Toward AI-Mediated Immersive Sensemaking with Gaze-Aware Semantic Interaction"],"dc:type":["Dissertation"],"thesis:degree_discipline":["Computer Science & Applications"],"thesis:degree_level":["doctoral"],"thesis:degree_name":["Doctor of Philosophy"],"thesis:institution_name":["Virginia Polytechnic Institute and State University"]},"updated_at":"2026-07-22T22:19:29Z"}