University of Northampton
Multimodal Eye-Tracking and Predictive Gaze Modelling for Assistive Human–Computer Interaction in Augmentative and Alternative Communication Systems
Abstract
dc:description.abstractAugmentative and Alternative Communication (AAC) systems are critical for individuals with neurological disorders such as Amyotrophic Lateral Sclerosis (ALS) and other motor impairments. These systems provide a means to express intent when conventional communication methods such as natural speech or manual typing are unavailable. Current eye-tracking solutions frequently fail due to three critical issues: low input speed, the Midas Touch problem of unintended selections during navigation, and prohibitive hardware costs. This thesis addresses these challenges by developing a multimodal and software-driven AAC framework that integrates eye-tracking, automatic speech recognition (ASR), speech synthesis, and gaze prediction to improve communication efficiency and user autonomy. To resolve the accuracy and latency trade-off, the framework implements a Bi-directional Long Short-Term Memory (LSTM) gaze prediction model. By training on the eye-tracking sequences of the EEGET-ALS dataset, which contains about 6million samples of eye-tracking data, the model achieves a mean prediction error of 0.00948 pixels. This high-precision prediction stabilizes the cursor and reduces the cognitive load of dwell-time selection, which directly addresses the bottleneck of low input speed. To solve the Midas Touch problem, the system departs from traditional gaze-only interfaces by employing a multimodal confirmation logic. By decoupling navigation (gaze) from selection (blink detection), the system allows users to scan the interface freely without triggering unintended actions. This architecture is powered by a real-time pipeline using MediaPipe for gaze estimation, Speech-to-Text for ASR and Text-to-Speech. This ensures a reduced cost by sidelining the reliance on high-end eye-tracking devices and hardware-agnostic solution that operates on standard consumer devices. Results indicate that the integrated ASR system achieves a 2% word error rate while the LSTM model yields a high correlation between predicted and actual gaze coordinates with an R2 of 0.99915. Usability testing across diverse experimental scenarios yielded an overall system rating of 9.08/10. These findings suggest that a multimodal and software-driven approach can democratize access to high-performance AAC while maintaining high user satisfaction.<br/>
Degree
thesis:*- Name dc:type.qualificationname
- Doctoral Thesis
- Level dc:type.qualificationlevel
- Student thesis
- Grantor dc:publisher.institution
- University of Northampton
- Year dc:date.issued
- 2026
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Edughele, Hilary
- Advisors dc:contributor.advisor
-
- Opoku Agyeman, Michael
- Zhang, Yinghui
- Khatri, Roshni Rama
Subjects
dc:subject × 6Rights
- Language dc:language
- eng
Identifiers
dc:identifier.*- Identifier
- oai:pure.atira.dk:studenttheses/8a1c95cf-5a5c-4fc4-8016-ac54f14b7b2c
- OAI identifier oai:identifier
- oai:pure.atira.dk:studenttheses/8a1c95cf-5a5c-4fc4-8016-ac54f14b7b2c