{"id":{"repo_id":"northampton","oai_identifier":"oai:pure.atira.dk:studenttheses/8a1c95cf-5a5c-4fc4-8016-ac54f14b7b2c"},"canonical_url":"https://search.dev.ndltd.org/etd/northampton/oai:pure.atira.dk:studenttheses/8a1c95cf-5a5c-4fc4-8016-ac54f14b7b2c","repository":{"repo_id":"northampton","name":"University of Northampton","base_url":"https://pure.northampton.ac.uk/ws/oai"},"display":{"title":"Multimodal Eye-Tracking and Predictive Gaze Modelling for Assistive Human–Computer Interaction in Augmentative and Alternative Communication Systems","abstract":"Augmentative and Alternative Communication (AAC) systems are critical for individuals with neurological disorders such as Amyotrophic Lateral Sclerosis (ALS) and other motor impairments. These systems provide a means to express intent when conventional communication methods such as natural speech or manual typing are unavailable. Current eye-tracking solutions frequently fail due to three critical issues: low input speed, the Midas Touch problem of unintended selections during navigation, and prohibitive hardware costs. This thesis addresses these challenges by developing a multimodal and software-driven AAC framework that integrates eye-tracking, automatic speech recognition (ASR), speech synthesis, and gaze prediction to improve communication efficiency and user autonomy. To resolve the accuracy and latency trade-off, the framework implements a Bi-directional Long Short-Term Memory (LSTM) gaze prediction model. By training on the eye-tracking sequences of the EEGET-ALS dataset, which contains about 6million samples of eye-tracking data, the model achieves a mean prediction error of 0.00948 pixels. This high-precision prediction stabilizes the cursor and reduces the cognitive load of dwell-time selection, which directly addresses the bottleneck of low input speed. To solve the Midas Touch problem, the system departs from traditional gaze-only interfaces by employing a multimodal confirmation logic. By decoupling navigation (gaze) from selection (blink detection), the system allows users to scan the interface freely without triggering unintended actions. This architecture is powered by a real-time pipeline using MediaPipe for gaze estimation, Speech-to-Text for ASR and Text-to-Speech. This ensures a reduced cost by sidelining the reliance on high-end eye-tracking devices and hardware-agnostic solution that operates on standard consumer devices. Results indicate that the integrated ASR system achieves a 2% word error rate while the LSTM model yields a high correlation between predicted and actual gaze coordinates with an R2 of 0.99915. Usability testing across diverse experimental scenarios yielded an overall system rating of 9.08/10. These findings suggest that a multimodal and software-driven approach can democratize access to high-performance AAC while maintaining high user satisfaction.<br/>","abstract_html":"Augmentative and Alternative Communication (AAC) systems are critical for individuals with neurological disorders such as Amyotrophic Lateral Sclerosis (ALS) and other motor impairments. These systems provide a means to express intent when conventional communication methods such as natural speech or manual typing are unavailable. Current eye-tracking solutions frequently fail due to three critical issues: low input speed, the Midas Touch problem of unintended selections during navigation, and prohibitive hardware costs. This thesis addresses these challenges by developing a multimodal and software-driven AAC framework that integrates eye-tracking, automatic speech recognition (ASR), speech synthesis, and gaze prediction to improve communication efficiency and user autonomy. To resolve the accuracy and latency trade-off, the framework implements a Bi-directional Long Short-Term Memory (LSTM) gaze prediction model. By training on the eye-tracking sequences of the EEGET-ALS dataset, which contains about 6million samples of eye-tracking data, the model achieves a mean prediction error of 0.00948 pixels. This high-precision prediction stabilizes the cursor and reduces the cognitive load of dwell-time selection, which directly addresses the bottleneck of low input speed. To solve the Midas Touch problem, the system departs from traditional gaze-only interfaces by employing a multimodal confirmation logic. By decoupling navigation (gaze) from selection (blink detection), the system allows users to scan the interface freely without triggering unintended actions. This architecture is powered by a real-time pipeline using MediaPipe for gaze estimation, Speech-to-Text for ASR and Text-to-Speech. This ensures a reduced cost by sidelining the reliance on high-end eye-tracking devices and hardware-agnostic solution that operates on standard consumer devices. Results indicate that the integrated ASR system achieves a 2% word error rate while the LSTM model yields a high correlation between predicted and actual gaze coordinates with an R2 of 0.99915. Usability testing across diverse experimental scenarios yielded an overall system rating of 9.08/10. These findings suggest that a multimodal and software-driven approach can democratize access to high-performance AAC while maintaining high user satisfaction.&lt;br/&gt;","abstract_has_math":false,"creators":["Edughele, Hilary"],"institution":"University of Northampton","degree_name":"Doctoral Thesis","degree_level":"Student thesis","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Opoku Agyeman, Michael","Zhang, Yinghui","Khatri, Roshni Rama"],"committee_chairs":[],"committee_members":[],"year":2026,"date_issued":"2026-3-30","date_published":"2026-3-30","updated_at":"2026-07-24T03:26:59Z","subjects":["Augmentative and Alternative Communication Systems","Eye-Tracking","Gaze prediction","Gamification","Human–Computer Interaction","LSTM"],"languages":["eng"],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["oai:pure.atira.dk:studenttheses/8a1c95cf-5a5c-4fc4-8016-ac54f14b7b2c"],"render_values":[{"text":"oai:pure.atira.dk:studenttheses/8a1c95cf-5a5c-4fc4-8016-ac54f14b7b2c","href":null,"code":true}]}]},"links":{"outbound_url":"https://pure.northampton.ac.uk/en/studentTheses/8a1c95cf-5a5c-4fc4-8016-ac54f14b7b2c","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Opoku Agyeman, Michael","Zhang, Yinghui","Khatri, Roshni Rama"]},{"key":"dc:creator","label":"Author","values":["Edughele, Hilary"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2026-3-30"]},{"key":"dc:date.issued","label":"Date","values":["2026-3-30"]},{"key":"dc:publisher.department","label":"Dc Publisher Department","values":["Graduate School","Centre for Advanced and Smart Technologies","Faculty of Education, Arts, Science & Technology","Technology"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["University of Northampton"]},{"key":"dc:relation.isreferencedby","label":"Dc Relation Isreferencedby","values":["https://pure.northampton.ac.uk/en/studentTheses/8a1c95cf-5a5c-4fc4-8016-ac54f14b7b2c"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["Student thesis"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["Doctoral Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Augmentative and Alternative Communication Systems","Eye-Tracking","Gaze prediction","Gamification","Human–Computer Interaction","LSTM"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["oai:pure.atira.dk:studenttheses/8a1c95cf-5a5c-4fc4-8016-ac54f14b7b2c","https://pure.northampton.ac.uk/en/studentTheses/8a1c95cf-5a5c-4fc4-8016-ac54f14b7b2c"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://pure.northampton.ac.uk/files/92861615/Edughele_Hilary_2026_Multimodal_Eye_Tracking_and_Predictive_Gaze_Modelling_for_Assistive_Human_Computer_Interaction_in_Augmentative_and_Alternative_Communication_Systems.pdf"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Augmentative and Alternative Communication (AAC) systems are critical for individuals with neurological disorders such as Amyotrophic Lateral Sclerosis (ALS) and other motor impairments. These systems provide a means to express intent when conventional communication methods such as natural speech or manual typing are unavailable. Current eye-tracking solutions frequently fail due to three critical issues: low input speed, the Midas Touch problem of unintended selections during navigation, and prohibitive hardware costs. This thesis addresses these challenges by developing a multimodal and software-driven AAC framework that integrates eye-tracking, automatic speech recognition (ASR), speech synthesis, and gaze prediction to improve communication efficiency and user autonomy. To resolve the accuracy and latency trade-off, the framework implements a Bi-directional Long Short-Term Memory (LSTM) gaze prediction model. By training on the eye-tracking sequences of the EEGET-ALS dataset, which contains about 6million samples of eye-tracking data, the model achieves a mean prediction error of 0.00948 pixels. This high-precision prediction stabilizes the cursor and reduces the cognitive load of dwell-time selection, which directly addresses the bottleneck of low input speed. To solve the Midas Touch problem, the system departs from traditional gaze-only interfaces by employing a multimodal confirmation logic. By decoupling navigation (gaze) from selection (blink detection), the system allows users to scan the interface freely without triggering unintended actions. This architecture is powered by a real-time pipeline using MediaPipe for gaze estimation, Speech-to-Text for ASR and Text-to-Speech. This ensures a reduced cost by sidelining the reliance on high-end eye-tracking devices and hardware-agnostic solution that operates on standard consumer devices. Results indicate that the integrated ASR system achieves a 2% word error rate while the LSTM model yields a high correlation between predicted and actual gaze coordinates with an R2 of 0.99915. Usability testing across diverse experimental scenarios yielded an overall system rating of 9.08/10. These findings suggest that a multimodal and software-driven approach can democratize access to high-performance AAC while maintaining high user satisfaction.<br/>"]},{"key":"dc:title","label":"Title","values":["Multimodal Eye-Tracking and Predictive Gaze Modelling for Assistive Human–Computer Interaction in Augmentative and Alternative Communication Systems"]}]}],"canonical_facts":{"dc:contributor.advisor":["Opoku Agyeman, Michael","Zhang, Yinghui","Khatri, Roshni Rama"],"dc:creator":["Edughele, Hilary"],"dc:date":["2026-3-30"],"dc:date.issued":["2026-3-30"],"dc:description.abstract":["Augmentative and Alternative Communication (AAC) systems are critical for individuals with neurological disorders such as Amyotrophic Lateral Sclerosis (ALS) and other motor impairments. These systems provide a means to express intent when conventional communication methods such as natural speech or manual typing are unavailable. Current eye-tracking solutions frequently fail due to three critical issues: low input speed, the Midas Touch problem of unintended selections during navigation, and prohibitive hardware costs. This thesis addresses these challenges by developing a multimodal and software-driven AAC framework that integrates eye-tracking, automatic speech recognition (ASR), speech synthesis, and gaze prediction to improve communication efficiency and user autonomy. To resolve the accuracy and latency trade-off, the framework implements a Bi-directional Long Short-Term Memory (LSTM) gaze prediction model. By training on the eye-tracking sequences of the EEGET-ALS dataset, which contains about 6million samples of eye-tracking data, the model achieves a mean prediction error of 0.00948 pixels. This high-precision prediction stabilizes the cursor and reduces the cognitive load of dwell-time selection, which directly addresses the bottleneck of low input speed. To solve the Midas Touch problem, the system departs from traditional gaze-only interfaces by employing a multimodal confirmation logic. By decoupling navigation (gaze) from selection (blink detection), the system allows users to scan the interface freely without triggering unintended actions. This architecture is powered by a real-time pipeline using MediaPipe for gaze estimation, Speech-to-Text for ASR and Text-to-Speech. This ensures a reduced cost by sidelining the reliance on high-end eye-tracking devices and hardware-agnostic solution that operates on standard consumer devices. Results indicate that the integrated ASR system achieves a 2% word error rate while the LSTM model yields a high correlation between predicted and actual gaze coordinates with an R2 of 0.99915. Usability testing across diverse experimental scenarios yielded an overall system rating of 9.08/10. These findings suggest that a multimodal and software-driven approach can democratize access to high-performance AAC while maintaining high user satisfaction.<br/>"],"dc:identifier":["oai:pure.atira.dk:studenttheses/8a1c95cf-5a5c-4fc4-8016-ac54f14b7b2c","https://pure.northampton.ac.uk/en/studentTheses/8a1c95cf-5a5c-4fc4-8016-ac54f14b7b2c"],"dc:identifier.uri":["https://pure.northampton.ac.uk/files/92861615/Edughele_Hilary_2026_Multimodal_Eye_Tracking_and_Predictive_Gaze_Modelling_for_Assistive_Human_Computer_Interaction_in_Augmentative_and_Alternative_Communication_Systems.pdf"],"dc:language":["eng"],"dc:publisher.department":["Graduate School","Centre for Advanced and Smart Technologies","Faculty of Education, Arts, Science & Technology","Technology"],"dc:publisher.institution":["University of Northampton"],"dc:relation.isreferencedby":["https://pure.northampton.ac.uk/en/studentTheses/8a1c95cf-5a5c-4fc4-8016-ac54f14b7b2c"],"dc:subject":["Augmentative and Alternative Communication Systems","Eye-Tracking","Gaze prediction","Gamification","Human–Computer Interaction","LSTM"],"dc:title":["Multimodal Eye-Tracking and Predictive Gaze Modelling for Assistive Human–Computer Interaction in Augmentative and Alternative Communication Systems"],"dc:type":["Thesis"],"dc:type.qualificationlevel":["Student thesis"],"dc:type.qualificationname":["Doctoral Thesis"]},"updated_at":"2026-07-24T03:26:59Z"}