University of Denver
A Study on Multimodal AI for Mild Cognitive Impairment Detection
Abstract
dc:description.abstract<p>Mild Cognitive Impairment (MCI) is an early stage of memory loss or other cognitive ability loss in individuals who maintain the ability to independently perform most activities of daily living. It is considered a transitional stage between normal cognitive stage and more severe cognitive declines like dementia or Alzheimer’s. Based on the reports from the National Institute of Aging (NIA), people with MCI are at a greater risk of developing dementia, thus it is of great importance to detect MCI at the earliest possible to mitigate the transformation of MCI to Alzheimer’s and dementia. Recent studies have harnessed Artificial Intelligence (AI) to develop automated methods to predict and detect MCI. The majority of the existing research is based on unimodal data (e.g., only speech or prosody), but recent studies have shown that multimodality leads to a more accurate prediction of MCI. However, effectively exploiting different modalities is still a big challenge due to the lack of efficient fusion methods. This thesis proposes a mid-level fusion architecture to make use of multimodal data for MCI prediction. We introduce a multimodal speech-language-vision Deep Learning-based method to differentiate MCI from Normal Cognition (NC). Our proposed architecture includes co-attention blocks to fuse three different modalities at the embedding level to find the potential interactions between speech (audio), language (transcribed speech), and vision (facial videos) within the cross-Transformer layer. To study and evaluated the proposed mid-level fusion model, the I-CONECT dataset was used. It contains a large number of semi-structured conversations via the internet/webcam between participants aged 75+ years old and interviewers. Our experimental results show that the proposed fusion method can detect MCI from NC with an average AUC of (85.3%) which outperforms the unimodal and bimodal baseline models.</p> <p>This thesis demonstrates that multimodal deep learning models outperform unimodal models in detecting MCI in older adults. To generalize the applicability of these findings, further research employing larger datasets should be conducted.</p>
Degree
thesis:*- Name thesis:degree_name
- M.S. in Computer Engineering
- Level thesis:degree_level
- Masters Thesis
- Year
- 2024
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Far Poor, Farida
- Contributors dc:contributor
-
- Mohammad H. Mahoor
- Yun-Bo Yi
- Kerstin Sophie Haring
- Haluk Ogmen
Subjects
dc:subject × 6Rights
dc:rights- Statement dc:rights
-
- <p>Copyright is held by the author. Permanently suppressed.</p>
- Language dc:language
- English (eng)
Identifiers
dc:identifier.*- Repository record dc:identifier
- https://digitalcommons.du.edu/etd/2449
- OAI identifier oai:identifier
- oai:digitalcommons.du.edu:etd-3438