{"id":{"repo_id":"denver","oai_identifier":"oai:digitalcommons.du.edu:etd-3438"},"canonical_url":"https://search.dev.ndltd.org/etd/denver/oai:digitalcommons.du.edu:etd-3438","repository":{"repo_id":"denver","name":"University of Denver","base_url":"https://digitalcommons.du.edu/do/oai/"},"display":{"title":"A Study on Multimodal AI for Mild Cognitive Impairment Detection","abstract":"<p>Mild Cognitive Impairment (MCI) is an early stage of memory loss or other cognitive ability loss in individuals who maintain the ability to independently perform most activities of daily living. It is considered a transitional stage between normal cognitive stage and more severe cognitive declines like dementia or Alzheimer’s. Based on the reports from the National Institute of Aging (NIA), people with MCI are at a greater risk of developing dementia, thus it is of great importance to detect MCI at the earliest possible to mitigate the transformation of MCI to Alzheimer’s and dementia. Recent studies have harnessed Artificial Intelligence (AI) to develop automated methods to predict and detect MCI. The majority of the existing research is based on unimodal data (e.g., only speech or prosody), but recent studies have shown that multimodality leads to a more accurate prediction of MCI. However, effectively exploiting different modalities is still a big challenge due to the lack of efficient fusion methods. This thesis proposes a mid-level fusion architecture to make use of multimodal data for MCI prediction. We introduce a multimodal speech-language-vision Deep Learning-based method to differentiate MCI from Normal Cognition (NC). Our proposed architecture includes co-attention blocks to fuse three different modalities at the embedding level to find the potential interactions between speech (audio), language (transcribed speech), and vision (facial videos) within the cross-Transformer layer. To study and evaluated the proposed mid-level fusion model, the I-CONECT dataset was used. It contains a large number of semi-structured conversations via the internet/webcam between participants aged 75+ years old and interviewers. Our experimental results show that the proposed fusion method can detect MCI from NC with an average AUC of (85.3%) which outperforms the unimodal and bimodal baseline models.</p> <p>This thesis demonstrates that multimodal deep learning models outperform unimodal models in detecting MCI in older adults. To generalize the applicability of these findings, further research employing larger datasets should be conducted.</p>","abstract_html":"&lt;p&gt;Mild Cognitive Impairment (MCI) is an early stage of memory loss or other cognitive ability loss in individuals who maintain the ability to independently perform most activities of daily living. It is considered a transitional stage between normal cognitive stage and more severe cognitive declines like dementia or Alzheimer’s. Based on the reports from the National Institute of Aging (NIA), people with MCI are at a greater risk of developing dementia, thus it is of great importance to detect MCI at the earliest possible to mitigate the transformation of MCI to Alzheimer’s and dementia. Recent studies have harnessed Artificial Intelligence (AI) to develop automated methods to predict and detect MCI. The majority of the existing research is based on unimodal data (e.g., only speech or prosody), but recent studies have shown that multimodality leads to a more accurate prediction of MCI. However, effectively exploiting different modalities is still a big challenge due to the lack of efficient fusion methods. This thesis proposes a mid-level fusion architecture to make use of multimodal data for MCI prediction. We introduce a multimodal speech-language-vision Deep Learning-based method to differentiate MCI from Normal Cognition (NC). Our proposed architecture includes co-attention blocks to fuse three different modalities at the embedding level to find the potential interactions between speech (audio), language (transcribed speech), and vision (facial videos) within the cross-Transformer layer. To study and evaluated the proposed mid-level fusion model, the I-CONECT dataset was used. It contains a large number of semi-structured conversations via the internet/webcam between participants aged 75+ years old and interviewers. Our experimental results show that the proposed fusion method can detect MCI from NC with an average AUC of (85.3%) which outperforms the unimodal and bimodal baseline models.&lt;/p&gt; &lt;p&gt;This thesis demonstrates that multimodal deep learning models outperform unimodal models in detecting MCI in older adults. To generalize the applicability of these findings, further research employing larger datasets should be conducted.&lt;/p&gt;","abstract_has_math":false,"creators":["Far Poor, Farida"],"institution":null,"degree_name":"M.S. in Computer Engineering","degree_level":"Masters Thesis","degree_discipline":null,"degree_department":null,"school":null,"contributors":["Mohammad H. Mahoor","Yun-Bo Yi","Kerstin Sophie Haring","Haluk Ogmen"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-06-15T07:00:00Z","date_published":"2024-06-15T07:00:00Z","updated_at":"2026-07-24T02:01:48Z","subjects":["Artificial intelligence","Cognitive impairment","Artificial Intelligence and Robotics","Cognitive Science","Computer Engineering","Engineering"],"languages":["English (eng)"],"rights":["<p>Copyright is held by the author. Permanently suppressed.</p>"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://digitalcommons.du.edu/etd/2449","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Mohammad H. Mahoor","Yun-Bo Yi","Kerstin Sophie Haring","Haluk Ogmen"]},{"key":"dc:creator","label":"Author","values":["Far Poor, Farida"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"thesis:degree_level","label":"Degree Level","values":["Masters Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S. in Computer Engineering"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Artificial intelligence","Cognitive impairment","Artificial Intelligence and Robotics","Cognitive Science","Computer Engineering","Engineering"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["English (eng)"]},{"key":"dc:rights","label":"Dc Rights","values":["<p>Copyright is held by the author. Permanently suppressed.</p>"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://digitalcommons.du.edu/etd/2449"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>Mild Cognitive Impairment (MCI) is an early stage of memory loss or other cognitive ability loss in individuals who maintain the ability to independently perform most activities of daily living. It is considered a transitional stage between normal cognitive stage and more severe cognitive declines like dementia or Alzheimer’s. Based on the reports from the National Institute of Aging (NIA), people with MCI are at a greater risk of developing dementia, thus it is of great importance to detect MCI at the earliest possible to mitigate the transformation of MCI to Alzheimer’s and dementia. Recent studies have harnessed Artificial Intelligence (AI) to develop automated methods to predict and detect MCI. The majority of the existing research is based on unimodal data (e.g., only speech or prosody), but recent studies have shown that multimodality leads to a more accurate prediction of MCI. However, effectively exploiting different modalities is still a big challenge due to the lack of efficient fusion methods. This thesis proposes a mid-level fusion architecture to make use of multimodal data for MCI prediction. We introduce a multimodal speech-language-vision Deep Learning-based method to differentiate MCI from Normal Cognition (NC). Our proposed architecture includes co-attention blocks to fuse three different modalities at the embedding level to find the potential interactions between speech (audio), language (transcribed speech), and vision (facial videos) within the cross-Transformer layer. To study and evaluated the proposed mid-level fusion model, the I-CONECT dataset was used. It contains a large number of semi-structured conversations via the internet/webcam between participants aged 75+ years old and interviewers. Our experimental results show that the proposed fusion method can detect MCI from NC with an average AUC of (85.3%) which outperforms the unimodal and bimodal baseline models.</p> <p>This thesis demonstrates that multimodal deep learning models outperform unimodal models in detecting MCI in older adults. To generalize the applicability of these findings, further research employing larger datasets should be conducted.</p>"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["A Study on Multimodal AI for Mild Cognitive Impairment Detection"]}]}],"canonical_facts":{"dc:contributor":["Mohammad H. Mahoor","Yun-Bo Yi","Kerstin Sophie Haring","Haluk Ogmen"],"dc:creator":["Far Poor, Farida"],"dc:description.abstract":["<p>Mild Cognitive Impairment (MCI) is an early stage of memory loss or other cognitive ability loss in individuals who maintain the ability to independently perform most activities of daily living. It is considered a transitional stage between normal cognitive stage and more severe cognitive declines like dementia or Alzheimer’s. Based on the reports from the National Institute of Aging (NIA), people with MCI are at a greater risk of developing dementia, thus it is of great importance to detect MCI at the earliest possible to mitigate the transformation of MCI to Alzheimer’s and dementia. Recent studies have harnessed Artificial Intelligence (AI) to develop automated methods to predict and detect MCI. The majority of the existing research is based on unimodal data (e.g., only speech or prosody), but recent studies have shown that multimodality leads to a more accurate prediction of MCI. However, effectively exploiting different modalities is still a big challenge due to the lack of efficient fusion methods. This thesis proposes a mid-level fusion architecture to make use of multimodal data for MCI prediction. We introduce a multimodal speech-language-vision Deep Learning-based method to differentiate MCI from Normal Cognition (NC). Our proposed architecture includes co-attention blocks to fuse three different modalities at the embedding level to find the potential interactions between speech (audio), language (transcribed speech), and vision (facial videos) within the cross-Transformer layer. To study and evaluated the proposed mid-level fusion model, the I-CONECT dataset was used. It contains a large number of semi-structured conversations via the internet/webcam between participants aged 75+ years old and interviewers. Our experimental results show that the proposed fusion method can detect MCI from NC with an average AUC of (85.3%) which outperforms the unimodal and bimodal baseline models.</p> <p>This thesis demonstrates that multimodal deep learning models outperform unimodal models in detecting MCI in older adults. To generalize the applicability of these findings, further research employing larger datasets should be conducted.</p>"],"dc:format":["application/pdf"],"dc:identifier":["https://digitalcommons.du.edu/etd/2449"],"dc:language":["English (eng)"],"dc:rights":["<p>Copyright is held by the author. Permanently suppressed.</p>"],"dc:subject":["Artificial intelligence","Cognitive impairment","Artificial Intelligence and Robotics","Cognitive Science","Computer Engineering","Engineering"],"dc:title":["A Study on Multimodal AI for Mild Cognitive Impairment Detection"],"thesis:degree_level":["Masters Thesis"],"thesis:degree_name":["M.S. in Computer Engineering"]},"updated_at":"2026-07-24T02:01:48Z"}