{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/113054"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/113054","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Application of machine learning to the detection and prediction of Parkinson’s disease subtypes","abstract":"Background: The clinical manifestations of Parkinson’s disease (PD) are characterized by heterogeneity in age at onset, disease duration, rate of progression, and the constellation of motor versus non-motor features. As such, counseling of patients about their prognosis is guarded. Characterization of unique disease subtypes and enhanced, personalized disease course projections are both unmet needs. Machine learning’s ability to ﬁnd hidden patterns in complicated, multi-dimensional datasets has opened up unprecedented possibilities for meeting this crucial demand. Methods and Findings: We used unsupervised and supervised machine learning methods on comprehensive, longitudinal clinical data from the Parkinson’s Disease Progression Marker Initiative (PPMI) (n = 294 cases) to identify patient subtypes and predict dis-ease progression. The resulting models were validated in an independent, clinically well-characterized cohort from the Parkinson’s Disease Biomarker Program (PDBP) (n = 263 cases). Our analysis distinguished three distinct disease subtypes with highly predictable progression rates, corresponding to slow, moderate, and fast disease progressors. In addition, we achieved highly accurate projections of disease progression ﬁve years after initial diagnosis with an average area under the curve (AUC) of 0.92 (95% Conﬁdence Interval (CI): 0.95 ± 0.01 for slow PD progressors (PDvec1), 0.87 ± 0.03 for model PD progressors (PDvec2), and 0.95 ± 0.02 for fast disease progressors (PDvec3)). Then, we replicated these ﬁndings in an independent validation cohort, released the analytical code, and developed models in an open science manner. Finally, we found that daytime sleepiness has the highest importance in clinical progression, followed by hobbies, activities, and urinary problems. Furthermore, we observe that Serum neuroﬁlament light levels and genetic risk score also play an essential role in predicting PD subtypes. Conclusions: These data-driven results could help deconstruct the heterogeneity of patients. This discovery could have immediate clinical trial consequences by increasing the discovery of signiﬁcant clinical outcomes previously hidden due to cohort heterogeneity. Machine learning models are expected to improve patient counseling, clinical trial design, healthcare resource allocation, and, ultimately, personalized patient care. To achieve more powerful predictions, we propose that these datasets should include standardized phenotype collection and recording.","abstract_html":"Background: The clinical manifestations of Parkinson’s disease (PD) are characterized by heterogeneity in age at onset, disease duration, rate of progression, and the constellation of motor versus non-motor features. As such, counseling of patients about their prognosis is guarded. Characterization of unique disease subtypes and enhanced, personalized disease course projections are both unmet needs. Machine learning’s ability to ﬁnd hidden patterns in complicated, multi-dimensional datasets has opened up unprecedented possibilities for meeting this crucial demand. Methods and Findings: We used unsupervised and supervised machine learning methods on comprehensive, longitudinal clinical data from the Parkinson’s Disease Progression Marker Initiative (PPMI) (n = 294 cases) to identify patient subtypes and predict dis-ease progression. The resulting models were validated in an independent, clinically well-characterized cohort from the Parkinson’s Disease Biomarker Program (PDBP) (n = 263 cases). Our analysis distinguished three distinct disease subtypes with highly predictable progression rates, corresponding to slow, moderate, and fast disease progressors. In addition, we achieved highly accurate projections of disease progression ﬁve years after initial diagnosis with an average area under the curve (AUC) of 0.92 (95% Conﬁdence Interval (CI): 0.95 ± 0.01 for slow PD progressors (PDvec1), 0.87 ± 0.03 for model PD progressors (PDvec2), and 0.95 ± 0.02 for fast disease progressors (PDvec3)). Then, we replicated these ﬁndings in an independent validation cohort, released the analytical code, and developed models in an open science manner. Finally, we found that daytime sleepiness has the highest importance in clinical progression, followed by hobbies, activities, and urinary problems. Furthermore, we observe that Serum neuroﬁlament light levels and genetic risk score also play an essential role in predicting PD subtypes. Conclusions: These data-driven results could help deconstruct the heterogeneity of patients. This discovery could have immediate clinical trial consequences by increasing the discovery of signiﬁcant clinical outcomes previously hidden due to cohort heterogeneity. Machine learning models are expected to improve patient counseling, clinical trial design, healthcare resource allocation, and, ultimately, personalized patient care. To achieve more powerful predictions, we propose that these datasets should include standardized phenotype collection and recording.","abstract_has_math":false,"creators":["Dadu, Anant"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Campbell, Roy H."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-01-12T21:46:49Z","date_published":"2022-01-12T21:46:49Z","updated_at":"2026-07-22T22:24:52Z","subjects":["Neurodegenerative Disorders, Machine Learning, Disease Progression"],"languages":["en"],"rights":["Copyright 2021 Anant Dadu"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/113054","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Campbell, Roy H."]},{"key":"dc:creator","label":"Author","values":["Dadu, Anant"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022-01-12T21:46:49Z","2021-07-19","2021-08"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Neurodegenerative Disorders, Machine Learning, Disease Progression"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2021 Anant Dadu"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/113054"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Background: The clinical manifestations of Parkinson’s disease (PD) are characterized by heterogeneity in age at onset, disease duration, rate of progression, and the constellation of motor versus non-motor features. As such, counseling of patients about their prognosis is guarded. Characterization of unique disease subtypes and enhanced, personalized disease course projections are both unmet needs. Machine learning’s ability to ﬁnd hidden patterns in complicated, multi-dimensional datasets has opened up unprecedented possibilities for meeting this crucial demand. Methods and Findings: We used unsupervised and supervised machine learning methods on comprehensive, longitudinal clinical data from the Parkinson’s Disease Progression Marker Initiative (PPMI) (n = 294 cases) to identify patient subtypes and predict dis-ease progression. The resulting models were validated in an independent, clinically well-characterized cohort from the Parkinson’s Disease Biomarker Program (PDBP) (n = 263 cases). Our analysis distinguished three distinct disease subtypes with highly predictable progression rates, corresponding to slow, moderate, and fast disease progressors. In addition, we achieved highly accurate projections of disease progression ﬁve years after initial diagnosis with an average area under the curve (AUC) of 0.92 (95% Conﬁdence Interval (CI): 0.95 ± 0.01 for slow PD progressors (PDvec1), 0.87 ± 0.03 for model PD progressors (PDvec2), and 0.95 ± 0.02 for fast disease progressors (PDvec3)). Then, we replicated these ﬁndings in an independent validation cohort, released the analytical code, and developed models in an open science manner. Finally, we found that daytime sleepiness has the highest importance in clinical progression, followed by hobbies, activities, and urinary problems. Furthermore, we observe that Serum neuroﬁlament light levels and genetic risk score also play an essential role in predicting PD subtypes. Conclusions: These data-driven results could help deconstruct the heterogeneity of patients. This discovery could have immediate clinical trial consequences by increasing the discovery of signiﬁcant clinical outcomes previously hidden due to cohort heterogeneity. Machine learning models are expected to improve patient counseling, clinical trial design, healthcare resource allocation, and, ultimately, personalized patient care. To achieve more powerful predictions, we propose that these datasets should include standardized phenotype collection and recording.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-01-12 without embargo terms","The student, Anant Dadu, accepted the attached license on 2021-07-15 at 13:09.","The student, Anant Dadu, submitted this Thesis for approval on 2021-07-15 at 13:11.","This Thesis was approved for publication on 2021-07-19 at 13:29.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16956 on 2022-01-12 at 12:45:54","Made available in DSpace on 2022-01-12T21:46:49Z (GMT). No. of bitstreams: 2 DADU-THESIS-2021.pdf: 4881619 bytes, checksum: 0c143fe35ca947c57e70f1e54a5e9185 (MD5) LICENSE.txt: 4207 bytes, checksum: b9b18612d0bc4b9e8a21f5e2e151f2ef (MD5) Previous issue date: 2021-07-19"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Application of machine learning to the detection and prediction of Parkinson’s disease subtypes"]}]}],"canonical_facts":{"dc:contributor":["Campbell, Roy H."],"dc:creator":["Dadu, Anant"],"dc:date":["2022-01-12T21:46:49Z","2021-07-19","2021-08"],"dc:description":["Background: The clinical manifestations of Parkinson’s disease (PD) are characterized by heterogeneity in age at onset, disease duration, rate of progression, and the constellation of motor versus non-motor features. As such, counseling of patients about their prognosis is guarded. Characterization of unique disease subtypes and enhanced, personalized disease course projections are both unmet needs. Machine learning’s ability to ﬁnd hidden patterns in complicated, multi-dimensional datasets has opened up unprecedented possibilities for meeting this crucial demand. Methods and Findings: We used unsupervised and supervised machine learning methods on comprehensive, longitudinal clinical data from the Parkinson’s Disease Progression Marker Initiative (PPMI) (n = 294 cases) to identify patient subtypes and predict dis-ease progression. The resulting models were validated in an independent, clinically well-characterized cohort from the Parkinson’s Disease Biomarker Program (PDBP) (n = 263 cases). Our analysis distinguished three distinct disease subtypes with highly predictable progression rates, corresponding to slow, moderate, and fast disease progressors. In addition, we achieved highly accurate projections of disease progression ﬁve years after initial diagnosis with an average area under the curve (AUC) of 0.92 (95% Conﬁdence Interval (CI): 0.95 ± 0.01 for slow PD progressors (PDvec1), 0.87 ± 0.03 for model PD progressors (PDvec2), and 0.95 ± 0.02 for fast disease progressors (PDvec3)). Then, we replicated these ﬁndings in an independent validation cohort, released the analytical code, and developed models in an open science manner. Finally, we found that daytime sleepiness has the highest importance in clinical progression, followed by hobbies, activities, and urinary problems. Furthermore, we observe that Serum neuroﬁlament light levels and genetic risk score also play an essential role in predicting PD subtypes. Conclusions: These data-driven results could help deconstruct the heterogeneity of patients. This discovery could have immediate clinical trial consequences by increasing the discovery of signiﬁcant clinical outcomes previously hidden due to cohort heterogeneity. Machine learning models are expected to improve patient counseling, clinical trial design, healthcare resource allocation, and, ultimately, personalized patient care. To achieve more powerful predictions, we propose that these datasets should include standardized phenotype collection and recording.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-01-12 without embargo terms","The student, Anant Dadu, accepted the attached license on 2021-07-15 at 13:09.","The student, Anant Dadu, submitted this Thesis for approval on 2021-07-15 at 13:11.","This Thesis was approved for publication on 2021-07-19 at 13:29.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16956 on 2022-01-12 at 12:45:54","Made available in DSpace on 2022-01-12T21:46:49Z (GMT). No. of bitstreams: 2 DADU-THESIS-2021.pdf: 4881619 bytes, checksum: 0c143fe35ca947c57e70f1e54a5e9185 (MD5) LICENSE.txt: 4207 bytes, checksum: b9b18612d0bc4b9e8a21f5e2e151f2ef (MD5) Previous issue date: 2021-07-19"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/113054"],"dc:language":["en"],"dc:rights":["Copyright 2021 Anant Dadu"],"dc:subject":["Neurodegenerative Disorders, Machine Learning, Disease Progression"],"dc:title":["Application of machine learning to the detection and prediction of Parkinson’s disease subtypes"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:52Z"}