{"id":{"repo_id":"cambridge","oai_identifier":"oai:www.repository.cam.ac.uk:1810/399435"},"canonical_url":"https://search.dev.ndltd.org/etd/cambridge/oai:www.repository.cam.ac.uk:1810/399435","repository":{"repo_id":"cambridge","name":"Cambridge University","base_url":"https://api.repository.cam.ac.uk/server/oai/request"},"display":{"title":"Vitras, Proteomizer, GhostBuster - A journey from Wet Lab to Machine Learning and back, pursuing Parkinson’s Disease","abstract":"This project is highly cross-disciplinary, as it aims at bridging the divide between wet-lab-based molecular biology and pure computer science, with Parkinson’s Disease (PD) serving as the common scientific driving question. Growing evidence has shown that ~60% PD patients first develop alpha-synuclein (αS) pathology in the enteric nervous system, subsequently spreading to the vagal nuclei in the brainstem, and finally to the rest of the brain. To model such early disease stages, I contributed to the characterisation of the Vitras mouse model, which overexpresses human 1-120 truncated αS under the villin promoter. I showed that, unexpectedly, the Vitras also expresses αS in the olfactory system, and specifically in the vomeronasal organ; from there, the protein travels to the first synaptic station, that is the accessory olfactory bulb, but is not found in any downstream localisation. By subdiaphragmatic hemivagotomy, I also showed that αS is transported by the vagus nerve from the gut to the medulla. To the best of my knowledge, this provides the first direct, protein-level evidence of the vagal involvement in alpha-synucleinopathy, using endogenous αS. PD is associated with a dramatic mismatch between the gene expression levels measured at transcriptomics (Tx), and those measured and at proteomics (Px). While exacerbated in neurodegeneration, this phenomenon is true in general, as Tx-Px correlation is generally estimated at a mere r = 0.5 across genes, and r = 0.3 across samples. Building on this observation, I developed Proteomizer, a machine learning (ML) platform that utilises deep learning to infer a sample’s Px landscape, based on its Tx and miRNomic (Mx) ones. Trained on a combined 8,613 Tx-Mx-Px matched samples, Proteomizer improved Tx-Px correlation r from 0.27 to 0.68, which is the best performance available in literature for such task. I also developed a Monte Carlo method to simulate thousands of differential expression analyses, showing that proteomisation boosts differential expression agreement by up to 62×, and as high as 6 orders of magnitude for a select group of genes, mostly involved in aerobic respiration and ribosomes. Unfortunately, proteomisation benefit did not generalise to novel tissue types, or to datasets collected using a different protocol. I also employed three explainer architectures, which proved effective at picking up microRNAs involved in a gene’s Tx-Px mismatch, yielding a test dataset ROC-AUC of 0.74. Finally, a limiting factor in PD therapeutics lies in our limited knowledge of the genes involved in its mechanism. Worrying evidence shows that only a minority of human protein-coding genes are sufficiently characterised, the others acting as ‘ghost genes’. Unfortunately, due to research social structures, literature tends to focus increasingly more on genes that are already well annotated. To address this issue, I introduced GhostBuster, an encoder-decoder-based ML platform. GhostBuster was effective at predicting novel functions, disease involvement and interaction modalities for a given gene, agnostically of its current degree of characterisation, showcasing a test ROC-AUC of 0.8-0.95. When applied to PD, GhostBuster was capable of highlighting genes whose involvement in the disease was only established in the last 3 years. Overall, GhostBuster is the first literature-unbiased architecture of its kind, and represents an ideal tool for ghost gene characterisation.","abstract_html":"This project is highly cross-disciplinary, as it aims at bridging the divide between wet-lab-based molecular biology and pure computer science, with Parkinson’s Disease (PD) serving as the common scientific driving question. Growing evidence has shown that ~60% PD patients first develop alpha-synuclein (αS) pathology in the enteric nervous system, subsequently spreading to the vagal nuclei in the brainstem, and finally to the rest of the brain. To model such early disease stages, I contributed to the characterisation of the Vitras mouse model, which overexpresses human 1-120 truncated αS under the villin promoter. I showed that, unexpectedly, the Vitras also expresses αS in the olfactory system, and specifically in the vomeronasal organ; from there, the protein travels to the first synaptic station, that is the accessory olfactory bulb, but is not found in any downstream localisation. By subdiaphragmatic hemivagotomy, I also showed that αS is transported by the vagus nerve from the gut to the medulla. To the best of my knowledge, this provides the first direct, protein-level evidence of the vagal involvement in alpha-synucleinopathy, using endogenous αS. PD is associated with a dramatic mismatch between the gene expression levels measured at transcriptomics (Tx), and those measured and at proteomics (Px). While exacerbated in neurodegeneration, this phenomenon is true in general, as Tx-Px correlation is generally estimated at a mere r = 0.5 across genes, and r = 0.3 across samples. Building on this observation, I developed Proteomizer, a machine learning (ML) platform that utilises deep learning to infer a sample’s Px landscape, based on its Tx and miRNomic (Mx) ones. Trained on a combined 8,613 Tx-Mx-Px matched samples, Proteomizer improved Tx-Px correlation r from 0.27 to 0.68, which is the best performance available in literature for such task. I also developed a Monte Carlo method to simulate thousands of differential expression analyses, showing that proteomisation boosts differential expression agreement by up to 62×, and as high as 6 orders of magnitude for a select group of genes, mostly involved in aerobic respiration and ribosomes. Unfortunately, proteomisation benefit did not generalise to novel tissue types, or to datasets collected using a different protocol. I also employed three explainer architectures, which proved effective at picking up microRNAs involved in a gene’s Tx-Px mismatch, yielding a test dataset ROC-AUC of 0.74. Finally, a limiting factor in PD therapeutics lies in our limited knowledge of the genes involved in its mechanism. Worrying evidence shows that only a minority of human protein-coding genes are sufficiently characterised, the others acting as ‘ghost genes’. Unfortunately, due to research social structures, literature tends to focus increasingly more on genes that are already well annotated. To address this issue, I introduced GhostBuster, an encoder-decoder-based ML platform. GhostBuster was effective at predicting novel functions, disease involvement and interaction modalities for a given gene, agnostically of its current degree of characterisation, showcasing a test ROC-AUC of 0.8-0.95. When applied to PD, GhostBuster was capable of highlighting genes whose involvement in the disease was only established in the last 3 years. Overall, GhostBuster is the first literature-unbiased architecture of its kind, and represents an ideal tool for ghost gene characterisation.","abstract_has_math":false,"creators":["Deangeli, Giulio"],"institution":"University of Cambridge","degree_name":"Doctor of Philosophy (PhD)","degree_level":"Doctoral","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Spillantini, Maria Grazia","Lio, Pietro"],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-06-30","date_published":"2025-06-30","updated_at":"2026-07-22T22:23:59Z","subjects":["AI","Machine Learning","Neurodegeneration","Parkinson's disease","Proteomics","Transcriptomics","Multiomics","Deep learning"],"languages":["eng"],"rights":[],"rights_urls":["https://www.repository.cam.ac.uk/bitstreams/3a6b259d-d451-4943-83de-9d9cee3c5cb9/download","http://purl.org/NET/rdflicense/allrightsreserved"],"identifier_entries":[]},"links":{"outbound_url":"https://doi.org/10.17863/CAM.127983","outbound_label":"DOI","outbound_source":"dc:identifier.doi"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Spillantini, Maria Grazia","Lio, Pietro"]},{"key":"dc:contributor.sponsor","label":"Sponsor","values":["This research was supported by funding from the Cambridge Trust and the Medical Research Council Doctoral Training Partnership (MRC DTP), grant no. RG86932. Furthermore, the Vitras project was supported by the Aligning Science Across Parkinson’s (ASAP) initiative, through the Michael J. Fox Foundation for Parkinson’s Research. I would like to acknowledge their generous support in making this project possible."]},{"key":"dc:creator","label":"Author","values":["Deangeli, Giulio"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2025-06-30"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["University of Cambridge"]},{"key":"dc:relation.isreferencedby.uri","label":"Dc Relation Isreferencedby URI","values":["https://www.repository.cam.ac.uk/handle/1810/399435"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["Doctoral"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["Doctor of Philosophy (PhD)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["AI","Machine Learning","Neurodegeneration","Parkinson's disease","Proteomics","Transcriptomics","Multiomics","Deep learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["https://www.repository.cam.ac.uk/bitstreams/3a6b259d-d451-4943-83de-9d9cee3c5cb9/download","http://purl.org/NET/rdflicense/allrightsreserved"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.17863/CAM.127983"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://www.repository.cam.ac.uk/bitstreams/13cb0dec-3766-41e1-9160-0b7aa7c14ce2/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["This project is highly cross-disciplinary, as it aims at bridging the divide between wet-lab-based molecular biology and pure computer science, with Parkinson’s Disease (PD) serving as the common scientific driving question. Growing evidence has shown that ~60% PD patients first develop alpha-synuclein (αS) pathology in the enteric nervous system, subsequently spreading to the vagal nuclei in the brainstem, and finally to the rest of the brain. To model such early disease stages, I contributed to the characterisation of the Vitras mouse model, which overexpresses human 1-120 truncated αS under the villin promoter. I showed that, unexpectedly, the Vitras also expresses αS in the olfactory system, and specifically in the vomeronasal organ; from there, the protein travels to the first synaptic station, that is the accessory olfactory bulb, but is not found in any downstream localisation. By subdiaphragmatic hemivagotomy, I also showed that αS is transported by the vagus nerve from the gut to the medulla. To the best of my knowledge, this provides the first direct, protein-level evidence of the vagal involvement in alpha-synucleinopathy, using endogenous αS. PD is associated with a dramatic mismatch between the gene expression levels measured at transcriptomics (Tx), and those measured and at proteomics (Px). While exacerbated in neurodegeneration, this phenomenon is true in general, as Tx-Px correlation is generally estimated at a mere r = 0.5 across genes, and r = 0.3 across samples. Building on this observation, I developed Proteomizer, a machine learning (ML) platform that utilises deep learning to infer a sample’s Px landscape, based on its Tx and miRNomic (Mx) ones. Trained on a combined 8,613 Tx-Mx-Px matched samples, Proteomizer improved Tx-Px correlation r from 0.27 to 0.68, which is the best performance available in literature for such task. I also developed a Monte Carlo method to simulate thousands of differential expression analyses, showing that proteomisation boosts differential expression agreement by up to 62×, and as high as 6 orders of magnitude for a select group of genes, mostly involved in aerobic respiration and ribosomes. Unfortunately, proteomisation benefit did not generalise to novel tissue types, or to datasets collected using a different protocol. I also employed three explainer architectures, which proved effective at picking up microRNAs involved in a gene’s Tx-Px mismatch, yielding a test dataset ROC-AUC of 0.74. Finally, a limiting factor in PD therapeutics lies in our limited knowledge of the genes involved in its mechanism. Worrying evidence shows that only a minority of human protein-coding genes are sufficiently characterised, the others acting as ‘ghost genes’. Unfortunately, due to research social structures, literature tends to focus increasingly more on genes that are already well annotated. To address this issue, I introduced GhostBuster, an encoder-decoder-based ML platform. GhostBuster was effective at predicting novel functions, disease involvement and interaction modalities for a given gene, agnostically of its current degree of characterisation, showcasing a test ROC-AUC of 0.8-0.95. When applied to PD, GhostBuster was capable of highlighting genes whose involvement in the disease was only established in the last 3 years. Overall, GhostBuster is the first literature-unbiased architecture of its kind, and represents an ideal tool for ghost gene characterisation."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["1fa5fc2ee51b6119ff12bc2eb4a61b1c","87eda9de84448d1f82354d60eee3eb5f"]},{"key":"dc:title","label":"Title","values":["Vitras, Proteomizer, GhostBuster - A journey from Wet Lab to Machine Learning and back, pursuing Parkinson’s Disease"]}]}],"canonical_facts":{"dc:contributor.advisor":["Spillantini, Maria Grazia","Lio, Pietro"],"dc:contributor.sponsor":["This research was supported by funding from the Cambridge Trust and the Medical Research Council Doctoral Training Partnership (MRC DTP), grant no. RG86932. Furthermore, the Vitras project was supported by the Aligning Science Across Parkinson’s (ASAP) initiative, through the Michael J. Fox Foundation for Parkinson’s Research. I would like to acknowledge their generous support in making this project possible."],"dc:creator":["Deangeli, Giulio"],"dc:date.issued":["2025-06-30"],"dc:description.abstract":["This project is highly cross-disciplinary, as it aims at bridging the divide between wet-lab-based molecular biology and pure computer science, with Parkinson’s Disease (PD) serving as the common scientific driving question. Growing evidence has shown that ~60% PD patients first develop alpha-synuclein (αS) pathology in the enteric nervous system, subsequently spreading to the vagal nuclei in the brainstem, and finally to the rest of the brain. To model such early disease stages, I contributed to the characterisation of the Vitras mouse model, which overexpresses human 1-120 truncated αS under the villin promoter. I showed that, unexpectedly, the Vitras also expresses αS in the olfactory system, and specifically in the vomeronasal organ; from there, the protein travels to the first synaptic station, that is the accessory olfactory bulb, but is not found in any downstream localisation. By subdiaphragmatic hemivagotomy, I also showed that αS is transported by the vagus nerve from the gut to the medulla. To the best of my knowledge, this provides the first direct, protein-level evidence of the vagal involvement in alpha-synucleinopathy, using endogenous αS. PD is associated with a dramatic mismatch between the gene expression levels measured at transcriptomics (Tx), and those measured and at proteomics (Px). While exacerbated in neurodegeneration, this phenomenon is true in general, as Tx-Px correlation is generally estimated at a mere r = 0.5 across genes, and r = 0.3 across samples. Building on this observation, I developed Proteomizer, a machine learning (ML) platform that utilises deep learning to infer a sample’s Px landscape, based on its Tx and miRNomic (Mx) ones. Trained on a combined 8,613 Tx-Mx-Px matched samples, Proteomizer improved Tx-Px correlation r from 0.27 to 0.68, which is the best performance available in literature for such task. I also developed a Monte Carlo method to simulate thousands of differential expression analyses, showing that proteomisation boosts differential expression agreement by up to 62×, and as high as 6 orders of magnitude for a select group of genes, mostly involved in aerobic respiration and ribosomes. Unfortunately, proteomisation benefit did not generalise to novel tissue types, or to datasets collected using a different protocol. I also employed three explainer architectures, which proved effective at picking up microRNAs involved in a gene’s Tx-Px mismatch, yielding a test dataset ROC-AUC of 0.74. Finally, a limiting factor in PD therapeutics lies in our limited knowledge of the genes involved in its mechanism. Worrying evidence shows that only a minority of human protein-coding genes are sufficiently characterised, the others acting as ‘ghost genes’. Unfortunately, due to research social structures, literature tends to focus increasingly more on genes that are already well annotated. To address this issue, I introduced GhostBuster, an encoder-decoder-based ML platform. GhostBuster was effective at predicting novel functions, disease involvement and interaction modalities for a given gene, agnostically of its current degree of characterisation, showcasing a test ROC-AUC of 0.8-0.95. When applied to PD, GhostBuster was capable of highlighting genes whose involvement in the disease was only established in the last 3 years. Overall, GhostBuster is the first literature-unbiased architecture of its kind, and represents an ideal tool for ghost gene characterisation."],"dc:format.checksum.md5":["1fa5fc2ee51b6119ff12bc2eb4a61b1c","87eda9de84448d1f82354d60eee3eb5f"],"dc:identifier.doi":["https://doi.org/10.17863/CAM.127983"],"dc:identifier.uri":["https://www.repository.cam.ac.uk/bitstreams/13cb0dec-3766-41e1-9160-0b7aa7c14ce2/download"],"dc:language":["eng"],"dc:publisher.institution":["University of Cambridge"],"dc:relation.isreferencedby.uri":["https://www.repository.cam.ac.uk/handle/1810/399435"],"dc:rights":["https://www.repository.cam.ac.uk/bitstreams/3a6b259d-d451-4943-83de-9d9cee3c5cb9/download","http://purl.org/NET/rdflicense/allrightsreserved"],"dc:subject":["AI","Machine Learning","Neurodegeneration","Parkinson's disease","Proteomics","Transcriptomics","Multiomics","Deep learning"],"dc:title":["Vitras, Proteomizer, GhostBuster - A journey from Wet Lab to Machine Learning and back, pursuing Parkinson’s Disease"],"dc:type":["Thesis"],"dc:type.qualificationlevel":["Doctoral"],"dc:type.qualificationname":["Doctor of Philosophy (PhD)"]},"updated_at":"2026-07-22T22:23:59Z"}