{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/106165"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/106165","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Evolution of protein structure function","abstract":"Proteins are generated as a result of millions of evolutionary experiments over a billion years. These experiments are constrained by proteins structure and function. Investigating genomics data is a natural way for mining these constraints to address long-standing biological challenges. In this dissertation, first, we tried to address two specific biological questions by reconstructing the evolutionary pathways. We simulated Cyclin-dependent kinase 2 (CDK2) and its experimentally validated Cyclin-independent ancestor to understand the atomistic basis of Cyclin dependence in a kinase family. A set of conformational differences and the corresponding residues were identified to be the origin of different activation processes between CDK2 and its ancestor. The second biological question was to understand how does the wonder drug of century, Imatinib (Gleevec), exhibit high selectivity for a particular kinase involved in Chronic myelogenous leukemia. We investigated this question by simulating two modern kinases, a strong (Abl) and a weak (Src) binder, with five of their common ancestors to regenerate the evolutionary pathway between them. In the second effort to use evolutionary information, we exploited large-scale genomic sequences along with machine learning techniques to address one of the biggest challenges facing Molecular Dynamics (MD) simulations. MD simulations are shown to be valuable tools for study of proteins at atomic resolution. How- ever, in practice, running these simulations is computationally expensive due to the long timescales associated with biologically relevant conformational changes. Here, we developed a reinforcement learning-based sampling algorithm to enhance the MD simulation by prioritizing the most important reaction coordinates at each stage. Testing the proposed algorithm on multiple case studies showed a significant improvement over other unbiased sampling techniques. Furthermore, we showed the distances between evolutionary coupling residues can be a natural and effective set of reaction coordinates to be used for reinforcement learning based adaptive sampling. At last, we went one step further and applied transfer learning on a genomic-based model to predict the effects of mutation more efficiently. We evaluated the effectiveness of the proposed transfer learning techniques in three different cases.","abstract_html":"Proteins are generated as a result of millions of evolutionary experiments over a billion years. These experiments are constrained by proteins structure and function. Investigating genomics data is a natural way for mining these constraints to address long-standing biological challenges. In this dissertation, first, we tried to address two specific biological questions by reconstructing the evolutionary pathways. We simulated Cyclin-dependent kinase 2 (CDK2) and its experimentally validated Cyclin-independent ancestor to understand the atomistic basis of Cyclin dependence in a kinase family. A set of conformational differences and the corresponding residues were identified to be the origin of different activation processes between CDK2 and its ancestor. The second biological question was to understand how does the wonder drug of century, Imatinib (Gleevec), exhibit high selectivity for a particular kinase involved in Chronic myelogenous leukemia. We investigated this question by simulating two modern kinases, a strong (Abl) and a weak (Src) binder, with five of their common ancestors to regenerate the evolutionary pathway between them. In the second effort to use evolutionary information, we exploited large-scale genomic sequences along with machine learning techniques to address one of the biggest challenges facing Molecular Dynamics (MD) simulations. MD simulations are shown to be valuable tools for study of proteins at atomic resolution. How- ever, in practice, running these simulations is computationally expensive due to the long timescales associated with biologically relevant conformational changes. Here, we developed a reinforcement learning-based sampling algorithm to enhance the MD simulation by prioritizing the most important reaction coordinates at each stage. Testing the proposed algorithm on multiple case studies showed a significant improvement over other unbiased sampling techniques. Furthermore, we showed the distances between evolutionary coupling residues can be a natural and effective set of reaction coordinates to be used for reinforcement learning based adaptive sampling. At last, we went one step further and applied transfer learning on a genomic-based model to predict the effects of mutation more efficiently. We evaluated the effectiveness of the proposed transfer learning techniques in three different cases.","abstract_has_math":false,"creators":["Shamsi, Zahra"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Chemical Engineering","degree_department":null,"school":null,"contributors":["Shukla, Diwakar","Kraft, Mary","Procko, Eric","Harley, Brendan"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2020,"date_issued":"2020-03-02T21:58:02Z","date_published":"2020-03-02T21:58:02Z","updated_at":"2026-07-22T22:24:45Z","subjects":["Cancer, Kinase, Imatinib, Molecular Dynamics Simulations, Computational Biology, Reinforcement Learning, Transfer Learning, Evolutionary Coupling, Evolution, Proteomics"],"languages":["en"],"rights":["Copyright 2019 Zahra Shamsi"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/106165","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Shukla, Diwakar","Kraft, Mary","Procko, Eric","Harley, Brendan"]},{"key":"dc:creator","label":"Author","values":["Shamsi, Zahra"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2020-03-02T21:58:02Z","2019-10-04","2019-12"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Chemical Engineering"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Cancer, Kinase, Imatinib, Molecular Dynamics Simulations, Computational Biology, Reinforcement Learning, Transfer Learning, Evolutionary Coupling, Evolution, Proteomics"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2019 Zahra Shamsi"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/106165"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Proteins are generated as a result of millions of evolutionary experiments over a billion years. These experiments are constrained by proteins structure and function. Investigating genomics data is a natural way for mining these constraints to address long-standing biological challenges. In this dissertation, first, we tried to address two specific biological questions by reconstructing the evolutionary pathways. We simulated Cyclin-dependent kinase 2 (CDK2) and its experimentally validated Cyclin-independent ancestor to understand the atomistic basis of Cyclin dependence in a kinase family. A set of conformational differences and the corresponding residues were identified to be the origin of different activation processes between CDK2 and its ancestor. The second biological question was to understand how does the wonder drug of century, Imatinib (Gleevec), exhibit high selectivity for a particular kinase involved in Chronic myelogenous leukemia. We investigated this question by simulating two modern kinases, a strong (Abl) and a weak (Src) binder, with five of their common ancestors to regenerate the evolutionary pathway between them. In the second effort to use evolutionary information, we exploited large-scale genomic sequences along with machine learning techniques to address one of the biggest challenges facing Molecular Dynamics (MD) simulations. MD simulations are shown to be valuable tools for study of proteins at atomic resolution. How- ever, in practice, running these simulations is computationally expensive due to the long timescales associated with biologically relevant conformational changes. Here, we developed a reinforcement learning-based sampling algorithm to enhance the MD simulation by prioritizing the most important reaction coordinates at each stage. Testing the proposed algorithm on multiple case studies showed a significant improvement over other unbiased sampling techniques. Furthermore, we showed the distances between evolutionary coupling residues can be a natural and effective set of reaction coordinates to be used for reinforcement learning based adaptive sampling. At last, we went one step further and applied transfer learning on a genomic-based model to predict the effects of mutation more efficiently. We evaluated the effectiveness of the proposed transfer learning techniques in three different cases.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2020-02-28 without embargo terms","The student, Zahra Shamsi, accepted the attached license on 2019-10-01 at 21:15.","The student, Zahra Shamsi, submitted this Dissertation for approval on 2019-10-01 at 21:30.","This Dissertation was approved for publication on 2019-10-04 at 15:48.","DSpace SAF Submission Ingestion Package generated from Vireo submission #14481 on 2020-02-28 at 17:12:07","Made available in DSpace on 2020-03-02T21:58:02Z (GMT). No. of bitstreams: 3 SHAMSI-DISSERTATION-2019.pdf: 181857577 bytes, checksum: 1db2845f72d4a2191add1056bc954b2a (MD5) LICENSE.txt: 4209 bytes, checksum: 46328fbf0e8a3cd40a978a984a02d0ec (MD5) PROQUEST_LICENSE.txt: 4555 bytes, checksum: 1cff2becf9514a2a08759b9a021ee49c (MD5) Previous issue date: 2019-10-04"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Evolution of protein structure function"]}]}],"canonical_facts":{"dc:contributor":["Shukla, Diwakar","Kraft, Mary","Procko, Eric","Harley, Brendan"],"dc:creator":["Shamsi, Zahra"],"dc:date":["2020-03-02T21:58:02Z","2019-10-04","2019-12"],"dc:description":["Proteins are generated as a result of millions of evolutionary experiments over a billion years. These experiments are constrained by proteins structure and function. Investigating genomics data is a natural way for mining these constraints to address long-standing biological challenges. In this dissertation, first, we tried to address two specific biological questions by reconstructing the evolutionary pathways. We simulated Cyclin-dependent kinase 2 (CDK2) and its experimentally validated Cyclin-independent ancestor to understand the atomistic basis of Cyclin dependence in a kinase family. A set of conformational differences and the corresponding residues were identified to be the origin of different activation processes between CDK2 and its ancestor. The second biological question was to understand how does the wonder drug of century, Imatinib (Gleevec), exhibit high selectivity for a particular kinase involved in Chronic myelogenous leukemia. We investigated this question by simulating two modern kinases, a strong (Abl) and a weak (Src) binder, with five of their common ancestors to regenerate the evolutionary pathway between them. In the second effort to use evolutionary information, we exploited large-scale genomic sequences along with machine learning techniques to address one of the biggest challenges facing Molecular Dynamics (MD) simulations. MD simulations are shown to be valuable tools for study of proteins at atomic resolution. How- ever, in practice, running these simulations is computationally expensive due to the long timescales associated with biologically relevant conformational changes. Here, we developed a reinforcement learning-based sampling algorithm to enhance the MD simulation by prioritizing the most important reaction coordinates at each stage. Testing the proposed algorithm on multiple case studies showed a significant improvement over other unbiased sampling techniques. Furthermore, we showed the distances between evolutionary coupling residues can be a natural and effective set of reaction coordinates to be used for reinforcement learning based adaptive sampling. At last, we went one step further and applied transfer learning on a genomic-based model to predict the effects of mutation more efficiently. We evaluated the effectiveness of the proposed transfer learning techniques in three different cases.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2020-02-28 without embargo terms","The student, Zahra Shamsi, accepted the attached license on 2019-10-01 at 21:15.","The student, Zahra Shamsi, submitted this Dissertation for approval on 2019-10-01 at 21:30.","This Dissertation was approved for publication on 2019-10-04 at 15:48.","DSpace SAF Submission Ingestion Package generated from Vireo submission #14481 on 2020-02-28 at 17:12:07","Made available in DSpace on 2020-03-02T21:58:02Z (GMT). No. of bitstreams: 3 SHAMSI-DISSERTATION-2019.pdf: 181857577 bytes, checksum: 1db2845f72d4a2191add1056bc954b2a (MD5) LICENSE.txt: 4209 bytes, checksum: 46328fbf0e8a3cd40a978a984a02d0ec (MD5) PROQUEST_LICENSE.txt: 4555 bytes, checksum: 1cff2becf9514a2a08759b9a021ee49c (MD5) Previous issue date: 2019-10-04"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/106165"],"dc:language":["en"],"dc:rights":["Copyright 2019 Zahra Shamsi"],"dc:subject":["Cancer, Kinase, Imatinib, Molecular Dynamics Simulations, Computational Biology, Reinforcement Learning, Transfer Learning, Evolutionary Coupling, Evolution, Proteomics"],"dc:title":["Evolution of protein structure function"],"dc:type":["text"],"thesis:degree_discipline":["Chemical Engineering"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:45Z"}