{"id":{"repo_id":"chapman","oai_identifier":"oai:digitalcommons.chapman.edu:cads_theses-1008"},"canonical_url":"https://search.dev.ndltd.org/etd/chapman/oai:digitalcommons.chapman.edu:cads_theses-1008","repository":{"repo_id":"chapman","name":"Chapman University","base_url":"https://digitalcommons.chapman.edu/do/oai/"},"display":{"title":"Automated Parsing of Flexible Molecular Systems using Principal Component Analysis and K-Means Clustering Techniques","abstract":"<p>Computational investigation of molecular structures and reactions of biological and pharmaceutical interests remains a grand scientific challenge due to the size and conformational flexibility of these systems. The work requires parsing and analyzing thousands of conformations in each molecular state for meaningful chemical information and subjecting the ensemble to costly quantum chemical calculations. The current status quo typically involves a manual process where the investigator must look at each conformation, separating each into structural families. This process is time-intensive and tedious, making this process infeasible in some cases, and limiting the ability of theoreticians to study these systems. However, the use of computational software allows for the necessary exhaustive investigation without the bottlenecks of a brute force approach to each flexible system.</p> <p>I aim to create the solution to this problem. In my thesis project, I seek to develop a Python software that will (i) automate the parsing of each conformation within a conformational ensemble, (ii) use principal component analysis (PCA) and clustering to find and investigate conformational families within the ensemble, (iii) separate and visualize conformational families in a user-friendly manner, and (iv) convey to the user how conformational families were delineated by way of features found within data. Results explored this work show that the program has the ability to separate conformational families with varying ranges of difficulty.</p>","abstract_html":"&lt;p&gt;Computational investigation of molecular structures and reactions of biological and pharmaceutical interests remains a grand scientific challenge due to the size and conformational flexibility of these systems. The work requires parsing and analyzing thousands of conformations in each molecular state for meaningful chemical information and subjecting the ensemble to costly quantum chemical calculations. The current status quo typically involves a manual process where the investigator must look at each conformation, separating each into structural families. This process is time-intensive and tedious, making this process infeasible in some cases, and limiting the ability of theoreticians to study these systems. However, the use of computational software allows for the necessary exhaustive investigation without the bottlenecks of a brute force approach to each flexible system.&lt;/p&gt; &lt;p&gt;I aim to create the solution to this problem. In my thesis project, I seek to develop a Python software that will (i) automate the parsing of each conformation within a conformational ensemble, (ii) use principal component analysis (PCA) and clustering to find and investigate conformational families within the ensemble, (iii) separate and visualize conformational families in a user-friendly manner, and (iv) convey to the user how conformational families were delineated by way of features found within data. Results explored this work show that the program has the ability to separate conformational families with varying ranges of difficulty.&lt;/p&gt;","abstract_has_math":false,"creators":["Nwerem, Matthew J"],"institution":null,"degree_name":"Master of Science (MS)","degree_level":"Thesis","degree_discipline":"Computational and Data Sciences","degree_department":null,"school":null,"contributors":["Dr. O. Maduka Ogba","Dr. Gennady Verkhivker","Dr. Lindsay Waldrop"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2021,"date_issued":"2021-08-01T07:00:00Z","date_published":"2021-08-01T07:00:00Z","updated_at":"2026-07-24T01:38:16Z","subjects":["conformational analysis","PCA","clustering","computational chemistry","Organic Chemistry","Other Computer Sciences","Structural Biology"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://digitalcommons.chapman.edu/cads_theses/9","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Dr. O. Maduka Ogba","Dr. Gennady Verkhivker","Dr. Lindsay Waldrop"]},{"key":"dc:creator","label":"Author","values":["Nwerem, Matthew J"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"thesis:degree_discipline","label":"Discipline","values":["Computational and Data Sciences"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science (MS)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["conformational analysis","PCA","clustering","computational chemistry","Organic Chemistry","Other Computer Sciences","Structural Biology"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://digitalcommons.chapman.edu/cads_theses/9"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>Computational investigation of molecular structures and reactions of biological and pharmaceutical interests remains a grand scientific challenge due to the size and conformational flexibility of these systems. The work requires parsing and analyzing thousands of conformations in each molecular state for meaningful chemical information and subjecting the ensemble to costly quantum chemical calculations. The current status quo typically involves a manual process where the investigator must look at each conformation, separating each into structural families. This process is time-intensive and tedious, making this process infeasible in some cases, and limiting the ability of theoreticians to study these systems. However, the use of computational software allows for the necessary exhaustive investigation without the bottlenecks of a brute force approach to each flexible system.</p> <p>I aim to create the solution to this problem. In my thesis project, I seek to develop a Python software that will (i) automate the parsing of each conformation within a conformational ensemble, (ii) use principal component analysis (PCA) and clustering to find and investigate conformational families within the ensemble, (iii) separate and visualize conformational families in a user-friendly manner, and (iv) convey to the user how conformational families were delineated by way of features found within data. Results explored this work show that the program has the ability to separate conformational families with varying ranges of difficulty.</p>"]},{"key":"dc:source","label":"Dc Source","values":["M. Nwerem, \"Automated parsing of flexible molecular systems using principal component analysis and K-means clustering techniques,\" M. S. thesis, Chapman University, Orange, CA, 2021. <a href=\"https://doi.org/10.36837/chapman.000293\">https://doi.org/10.36837/chapman.000293</a>"]},{"key":"dc:title","label":"Title","values":["Automated Parsing of Flexible Molecular Systems using Principal Component Analysis and K-Means Clustering Techniques"]}]}],"canonical_facts":{"dc:contributor":["Dr. O. Maduka Ogba","Dr. Gennady Verkhivker","Dr. Lindsay Waldrop"],"dc:creator":["Nwerem, Matthew J"],"dc:description.abstract":["<p>Computational investigation of molecular structures and reactions of biological and pharmaceutical interests remains a grand scientific challenge due to the size and conformational flexibility of these systems. The work requires parsing and analyzing thousands of conformations in each molecular state for meaningful chemical information and subjecting the ensemble to costly quantum chemical calculations. The current status quo typically involves a manual process where the investigator must look at each conformation, separating each into structural families. This process is time-intensive and tedious, making this process infeasible in some cases, and limiting the ability of theoreticians to study these systems. However, the use of computational software allows for the necessary exhaustive investigation without the bottlenecks of a brute force approach to each flexible system.</p> <p>I aim to create the solution to this problem. In my thesis project, I seek to develop a Python software that will (i) automate the parsing of each conformation within a conformational ensemble, (ii) use principal component analysis (PCA) and clustering to find and investigate conformational families within the ensemble, (iii) separate and visualize conformational families in a user-friendly manner, and (iv) convey to the user how conformational families were delineated by way of features found within data. Results explored this work show that the program has the ability to separate conformational families with varying ranges of difficulty.</p>"],"dc:identifier":["https://digitalcommons.chapman.edu/cads_theses/9"],"dc:source":["M. Nwerem, \"Automated parsing of flexible molecular systems using principal component analysis and K-means clustering techniques,\" M. S. thesis, Chapman University, Orange, CA, 2021. <a href=\"https://doi.org/10.36837/chapman.000293\">https://doi.org/10.36837/chapman.000293</a>"],"dc:subject":["conformational analysis","PCA","clustering","computational chemistry","Organic Chemistry","Other Computer Sciences","Structural Biology"],"dc:title":["Automated Parsing of Flexible Molecular Systems using Principal Component Analysis and K-Means Clustering Techniques"],"thesis:degree_discipline":["Computational and Data Sciences"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["Master of Science (MS)"]},"updated_at":"2026-07-24T01:38:16Z"}