{"id":{"repo_id":"colo-mines","oai_identifier":"oai:repository.mines.edu:11124/171241"},"canonical_url":"https://search.dev.ndltd.org/etd/colo-mines/oai:repository.mines.edu:11124/171241","repository":{"repo_id":"colo-mines","name":"Colorado School of Mines","base_url":"https://repository.mines.edu/server/oai/request"},"display":{"title":"Faster isomer network generation","abstract":"Isomer networks provide a mechanism to understand and interpret relationships between organic molecules with applications in medicinal chemistry and drug design. The extraction of isomer networks is a time and data-intensive computation. The contributions of this dissertation are a variety of techniques to more efficiently (with respect to time and memory) compute isomers networks. Specifically, we describe our efforts to improve the network extraction process by 1) Using the symmetry present in most molecules to reduce run time and memory and streamlining the algorithm used for the detection of duplicate canonical names, a key step in determining the bond count distances between pairs of isomers. Together, these techniques result in reductions in memory of up to 60% and improvements in runtime of up to a factor of 100. 2) Developing an optimal grouping algorithm to subdivide an all-all computation with large memory requirements. The algorithm provides a solution to sub divide the \"big data\" problem that arises in the construction of isomer networks into several independent \"small data\" problems. Our results show that using the grouping algorithm can help divide large data sets into independent smaller ones that can be processed in parallel. 3) Generating the isomer network for 1,050,125 isomers of Nicotine (with a preliminary analysis of the same) using the cloud computing capabilities of Amazon Web Services and Microsoft Azure. These techniques can also be employed to successfully compute isomers networks for other chemical compounds.","abstract_html":"Isomer networks provide a mechanism to understand and interpret relationships between organic molecules with applications in medicinal chemistry and drug design. The extraction of isomer networks is a time and data-intensive computation. The contributions of this dissertation are a variety of techniques to more efficiently (with respect to time and memory) compute isomers networks. Specifically, we describe our efforts to improve the network extraction process by 1) Using the symmetry present in most molecules to reduce run time and memory and streamlining the algorithm used for the detection of duplicate canonical names, a key step in determining the bond count distances between pairs of isomers. Together, these techniques result in reductions in memory of up to 60% and improvements in runtime of up to a factor of 100. 2) Developing an optimal grouping algorithm to subdivide an all-all computation with large memory requirements. The algorithm provides a solution to sub divide the &quot;big data&quot; problem that arises in the construction of isomer networks into several independent &quot;small data&quot; problems. Our results show that using the grouping algorithm can help divide large data sets into independent smaller ones that can be processed in parallel. 3) Generating the isomer network for 1,050,125 isomers of Nicotine (with a preliminary analysis of the same) using the cloud computing capabilities of Amazon Web Services and Microsoft Azure. These techniques can also be employed to successfully compute isomers networks for other chemical compounds.","abstract_has_math":false,"creators":["Thiagarajan, Dheivya"],"institution":"Colorado School of Mines. Arthur Lakes Library","degree_name":"Doctor of Philosophy (Ph.D.)","degree_level":"Doctoral","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":[],"advisors":["Mehta, Dinesh P."],"committee_chairs":[],"committee_members":["Han, Qi","Wu, Bo","Ciobanu, Cristian V."],"year":2017,"date_issued":"2017","date_published":"2017","updated_at":"2026-07-24T01:43:08Z","subjects":["cheminformatics","network analytics","big data","symmetry","isomers"],"languages":["eng","English"],"rights":["Copyright of the original work is retained by the author."],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["T 8330"],"render_values":[{"text":"T 8330","href":null,"code":true}]}]},"links":{"outbound_url":"https://hdl.handle.net/11124/171241","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Mehta, Dinesh P."]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Han, Qi","Wu, Bo","Ciobanu, Cristian V."]},{"key":"dc:creator","label":"Author","values":["Thiagarajan, Dheivya"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2017-07-31T16:32:51Z","2022-02-03T13:00:56Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2017-07-31T16:32:51Z","2022-02-03T13:00:56Z"]},{"key":"dc:date.issued","label":"Date","values":["2017"]},{"key":"dc:publisher","label":"Institution","values":["Colorado School of Mines. Arthur Lakes Library"]},{"key":"dc:type","label":"Dc Type","values":["Text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Doctoral"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctor of Philosophy (Ph.D.)"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Colorado School of Mines"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["cheminformatics","network analytics","big data","symmetry","isomers"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["English"]},{"key":"dc:language.iso","label":"Language (ISO)","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright of the original work is retained by the author."]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["Thiagarajan_mines_0052E_11303.pdf","T 8330"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/11124/171241"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Includes bibliographical references.","2017 Summer."]},{"key":"dc:description.abstract","label":"Abstract","values":["Isomer networks provide a mechanism to understand and interpret relationships between organic molecules with applications in medicinal chemistry and drug design. The extraction of isomer networks is a time and data-intensive computation. The contributions of this dissertation are a variety of techniques to more efficiently (with respect to time and memory) compute isomers networks. Specifically, we describe our efforts to improve the network extraction process by 1) Using the symmetry present in most molecules to reduce run time and memory and streamlining the algorithm used for the detection of duplicate canonical names, a key step in determining the bond count distances between pairs of isomers. Together, these techniques result in reductions in memory of up to 60% and improvements in runtime of up to a factor of 100. 2) Developing an optimal grouping algorithm to subdivide an all-all computation with large memory requirements. The algorithm provides a solution to sub divide the \"big data\" problem that arises in the construction of isomer networks into several independent \"small data\" problems. Our results show that using the grouping algorithm can help divide large data sets into independent smaller ones that can be processed in parallel. 3) Generating the isomer network for 1,050,125 isomers of Nicotine (with a preliminary analysis of the same) using the cloud computing capabilities of Amazon Web Services and Microsoft Azure. These techniques can also be employed to successfully compute isomers networks for other chemical compounds."]},{"key":"dc:format.medium","label":"Dc Format Medium","values":["born digital","doctoral dissertations"]},{"key":"dc:title","label":"Title","values":["Faster isomer network generation"]}]}],"canonical_facts":{"dc:contributor.advisor":["Mehta, Dinesh P."],"dc:contributor.committeemember":["Han, Qi","Wu, Bo","Ciobanu, Cristian V."],"dc:creator":["Thiagarajan, Dheivya"],"dc:date.accessioned":["2017-07-31T16:32:51Z","2022-02-03T13:00:56Z"],"dc:date.available":["2017-07-31T16:32:51Z","2022-02-03T13:00:56Z"],"dc:date.issued":["2017"],"dc:description":["Includes bibliographical references.","2017 Summer."],"dc:description.abstract":["Isomer networks provide a mechanism to understand and interpret relationships between organic molecules with applications in medicinal chemistry and drug design. The extraction of isomer networks is a time and data-intensive computation. The contributions of this dissertation are a variety of techniques to more efficiently (with respect to time and memory) compute isomers networks. Specifically, we describe our efforts to improve the network extraction process by 1) Using the symmetry present in most molecules to reduce run time and memory and streamlining the algorithm used for the detection of duplicate canonical names, a key step in determining the bond count distances between pairs of isomers. Together, these techniques result in reductions in memory of up to 60% and improvements in runtime of up to a factor of 100. 2) Developing an optimal grouping algorithm to subdivide an all-all computation with large memory requirements. The algorithm provides a solution to sub divide the \"big data\" problem that arises in the construction of isomer networks into several independent \"small data\" problems. Our results show that using the grouping algorithm can help divide large data sets into independent smaller ones that can be processed in parallel. 3) Generating the isomer network for 1,050,125 isomers of Nicotine (with a preliminary analysis of the same) using the cloud computing capabilities of Amazon Web Services and Microsoft Azure. These techniques can also be employed to successfully compute isomers networks for other chemical compounds."],"dc:format.medium":["born digital","doctoral dissertations"],"dc:identifier":["Thiagarajan_mines_0052E_11303.pdf","T 8330"],"dc:identifier.uri":["https://hdl.handle.net/11124/171241"],"dc:language":["English"],"dc:language.iso":["eng"],"dc:publisher":["Colorado School of Mines. Arthur Lakes Library"],"dc:rights":["Copyright of the original work is retained by the author."],"dc:subject":["cheminformatics","network analytics","big data","symmetry","isomers"],"dc:title":["Faster isomer network generation"],"dc:type":["Text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Doctoral"],"thesis:degree_name":["Doctor of Philosophy (Ph.D.)"],"thesis:institution_name":["Colorado School of Mines"]},"updated_at":"2026-07-24T01:43:08Z"}