{"id":{"repo_id":"unsw","oai_identifier":"oai:unsworks.library.unsw.edu.au:1959.4/57471"},"canonical_url":"https://search.dev.ndltd.org/etd/unsw/oai:unsworks.library.unsw.edu.au:1959.4/57471","repository":{"repo_id":"unsw","name":"University of New South Wales","base_url":"https://unsworks.unsw.edu.au/oai/provider"},"display":{"title":"Novel pharmacophore clustering methods for protein binding site comparison","abstract":"Proteins perform diverse functions within cells. Some of the functions depend on the protein being involved in a protein complex, interacting with other proteins or with other entities (ligands) through specific binding sites on their surface. Comparison of protein binding sites has potential benefits in many research fields, including drug promiscuity studies, polypharmacology and immunology. While multiple methods have been proposed for comparing binding sites, they tend to focus on comparing very similar proteins and have only been developed for small specific datasets or very targeted applications. None of these methods make use of the powerful representation afforded by 3D complex-based pharmacophores. A pharmacophore model provides a description of a binding site, consisting of a group of chemical features arranged in three-dimensional space, that can be used to represent biological activities. Two different pharmacophore comparison and clustering methods based on the Iterative Closest Point (ICP) algorithm are proposed: a 3-dimensional ICP pharmacophore clustering method, and an N-dimensional ICP pharmacophore clustering method. These methods are complemented by a series of data pre-processing methods for input data preparation. The implementation of the methods takes computational representations (pharmacophores) of single molecule or protein complexes as input and produces distance matrices that can be visualised as dendrograms. The methods integrate both alignment-dependent and alignment-independent concepts. Both clustering methods were successfully evaluated using a 31 globulin-binding steroid dataset and a 41 antibody-antigen dataset, and were able to handle a larger dataset of 159 protein homodimers. For the steroid dataset, the resulting classification of ligands shows good correspondence with a classification based on binding affinity. For the antibody-antigen dataset, the classification of antigens reflected both antigen type and binding antibody. The applications to homodimers demonstrated the ability of both clustering methods to handle a larger dataset, and the possibility to visualise N-D pairwise comparisons using structural superposition of binding sites.","abstract_html":"Proteins perform diverse functions within cells. Some of the functions depend on the protein being involved in a protein complex, interacting with other proteins or with other entities (ligands) through specific binding sites on their surface. Comparison of protein binding sites has potential benefits in many research fields, including drug promiscuity studies, polypharmacology and immunology. While multiple methods have been proposed for comparing binding sites, they tend to focus on comparing very similar proteins and have only been developed for small specific datasets or very targeted applications. None of these methods make use of the powerful representation afforded by 3D complex-based pharmacophores. A pharmacophore model provides a description of a binding site, consisting of a group of chemical features arranged in three-dimensional space, that can be used to represent biological activities. Two different pharmacophore comparison and clustering methods based on the Iterative Closest Point (ICP) algorithm are proposed: a 3-dimensional ICP pharmacophore clustering method, and an N-dimensional ICP pharmacophore clustering method. These methods are complemented by a series of data pre-processing methods for input data preparation. The implementation of the methods takes computational representations (pharmacophores) of single molecule or protein complexes as input and produces distance matrices that can be visualised as dendrograms. The methods integrate both alignment-dependent and alignment-independent concepts. Both clustering methods were successfully evaluated using a 31 globulin-binding steroid dataset and a 41 antibody-antigen dataset, and were able to handle a larger dataset of 159 protein homodimers. For the steroid dataset, the resulting classification of ligands shows good correspondence with a classification based on binding affinity. For the antibody-antigen dataset, the classification of antigens reflected both antigen type and binding antibody. The applications to homodimers demonstrated the ability of both clustering methods to handle a larger dataset, and the possibility to visualise N-D pairwise comparisons using structural superposition of binding sites.","abstract_has_math":false,"creators":["Zhou, Lingxiao"],"institution":"UNSW, Sydney","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2017,"date_issued":"2017","date_published":"2017","updated_at":"2026-07-24T05:34:44Z","subjects":["Bioinformatics","ICP","Pharmacophores","Cheminformatics","Machine learning"],"languages":["EN"],"rights":["open access","CC BY-NC-ND 3.0","free_to_read"],"rights_urls":["https://purl.org/coar/access_right/c_abf2","https://creativecommons.org/licenses/by-nc-nd/3.0/au/"],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["https://doi.org/10.26190/unsworks/19478"],"render_values":[{"text":"https://doi.org/10.26190/unsworks/19478","href":"https://doi.org/10.26190/unsworks/19478","code":true}]}]},"links":{"outbound_url":"http://hdl.handle.net/1959.4/57471","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["Zhou, Lingxiao"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2017"]},{"key":"dc:publisher","label":"Institution","values":["UNSW, Sydney"]},{"key":"dc:type","label":"Dc Type","values":["doctoral thesis","http://purl.org/coar/resource_type/c_db06"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Bioinformatics","ICP","Pharmacophores","Cheminformatics","Machine learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["EN"]},{"key":"dc:rights","label":"Dc Rights","values":["open access","https://purl.org/coar/access_right/c_abf2","CC BY-NC-ND 3.0","https://creativecommons.org/licenses/by-nc-nd/3.0/au/","free_to_read"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/1959.4/57471","https://unsworks.unsw.edu.au/bitstreams/6ef0082e-11fb-4ba1-8f3e-93ec04a9bf2d/download","https://doi.org/10.26190/unsworks/19478"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Proteins perform diverse functions within cells. Some of the functions depend on the protein being involved in a protein complex, interacting with other proteins or with other entities (ligands) through specific binding sites on their surface. Comparison of protein binding sites has potential benefits in many research fields, including drug promiscuity studies, polypharmacology and immunology. While multiple methods have been proposed for comparing binding sites, they tend to focus on comparing very similar proteins and have only been developed for small specific datasets or very targeted applications. None of these methods make use of the powerful representation afforded by 3D complex-based pharmacophores. A pharmacophore model provides a description of a binding site, consisting of a group of chemical features arranged in three-dimensional space, that can be used to represent biological activities. Two different pharmacophore comparison and clustering methods based on the Iterative Closest Point (ICP) algorithm are proposed: a 3-dimensional ICP pharmacophore clustering method, and an N-dimensional ICP pharmacophore clustering method. These methods are complemented by a series of data pre-processing methods for input data preparation. The implementation of the methods takes computational representations (pharmacophores) of single molecule or protein complexes as input and produces distance matrices that can be visualised as dendrograms. The methods integrate both alignment-dependent and alignment-independent concepts. Both clustering methods were successfully evaluated using a 31 globulin-binding steroid dataset and a 41 antibody-antigen dataset, and were able to handle a larger dataset of 159 protein homodimers. For the steroid dataset, the resulting classification of ligands shows good correspondence with a classification based on binding affinity. For the antibody-antigen dataset, the classification of antigens reflected both antigen type and binding antibody. The applications to homodimers demonstrated the ability of both clustering methods to handle a larger dataset, and the possibility to visualise N-D pairwise comparisons using structural superposition of binding sites."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Novel pharmacophore clustering methods for protein binding site comparison"]}]}],"canonical_facts":{"dc:creator":["Zhou, Lingxiao"],"dc:date":["2017"],"dc:description":["Proteins perform diverse functions within cells. Some of the functions depend on the protein being involved in a protein complex, interacting with other proteins or with other entities (ligands) through specific binding sites on their surface. Comparison of protein binding sites has potential benefits in many research fields, including drug promiscuity studies, polypharmacology and immunology. While multiple methods have been proposed for comparing binding sites, they tend to focus on comparing very similar proteins and have only been developed for small specific datasets or very targeted applications. None of these methods make use of the powerful representation afforded by 3D complex-based pharmacophores. A pharmacophore model provides a description of a binding site, consisting of a group of chemical features arranged in three-dimensional space, that can be used to represent biological activities. Two different pharmacophore comparison and clustering methods based on the Iterative Closest Point (ICP) algorithm are proposed: a 3-dimensional ICP pharmacophore clustering method, and an N-dimensional ICP pharmacophore clustering method. These methods are complemented by a series of data pre-processing methods for input data preparation. The implementation of the methods takes computational representations (pharmacophores) of single molecule or protein complexes as input and produces distance matrices that can be visualised as dendrograms. The methods integrate both alignment-dependent and alignment-independent concepts. Both clustering methods were successfully evaluated using a 31 globulin-binding steroid dataset and a 41 antibody-antigen dataset, and were able to handle a larger dataset of 159 protein homodimers. For the steroid dataset, the resulting classification of ligands shows good correspondence with a classification based on binding affinity. For the antibody-antigen dataset, the classification of antigens reflected both antigen type and binding antibody. The applications to homodimers demonstrated the ability of both clustering methods to handle a larger dataset, and the possibility to visualise N-D pairwise comparisons using structural superposition of binding sites."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/1959.4/57471","https://unsworks.unsw.edu.au/bitstreams/6ef0082e-11fb-4ba1-8f3e-93ec04a9bf2d/download","https://doi.org/10.26190/unsworks/19478"],"dc:language":["EN"],"dc:publisher":["UNSW, Sydney"],"dc:rights":["open access","https://purl.org/coar/access_right/c_abf2","CC BY-NC-ND 3.0","https://creativecommons.org/licenses/by-nc-nd/3.0/au/","free_to_read"],"dc:subject":["Bioinformatics","ICP","Pharmacophores","Cheminformatics","Machine learning"],"dc:title":["Novel pharmacophore clustering methods for protein binding site comparison"],"dc:type":["doctoral thesis","http://purl.org/coar/resource_type/c_db06"]},"updated_at":"2026-07-24T05:34:44Z"}