{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/104910"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/104910","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Efficient algorithm for selecting protein residue-residue contacts","abstract":"\"The functions of proteins are largely determined by their structures. Determination of the protein three-dimensional structure is experimentally and computationally challenging. Since amino acids residues that are spatially close often co-evolve, the correlation allows us to predict the contacts from multiple sequence alignments. The predicted contacts can then be used as spatial constraints and offer guidance in protein structure prediction. The constraints can be used as inputs to a protein structure prediction algorithm to produce \"\"decoy\"\" models as tentative 3D structures for proteins. However, the computation power required for structural prediction grows exponentially with respect to the number of contacts selected. Thus selecting few and yet informative contacts are essential for producing high-quality models quickly. Existing contact prediction methods aim for improving precision and recall. However, not all contacts offer the same level of structural information in terms of structure prediction. Therefore, the strategy to select contacts of highest confidence may not be ideal for structure prediction. Here we present an efficient algorithm, ContactSel, to select contacts for assisting contact-guided ab inito folding. We take the key idea that contacts that involve residues far apart (long-ranged) and collections of contacts that are most diverse contains more information than contacts that are shorter ranged and closed by. We formulate the contact selection problem into an integer programming algorithm to select structurally diverse contacts. For evaluation, we generated decoy models using L/2 contacts selected by ContactSel and a naive selection baseline. We show that we achieved significant improvement on the CASP 12 domain set.\"","abstract_html":"&quot;The functions of proteins are largely determined by their structures. Determination of the protein three-dimensional structure is experimentally and computationally challenging. Since amino acids residues that are spatially close often co-evolve, the correlation allows us to predict the contacts from multiple sequence alignments. The predicted contacts can then be used as spatial constraints and offer guidance in protein structure prediction. The constraints can be used as inputs to a protein structure prediction algorithm to produce &quot;&quot;decoy&quot;&quot; models as tentative 3D structures for proteins. However, the computation power required for structural prediction grows exponentially with respect to the number of contacts selected. Thus selecting few and yet informative contacts are essential for producing high-quality models quickly. Existing contact prediction methods aim for improving precision and recall. However, not all contacts offer the same level of structural information in terms of structure prediction. Therefore, the strategy to select contacts of highest confidence may not be ideal for structure prediction. Here we present an efficient algorithm, ContactSel, to select contacts for assisting contact-guided ab inito folding. We take the key idea that contacts that involve residues far apart (long-ranged) and collections of contacts that are most diverse contains more information than contacts that are shorter ranged and closed by. We formulate the contact selection problem into an integer programming algorithm to select structurally diverse contacts. For evaluation, we generated decoy models using L/2 contacts selected by ContactSel and a naive selection baseline. We show that we achieved significant improvement on the CASP 12 domain set.&quot;","abstract_has_math":false,"creators":["Ye, Qing"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Peng, Jian"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019-08-23T20:01:15Z","date_published":"2019-08-23T20:01:15Z","updated_at":"2026-07-22T22:24:42Z","subjects":["Protein Contacts","Protein Structure","Contact Selection","Protein Folding","Integer Programming"],"languages":["en"],"rights":["Copyright 2019 Qing Ye"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/104910","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Peng, Jian"]},{"key":"dc:creator","label":"Author","values":["Ye, Qing"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2019-08-23T20:01:15Z","2019-04-26","2019-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Protein Contacts","Protein Structure","Contact Selection","Protein Folding","Integer Programming"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2019 Qing Ye"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/104910"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["\"The functions of proteins are largely determined by their structures. Determination of the protein three-dimensional structure is experimentally and computationally challenging. Since amino acids residues that are spatially close often co-evolve, the correlation allows us to predict the contacts from multiple sequence alignments. The predicted contacts can then be used as spatial constraints and offer guidance in protein structure prediction. The constraints can be used as inputs to a protein structure prediction algorithm to produce \"\"decoy\"\" models as tentative 3D structures for proteins. However, the computation power required for structural prediction grows exponentially with respect to the number of contacts selected. Thus selecting few and yet informative contacts are essential for producing high-quality models quickly. Existing contact prediction methods aim for improving precision and recall. However, not all contacts offer the same level of structural information in terms of structure prediction. Therefore, the strategy to select contacts of highest confidence may not be ideal for structure prediction. Here we present an efficient algorithm, ContactSel, to select contacts for assisting contact-guided ab inito folding. We take the key idea that contacts that involve residues far apart (long-ranged) and collections of contacts that are most diverse contains more information than contacts that are shorter ranged and closed by. We formulate the contact selection problem into an integer programming algorithm to select structurally diverse contacts. For evaluation, we generated decoy models using L/2 contacts selected by ContactSel and a naive selection baseline. We show that we achieved significant improvement on the CASP 12 domain set.\"","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2019-08-22 without embargo terms","The student, Qing Ye, accepted the attached license on 2019-04-25 at 15:43.","The student, Qing Ye, submitted this Thesis for approval on 2019-04-25 at 16:04.","This Thesis was approved for publication on 2019-04-26 at 10:28.","DSpace SAF Submission Ingestion Package generated from Vireo submission #13846 on 2019-08-22 at 14:46:14","Made available in DSpace on 2019-08-23T20:01:15Z (GMT). No. of bitstreams: 2 YE-THESIS-2019.pdf: 3966731 bytes, checksum: 53dbf56d9edb1ab6e676dab76cc77566 (MD5) LICENSE.txt: 4204 bytes, checksum: f74b7902a91280048e04caaaf17c1c7d (MD5) Previous issue date: 2019-04-26"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Efficient algorithm for selecting protein residue-residue contacts"]}]}],"canonical_facts":{"dc:contributor":["Peng, Jian"],"dc:creator":["Ye, Qing"],"dc:date":["2019-08-23T20:01:15Z","2019-04-26","2019-05"],"dc:description":["\"The functions of proteins are largely determined by their structures. Determination of the protein three-dimensional structure is experimentally and computationally challenging. Since amino acids residues that are spatially close often co-evolve, the correlation allows us to predict the contacts from multiple sequence alignments. The predicted contacts can then be used as spatial constraints and offer guidance in protein structure prediction. The constraints can be used as inputs to a protein structure prediction algorithm to produce \"\"decoy\"\" models as tentative 3D structures for proteins. However, the computation power required for structural prediction grows exponentially with respect to the number of contacts selected. Thus selecting few and yet informative contacts are essential for producing high-quality models quickly. Existing contact prediction methods aim for improving precision and recall. However, not all contacts offer the same level of structural information in terms of structure prediction. Therefore, the strategy to select contacts of highest confidence may not be ideal for structure prediction. Here we present an efficient algorithm, ContactSel, to select contacts for assisting contact-guided ab inito folding. We take the key idea that contacts that involve residues far apart (long-ranged) and collections of contacts that are most diverse contains more information than contacts that are shorter ranged and closed by. We formulate the contact selection problem into an integer programming algorithm to select structurally diverse contacts. For evaluation, we generated decoy models using L/2 contacts selected by ContactSel and a naive selection baseline. We show that we achieved significant improvement on the CASP 12 domain set.\"","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2019-08-22 without embargo terms","The student, Qing Ye, accepted the attached license on 2019-04-25 at 15:43.","The student, Qing Ye, submitted this Thesis for approval on 2019-04-25 at 16:04.","This Thesis was approved for publication on 2019-04-26 at 10:28.","DSpace SAF Submission Ingestion Package generated from Vireo submission #13846 on 2019-08-22 at 14:46:14","Made available in DSpace on 2019-08-23T20:01:15Z (GMT). No. of bitstreams: 2 YE-THESIS-2019.pdf: 3966731 bytes, checksum: 53dbf56d9edb1ab6e676dab76cc77566 (MD5) LICENSE.txt: 4204 bytes, checksum: f74b7902a91280048e04caaaf17c1c7d (MD5) Previous issue date: 2019-04-26"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/104910"],"dc:language":["en"],"dc:rights":["Copyright 2019 Qing Ye"],"dc:subject":["Protein Contacts","Protein Structure","Contact Selection","Protein Folding","Integer Programming"],"dc:title":["Efficient algorithm for selecting protein residue-residue contacts"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:42Z"}