{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/29776"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/29776","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Inference of degree of significance of single amino acids from the literature","abstract":"Several subfamilies of potassium channels are highly conserved along the vast majority of the protein sequence among a wide array of very distantly related animals. We call this characteristic “hyperconservation”. In this work we create a quantitative definition and explore the degree of hyperconservation and characterization in each of the well-known potassium channel subfamilies. In general the potassium channels seem to exhibit a large degree of hyperconservation within the subfamilies but a wide diversity (to the point of confounding alignment) between subfamilies. Here we examine the literature of one potassium channel subfamily (KCNA2) to determine whether or not all of the completely conserved residues have been noted and considered for functional inference. Out of several thousand papers, we find four residues that are completely conserved but unmentioned in any article; F85, E112, P156, S159. F85 and E112 are in fact completely conserved within and across different K channel subfamilies. The challenges encountered during this search, plus the fact that some completely conserved residues have been overlooked, make it clear that there needs to be a more automated method for extracting sequence-related information from literature articles. The work in this thesis emerged from considering the problem of how to intensively study a protein family based on the sequences for the family. In the first part of the thesis, we consider the issue of studying a family of potassium channels, residue-by-residue. This involves accounting for a history in which residue numbering systems and protein nomenclature are variable throughout the literature on this family. Discovering information in literature about single residues in any protein family can be daunting considering that the residues have a different number placement in each sequence. Then one must consider the change in numbers for each isoform or if an author renumbers them from a sequence section. This problem is greatly compounded when one wishes to consider orthologs and paralogs to these orthologs (homologs) in all species. This involves accounting for a history in which residue numbering systems and protein nomenclature are variable throughout the literature on any family. This has resulted in the creation of a program called FiSHAAL-Finding Single Homologous Amino Acids. It is offered as a prototype literature amino acid location determination program for partial automation of identifying homologous residues and linking any corresponding residues in an alignment column to their PubMed IDs. Ultimate Hypothesis: Can accurate homologous amino acid residue mention information be linked effectively to all PubMed articles in a semi-automated fashion?","abstract_html":"Several subfamilies of potassium channels are highly conserved along the vast majority of the protein sequence among a wide array of very distantly related animals. We call this characteristic “hyperconservation”. In this work we create a quantitative definition and explore the degree of hyperconservation and characterization in each of the well-known potassium channel subfamilies. In general the potassium channels seem to exhibit a large degree of hyperconservation within the subfamilies but a wide diversity (to the point of confounding alignment) between subfamilies. Here we examine the literature of one potassium channel subfamily (KCNA2) to determine whether or not all of the completely conserved residues have been noted and considered for functional inference. Out of several thousand papers, we find four residues that are completely conserved but unmentioned in any article; F85, E112, P156, S159. F85 and E112 are in fact completely conserved within and across different K channel subfamilies. The challenges encountered during this search, plus the fact that some completely conserved residues have been overlooked, make it clear that there needs to be a more automated method for extracting sequence-related information from literature articles. The work in this thesis emerged from considering the problem of how to intensively study a protein family based on the sequences for the family. In the first part of the thesis, we consider the issue of studying a family of potassium channels, residue-by-residue. This involves accounting for a history in which residue numbering systems and protein nomenclature are variable throughout the literature on this family. Discovering information in literature about single residues in any protein family can be daunting considering that the residues have a different number placement in each sequence. Then one must consider the change in numbers for each isoform or if an author renumbers them from a sequence section. This problem is greatly compounded when one wishes to consider orthologs and paralogs to these orthologs (homologs) in all species. This involves accounting for a history in which residue numbering systems and protein nomenclature are variable throughout the literature on any family. This has resulted in the creation of a program called FiSHAAL-Finding Single Homologous Amino Acids. It is offered as a prototype literature amino acid location determination program for partial automation of identifying homologous residues and linking any corresponding residues in an alignment column to their PubMed IDs. Ultimate Hypothesis: Can accurate homologous amino acid residue mention information be linked effectively to all PubMed articles in a semi-automated fashion?","abstract_has_math":false,"creators":["Becker, Anthony"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Molecular & Integrative Physi","degree_department":null,"school":null,"contributors":["Jakobsson, Eric","Nelson, Mark E.","Chung, Hee Jung","Anastasio, Thomas J.","Grosman, Claudio F."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2012,"date_issued":"2012-02-06T20:15:50Z","date_published":"2012-02-06T20:15:50Z","updated_at":"2026-07-22T22:25:29Z","subjects":["homologous amino acid residue","mutated residue search","homologous amino acid residue search","homologous amino acid residue and mutated residue search in journal article literature publications","amino acid","mutated residue","journal article literature","journal literature"],"languages":["en"],"rights":["Copyright 2011 Anthony Becker"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/29776","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Jakobsson, Eric","Nelson, Mark E.","Chung, Hee Jung","Anastasio, Thomas J.","Grosman, Claudio F."]},{"key":"dc:creator","label":"Author","values":["Becker, Anthony"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2012-02-06T20:15:50Z","2011-12"]},{"key":"dc:type","label":"Dc Type","values":["Dissertation / Thesis","text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Molecular & Integrative Physi"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["homologous amino acid residue","mutated residue search","homologous amino acid residue search","homologous amino acid residue and mutated residue search in journal article literature publications","amino acid","mutated residue","journal article literature","journal literature"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2011 Anthony Becker"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/29776"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Several subfamilies of potassium channels are highly conserved along the vast majority of the protein sequence among a wide array of very distantly related animals. We call this characteristic “hyperconservation”. In this work we create a quantitative definition and explore the degree of hyperconservation and characterization in each of the well-known potassium channel subfamilies. In general the potassium channels seem to exhibit a large degree of hyperconservation within the subfamilies but a wide diversity (to the point of confounding alignment) between subfamilies. Here we examine the literature of one potassium channel subfamily (KCNA2) to determine whether or not all of the completely conserved residues have been noted and considered for functional inference. Out of several thousand papers, we find four residues that are completely conserved but unmentioned in any article; F85, E112, P156, S159. F85 and E112 are in fact completely conserved within and across different K channel subfamilies. The challenges encountered during this search, plus the fact that some completely conserved residues have been overlooked, make it clear that there needs to be a more automated method for extracting sequence-related information from literature articles. The work in this thesis emerged from considering the problem of how to intensively study a protein family based on the sequences for the family. In the first part of the thesis, we consider the issue of studying a family of potassium channels, residue-by-residue. This involves accounting for a history in which residue numbering systems and protein nomenclature are variable throughout the literature on this family. Discovering information in literature about single residues in any protein family can be daunting considering that the residues have a different number placement in each sequence. Then one must consider the change in numbers for each isoform or if an author renumbers them from a sequence section. This problem is greatly compounded when one wishes to consider orthologs and paralogs to these orthologs (homologs) in all species. This involves accounting for a history in which residue numbering systems and protein nomenclature are variable throughout the literature on any family. This has resulted in the creation of a program called FiSHAAL-Finding Single Homologous Amino Acids. It is offered as a prototype literature amino acid location determination program for partial automation of identifying homologous residues and linking any corresponding residues in an alignment column to their PubMed IDs. Ultimate Hypothesis: Can accurate homologous amino acid residue mention information be linked effectively to all PubMed articles in a semi-automated fashion?","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2011-11-21T16:01:45Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 10 supplementreferences.doc: 344064 bytes, checksum: 5c5fcdcb9baedd1d531b7969981c04ba (MD5) SFFinderv2.0: 2930 bytes, checksum: 6386f5fe9608345712c7010751f388ec (MD5) Perlprogram.pl: 78123 bytes, checksum: 193cda83a2bf1c115776ba1f0f1d131a (MD5) pdbdistBF85.xls: 1025536 bytes, checksum: 4f6d54d0e3028956c4df2430e106ec20 (MD5) pdbdistBE112.xls: 973824 bytes, checksum: a5f1213a14271232951727495af2778a (MD5) kcna2vertalignment.xls: 2588160 bytes, checksum: 4fcef1c1d38a6afc281aabc868e7629b (MD5) completequerylist.txt: 48977 bytes, checksum: 06a1dad0d622be8a178ff366c6bcb642 (MD5) barplot.py: 1476 bytes, checksum: cea9554f258c8eb05e9f6dbe0731801d (MD5) fishaal.zip: 13795492 bytes, checksum: 8a8e007bb1313ef595ee9308ddd8b7d4 (MD5) Becker_Anthony.pdf: 46254758 bytes, checksum: 2a0c779d4f019eebaeefd4fa23e70bf8 (MD5)","Made available in DSpace on 2012-02-06T20:15:50Z (GMT). No. of bitstreams: 11 fishaal.zip: 13795492 bytes, checksum: 8a8e007bb1313ef595ee9308ddd8b7d4 (MD5) barplot.py: 1476 bytes, checksum: cea9554f258c8eb05e9f6dbe0731801d (MD5) completequerylist.txt: 48977 bytes, checksum: 06a1dad0d622be8a178ff366c6bcb642 (MD5) kcna2vertalignment.xls: 2588160 bytes, checksum: 4fcef1c1d38a6afc281aabc868e7629b (MD5) pdbdistBE112.xls: 973824 bytes, checksum: a5f1213a14271232951727495af2778a (MD5) pdbdistBF85.xls: 1025536 bytes, checksum: 4f6d54d0e3028956c4df2430e106ec20 (MD5) Perlprogram.pl: 78123 bytes, checksum: 193cda83a2bf1c115776ba1f0f1d131a (MD5) SFFinderv2.0: 2930 bytes, checksum: 6386f5fe9608345712c7010751f388ec (MD5) supplementreferences.doc: 344064 bytes, checksum: 5c5fcdcb9baedd1d531b7969981c04ba (MD5) Becker_Anthony.pdf: 46262748 bytes, checksum: 6221fd60350757f9dd75f1471dafc0b1 (MD5) license.txt: 4064 bytes, checksum: 00e0cc0e53b168e7593078e7c24d53c6 (MD5)"]},{"key":"dc:title","label":"Title","values":["Inference of degree of significance of single amino acids from the literature"]}]}],"canonical_facts":{"dc:contributor":["Jakobsson, Eric","Nelson, Mark E.","Chung, Hee Jung","Anastasio, Thomas J.","Grosman, Claudio F."],"dc:creator":["Becker, Anthony"],"dc:date":["2012-02-06T20:15:50Z","2011-12"],"dc:description":["Several subfamilies of potassium channels are highly conserved along the vast majority of the protein sequence among a wide array of very distantly related animals. We call this characteristic “hyperconservation”. In this work we create a quantitative definition and explore the degree of hyperconservation and characterization in each of the well-known potassium channel subfamilies. In general the potassium channels seem to exhibit a large degree of hyperconservation within the subfamilies but a wide diversity (to the point of confounding alignment) between subfamilies. Here we examine the literature of one potassium channel subfamily (KCNA2) to determine whether or not all of the completely conserved residues have been noted and considered for functional inference. Out of several thousand papers, we find four residues that are completely conserved but unmentioned in any article; F85, E112, P156, S159. F85 and E112 are in fact completely conserved within and across different K channel subfamilies. The challenges encountered during this search, plus the fact that some completely conserved residues have been overlooked, make it clear that there needs to be a more automated method for extracting sequence-related information from literature articles. The work in this thesis emerged from considering the problem of how to intensively study a protein family based on the sequences for the family. In the first part of the thesis, we consider the issue of studying a family of potassium channels, residue-by-residue. This involves accounting for a history in which residue numbering systems and protein nomenclature are variable throughout the literature on this family. Discovering information in literature about single residues in any protein family can be daunting considering that the residues have a different number placement in each sequence. Then one must consider the change in numbers for each isoform or if an author renumbers them from a sequence section. This problem is greatly compounded when one wishes to consider orthologs and paralogs to these orthologs (homologs) in all species. This involves accounting for a history in which residue numbering systems and protein nomenclature are variable throughout the literature on any family. This has resulted in the creation of a program called FiSHAAL-Finding Single Homologous Amino Acids. It is offered as a prototype literature amino acid location determination program for partial automation of identifying homologous residues and linking any corresponding residues in an alignment column to their PubMed IDs. Ultimate Hypothesis: Can accurate homologous amino acid residue mention information be linked effectively to all PubMed articles in a semi-automated fashion?","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2011-11-21T16:01:45Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 10 supplementreferences.doc: 344064 bytes, checksum: 5c5fcdcb9baedd1d531b7969981c04ba (MD5) SFFinderv2.0: 2930 bytes, checksum: 6386f5fe9608345712c7010751f388ec (MD5) Perlprogram.pl: 78123 bytes, checksum: 193cda83a2bf1c115776ba1f0f1d131a (MD5) pdbdistBF85.xls: 1025536 bytes, checksum: 4f6d54d0e3028956c4df2430e106ec20 (MD5) pdbdistBE112.xls: 973824 bytes, checksum: a5f1213a14271232951727495af2778a (MD5) kcna2vertalignment.xls: 2588160 bytes, checksum: 4fcef1c1d38a6afc281aabc868e7629b (MD5) completequerylist.txt: 48977 bytes, checksum: 06a1dad0d622be8a178ff366c6bcb642 (MD5) barplot.py: 1476 bytes, checksum: cea9554f258c8eb05e9f6dbe0731801d (MD5) fishaal.zip: 13795492 bytes, checksum: 8a8e007bb1313ef595ee9308ddd8b7d4 (MD5) Becker_Anthony.pdf: 46254758 bytes, checksum: 2a0c779d4f019eebaeefd4fa23e70bf8 (MD5)","Made available in DSpace on 2012-02-06T20:15:50Z (GMT). No. of bitstreams: 11 fishaal.zip: 13795492 bytes, checksum: 8a8e007bb1313ef595ee9308ddd8b7d4 (MD5) barplot.py: 1476 bytes, checksum: cea9554f258c8eb05e9f6dbe0731801d (MD5) completequerylist.txt: 48977 bytes, checksum: 06a1dad0d622be8a178ff366c6bcb642 (MD5) kcna2vertalignment.xls: 2588160 bytes, checksum: 4fcef1c1d38a6afc281aabc868e7629b (MD5) pdbdistBE112.xls: 973824 bytes, checksum: a5f1213a14271232951727495af2778a (MD5) pdbdistBF85.xls: 1025536 bytes, checksum: 4f6d54d0e3028956c4df2430e106ec20 (MD5) Perlprogram.pl: 78123 bytes, checksum: 193cda83a2bf1c115776ba1f0f1d131a (MD5) SFFinderv2.0: 2930 bytes, checksum: 6386f5fe9608345712c7010751f388ec (MD5) supplementreferences.doc: 344064 bytes, checksum: 5c5fcdcb9baedd1d531b7969981c04ba (MD5) Becker_Anthony.pdf: 46262748 bytes, checksum: 6221fd60350757f9dd75f1471dafc0b1 (MD5) license.txt: 4064 bytes, checksum: 00e0cc0e53b168e7593078e7c24d53c6 (MD5)"],"dc:identifier":["http://hdl.handle.net/2142/29776"],"dc:language":["en"],"dc:rights":["Copyright 2011 Anthony Becker"],"dc:subject":["homologous amino acid residue","mutated residue search","homologous amino acid residue search","homologous amino acid residue and mutated residue search in journal article literature publications","amino acid","mutated residue","journal article literature","journal literature"],"dc:title":["Inference of degree of significance of single amino acids from the literature"],"dc:type":["Dissertation / Thesis","text"],"thesis:degree_discipline":["Molecular & Integrative Physi"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:29Z"}