{"id":{"repo_id":"unr","oai_identifier":"oai:scholarwolf.unr.edu:11714/589"},"canonical_url":"https://search.dev.ndltd.org/etd/unr/oai:scholarwolf.unr.edu:11714/589","repository":{"repo_id":"unr","name":"University of Nevada - Reno","base_url":"https://scholarwolf.unr.edu/server/oai/request"},"display":{"title":"Use of Short Amino Acid Motifs in the Computational Analysis of Protein Diversity and Function","abstract":"The explosion of whole genome sequence and environmental sequence data afford us the opportunity to explore protein diversity and protein function. This is particularly exciting given the nascent field of synthetic biology. A comprehensive computational analysis of extant proteins is needed in order to define the limitations on protein structure and diversity from a bioengineering perspective. This paper focuses on defining an upper limit for protein diversity using computational approaches derived from linguistic analyses. These methods are used to make a prediction on the upper limit of unique proteins and number of highly conserved motifs. Motifs deemed highly conserved will, more than likely represent important structural components of basic proteins. Results were gathered from two large data sets: all of the currently available microbial genome sequences available from NCBI and the Global Ocean Survey data set. There were 6.6 million unique proteins at 95% amino acid identity. The majority of unique motifs in these data sets were only found once. The motifs deemed highly conserved in lifestyle groupings of organisms and individual organisms were analyzed for function based on a conserved domain search. The importance between pathogenicity and cell motility and secretion related genes and proteins was observed. These motifs represent potential new drug targets or areas of future experimentation.","abstract_html":"The explosion of whole genome sequence and environmental sequence data afford us the opportunity to explore protein diversity and protein function. This is particularly exciting given the nascent field of synthetic biology. A comprehensive computational analysis of extant proteins is needed in order to define the limitations on protein structure and diversity from a bioengineering perspective. This paper focuses on defining an upper limit for protein diversity using computational approaches derived from linguistic analyses. These methods are used to make a prediction on the upper limit of unique proteins and number of highly conserved motifs. Motifs deemed highly conserved will, more than likely represent important structural components of basic proteins. Results were gathered from two large data sets: all of the currently available microbial genome sequences available from NCBI and the Global Ocean Survey data set. There were 6.6 million unique proteins at 95% amino acid identity. The majority of unique motifs in these data sets were only found once. The motifs deemed highly conserved in lifestyle groupings of organisms and individual organisms were analyzed for function based on a conserved domain search. The importance between pathogenicity and cell motility and secretion related genes and proteins was observed. These motifs represent potential new drug targets or areas of future experimentation.","abstract_has_math":false,"creators":["Dussaq, Alex M."],"institution":"University of Nevada, Reno","degree_name":"Journalism","degree_level":"Honors Thesis","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Grzymski, Joseph J."],"committee_chairs":[],"committee_members":[],"year":2010,"date_issued":"2010","date_published":"2010","updated_at":"2026-07-27T21:46:41Z","subjects":[],"languages":["en_US","English"],"rights":["In Copyright(All Rights Reserved)"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/11714/589","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Grzymski, Joseph J."]},{"key":"dc:creator","label":"Author","values":["Dussaq, Alex M."]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2017-01-24T23:09:14Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2017-01-24T23:09:14Z"]},{"key":"dc:date.issued","label":"Date","values":["2010"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Honors Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Journalism"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Nevada, Reno"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["English"]},{"key":"dc:language.iso","label":"Language (ISO)","values":["en_US"]},{"key":"dc:rights","label":"Dc Rights","values":["In Copyright(All Rights Reserved)"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["http://hdl.handle.net/11714/589"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The University of Nevada, Reno Libraries will promptly respond to removal requests related to content that violates intellectual property laws, data protections, or has been uploaded without creator consent. Takedown notices should be directed to our ScholarWolf team (scholarwolf@library.unr.edu) with information about the object, including its full URL and the nature of your complaint."]},{"key":"dc:description.abstract","label":"Abstract","values":["The explosion of whole genome sequence and environmental sequence data afford us the opportunity to explore protein diversity and protein function. This is particularly exciting given the nascent field of synthetic biology. A comprehensive computational analysis of extant proteins is needed in order to define the limitations on protein structure and diversity from a bioengineering perspective. This paper focuses on defining an upper limit for protein diversity using computational approaches derived from linguistic analyses. These methods are used to make a prediction on the upper limit of unique proteins and number of highly conserved motifs. Motifs deemed highly conserved will, more than likely represent important structural components of basic proteins. Results were gathered from two large data sets: all of the currently available microbial genome sequences available from NCBI and the Global Ocean Survey data set. There were 6.6 million unique proteins at 95% amino acid identity. The majority of unique motifs in these data sets were only found once. The motifs deemed highly conserved in lifestyle groupings of organisms and individual organisms were analyzed for function based on a conserved domain search. The importance between pathogenicity and cell motility and secretion related genes and proteins was observed. These motifs represent potential new drug targets or areas of future experimentation."]},{"key":"dc:format","label":"Dc Format","values":["PDF"]},{"key":"dc:title","label":"Title","values":["Use of Short Amino Acid Motifs in the Computational Analysis of Protein Diversity and Function"]}]}],"canonical_facts":{"dc:contributor.advisor":["Grzymski, Joseph J."],"dc:creator":["Dussaq, Alex M."],"dc:date.accessioned":["2017-01-24T23:09:14Z"],"dc:date.available":["2017-01-24T23:09:14Z"],"dc:date.issued":["2010"],"dc:description":["The University of Nevada, Reno Libraries will promptly respond to removal requests related to content that violates intellectual property laws, data protections, or has been uploaded without creator consent. Takedown notices should be directed to our ScholarWolf team (scholarwolf@library.unr.edu) with information about the object, including its full URL and the nature of your complaint."],"dc:description.abstract":["The explosion of whole genome sequence and environmental sequence data afford us the opportunity to explore protein diversity and protein function. This is particularly exciting given the nascent field of synthetic biology. A comprehensive computational analysis of extant proteins is needed in order to define the limitations on protein structure and diversity from a bioengineering perspective. This paper focuses on defining an upper limit for protein diversity using computational approaches derived from linguistic analyses. These methods are used to make a prediction on the upper limit of unique proteins and number of highly conserved motifs. Motifs deemed highly conserved will, more than likely represent important structural components of basic proteins. Results were gathered from two large data sets: all of the currently available microbial genome sequences available from NCBI and the Global Ocean Survey data set. There were 6.6 million unique proteins at 95% amino acid identity. The majority of unique motifs in these data sets were only found once. The motifs deemed highly conserved in lifestyle groupings of organisms and individual organisms were analyzed for function based on a conserved domain search. The importance between pathogenicity and cell motility and secretion related genes and proteins was observed. These motifs represent potential new drug targets or areas of future experimentation."],"dc:format":["PDF"],"dc:identifier.uri":["http://hdl.handle.net/11714/589"],"dc:language":["English"],"dc:language.iso":["en_US"],"dc:rights":["In Copyright(All Rights Reserved)"],"dc:title":["Use of Short Amino Acid Motifs in the Computational Analysis of Protein Diversity and Function"],"dc:type":["Thesis"],"thesis:degree_level":["Honors Thesis"],"thesis:degree_name":["Journalism"],"thesis:institution_name":["University of Nevada, Reno"]},"updated_at":"2026-07-27T21:46:41Z"}