{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/88275"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/88275","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Functional discovery in the oxidative D-galacturonate assimilation pathway and development of the enzyme similarity web tool","abstract":"Sequencing technology has improved dramatically over the past few decades. Before the sequencing of complete genomes was possible, the sequencing of a gene was directly linked to the biochemical characterization of its product [1], however biochemical and genetic characterization has not benefited from being scaled up in the same way as has sequencing. Thus, the scientific community is confronted with exponentially growing sequence databases in which roughly half of the entries are either annotated incorrectly or not at all. Therefore, in order to realize the true potential of the data being generated by sequencing projects, something must be done about the way the functions of those sequences are being discovered and identified. One approach to addressing the problem of the growing number of sequences without a known function is that set forth by the Enzyme Function Initiative (EFI). The goal of the EFI is to develop tools and strategies to characterize enzymes discovered in genome projects, and the EFI uses an interdisciplinary approach to address the problem. EFI labs include those with expertise in bioinformatics, computational biology, structural biology, enzymology, and biology, that work together to develop a systematic approach that starts with using bioinformatics to select enzyme candidates for structural elucidation, ligand docking to identify potential substrates, in vitro biochemistry to test those predictions, and microbiology to test for the physiological role of activities identified in vitro. The approach just described is the general approach taken, but other tools and approaches also have been tested and developed in each of the areas mentioned (e.g., bioinformatics, computational biology). Bioinformatics tools that have been further developed include sequence similarity networks (SSNs) and genomic context networks. SSNs have a long history and are useful in visualizing trends across groups of related protein sequences, namely function. Before this work, access to SSNs by experimentalists with little bioinformatics training was limited. To provide the ability for experimentalist to generate an SSN for any protein family (~16,000 now in Pfam), we developed a web tool to generate SSNs quickly and easily. The networks can be viewed in Cytoscape and contain an aggregate of annotation data pulled from different sources (e.g., UniProt, GenomesOnline). The first part of this work (Chapter 2) describes the web tool and provides an example in which members of the enolase superfamily from Agrobacterium tumefaciens strain C58 are mined in a shotgun approach to discover novel enzymatic activities. In the second part of this work, combined bioinformatics and experimental approaches are used to identify two novel enzymes in the oxidative pathway to degrade pectin, the abundant plant cell wall polysaccharide. In the first example (Chapter 3), genomic context and pathway reconstruction combined with in vitro biochemistry and gene expression analysis reveal a novel enzymatic activity of isomerizing the 6-member ring lactone of D-galacturonate (D-galA) to its 5-member ring lactone counterpart. An enzyme to catalyze this reaction had not been identified before this work. In the second example (Chapter 4), in a large scale screening of transporters we were lead to microbial gene neighborhoods containing many enzymes in the known D-galA oxidative pathway but noticed in a number of cases components of the known pathway were missing; in their place candidate enzymes were likely involved in an alternative pathway for metabolizing D-galA. This work lead us to the discovery of an enzyme that hydrolyzed the 6-member ring lactone of D-galA to its acyclic diacid counterpart, meso-galactarate.","abstract_html":"Sequencing technology has improved dramatically over the past few decades. Before the sequencing of complete genomes was possible, the sequencing of a gene was directly linked to the biochemical characterization of its product [1], however biochemical and genetic characterization has not benefited from being scaled up in the same way as has sequencing. Thus, the scientific community is confronted with exponentially growing sequence databases in which roughly half of the entries are either annotated incorrectly or not at all. Therefore, in order to realize the true potential of the data being generated by sequencing projects, something must be done about the way the functions of those sequences are being discovered and identified. One approach to addressing the problem of the growing number of sequences without a known function is that set forth by the Enzyme Function Initiative (EFI). The goal of the EFI is to develop tools and strategies to characterize enzymes discovered in genome projects, and the EFI uses an interdisciplinary approach to address the problem. EFI labs include those with expertise in bioinformatics, computational biology, structural biology, enzymology, and biology, that work together to develop a systematic approach that starts with using bioinformatics to select enzyme candidates for structural elucidation, ligand docking to identify potential substrates, in vitro biochemistry to test those predictions, and microbiology to test for the physiological role of activities identified in vitro. The approach just described is the general approach taken, but other tools and approaches also have been tested and developed in each of the areas mentioned (e.g., bioinformatics, computational biology). Bioinformatics tools that have been further developed include sequence similarity networks (SSNs) and genomic context networks. SSNs have a long history and are useful in visualizing trends across groups of related protein sequences, namely function. Before this work, access to SSNs by experimentalists with little bioinformatics training was limited. To provide the ability for experimentalist to generate an SSN for any protein family (~16,000 now in Pfam), we developed a web tool to generate SSNs quickly and easily. The networks can be viewed in Cytoscape and contain an aggregate of annotation data pulled from different sources (e.g., UniProt, GenomesOnline). The first part of this work (Chapter 2) describes the web tool and provides an example in which members of the enolase superfamily from Agrobacterium tumefaciens strain C58 are mined in a shotgun approach to discover novel enzymatic activities. In the second part of this work, combined bioinformatics and experimental approaches are used to identify two novel enzymes in the oxidative pathway to degrade pectin, the abundant plant cell wall polysaccharide. In the first example (Chapter 3), genomic context and pathway reconstruction combined with in vitro biochemistry and gene expression analysis reveal a novel enzymatic activity of isomerizing the 6-member ring lactone of D-galacturonate (D-galA) to its 5-member ring lactone counterpart. An enzyme to catalyze this reaction had not been identified before this work. In the second example (Chapter 4), in a large scale screening of transporters we were lead to microbial gene neighborhoods containing many enzymes in the known D-galA oxidative pathway but noticed in a number of cases components of the known pathway were missing; in their place candidate enzymes were likely involved in an alternative pathway for metabolizing D-galA. This work lead us to the discovery of an enzyme that hydrolyzed the 6-member ring lactone of D-galA to its acyclic diacid counterpart, meso-galactarate.","abstract_has_math":false,"creators":["Bouvier, Jason T"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Biochemistry","degree_department":null,"school":null,"contributors":["Gerlt, John A.","Cronan, John E.","Nair, Satish K.","Orlean, Peter A."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2015,"date_issued":"2015-09-29T21:03:12Z","date_published":"2015-09-29T21:03:12Z","updated_at":"2026-07-22T22:26:31Z","subjects":["Hexuronate degradation","sequence similarity network","Enzyme Function Initiative"],"languages":["en"],"rights":["Copyright 2015 Jason Bouvier"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/88275","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Gerlt, John A.","Cronan, John E.","Nair, Satish K.","Orlean, Peter A."]},{"key":"dc:creator","label":"Author","values":["Bouvier, Jason T"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2015-09-29T21:03:12Z","2017-09-30T09:15:38Z","2015-08","2015-07-15","2015-8"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Biochemistry"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Hexuronate degradation","sequence similarity network","Enzyme Function Initiative"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2015 Jason Bouvier"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/88275"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Sequencing technology has improved dramatically over the past few decades. Before the sequencing of complete genomes was possible, the sequencing of a gene was directly linked to the biochemical characterization of its product [1], however biochemical and genetic characterization has not benefited from being scaled up in the same way as has sequencing. Thus, the scientific community is confronted with exponentially growing sequence databases in which roughly half of the entries are either annotated incorrectly or not at all. Therefore, in order to realize the true potential of the data being generated by sequencing projects, something must be done about the way the functions of those sequences are being discovered and identified. One approach to addressing the problem of the growing number of sequences without a known function is that set forth by the Enzyme Function Initiative (EFI). The goal of the EFI is to develop tools and strategies to characterize enzymes discovered in genome projects, and the EFI uses an interdisciplinary approach to address the problem. EFI labs include those with expertise in bioinformatics, computational biology, structural biology, enzymology, and biology, that work together to develop a systematic approach that starts with using bioinformatics to select enzyme candidates for structural elucidation, ligand docking to identify potential substrates, in vitro biochemistry to test those predictions, and microbiology to test for the physiological role of activities identified in vitro. The approach just described is the general approach taken, but other tools and approaches also have been tested and developed in each of the areas mentioned (e.g., bioinformatics, computational biology). Bioinformatics tools that have been further developed include sequence similarity networks (SSNs) and genomic context networks. SSNs have a long history and are useful in visualizing trends across groups of related protein sequences, namely function. Before this work, access to SSNs by experimentalists with little bioinformatics training was limited. To provide the ability for experimentalist to generate an SSN for any protein family (~16,000 now in Pfam), we developed a web tool to generate SSNs quickly and easily. The networks can be viewed in Cytoscape and contain an aggregate of annotation data pulled from different sources (e.g., UniProt, GenomesOnline). The first part of this work (Chapter 2) describes the web tool and provides an example in which members of the enolase superfamily from Agrobacterium tumefaciens strain C58 are mined in a shotgun approach to discover novel enzymatic activities. In the second part of this work, combined bioinformatics and experimental approaches are used to identify two novel enzymes in the oxidative pathway to degrade pectin, the abundant plant cell wall polysaccharide. In the first example (Chapter 3), genomic context and pathway reconstruction combined with in vitro biochemistry and gene expression analysis reveal a novel enzymatic activity of isomerizing the 6-member ring lactone of D-galacturonate (D-galA) to its 5-member ring lactone counterpart. An enzyme to catalyze this reaction had not been identified before this work. In the second example (Chapter 4), in a large scale screening of transporters we were lead to microbial gene neighborhoods containing many enzymes in the known D-galA oxidative pathway but noticed in a number of cases components of the known pathway were missing; in their place candidate enzymes were likely involved in an alternative pathway for metabolizing D-galA. This work lead us to the discovery of an enzyme that hydrolyzed the 6-member ring lactone of D-galA to its acyclic diacid counterpart, meso-galactarate.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2017-08-01","The student, Jason Bouvier, accepted the attached license on 2015-07-10 at 14:47.","The student, Jason Bouvier, submitted this Dissertation for approval on 2015-07-10 at 14:49.","This Dissertation was approved for publication on 2015-07-15 at 13:28.","DSpace SAF Submission Ingestion Package generated from Vireo submission #8409 on 2015-09-29 at 15:05:57","Made available in DSpace on 2015-09-29T21:03:12Z (GMT). No. of bitstreams: 2 BOUVIER-DISSERTATION-2015.pdf: 6029594 bytes, checksum: 50aa0f1af0cb8e364c942c67217b70c9 (MD5) LICENSE.txt: 4210 bytes, checksum: 605337e9b2237c2297a48648e54d301a (MD5) Previous issue date: 2015-07-15","Embargo set by: Seth Robbins for item 89555 Lift date: 2017-09-29T21:03:28Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 89555 Lift date: 2017-09-29T21:08:35Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 89555 on 2017-09-30T09:15:38Z."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Functional discovery in the oxidative D-galacturonate assimilation pathway and development of the enzyme similarity web tool"]}]}],"canonical_facts":{"dc:contributor":["Gerlt, John A.","Cronan, John E.","Nair, Satish K.","Orlean, Peter A."],"dc:creator":["Bouvier, Jason T"],"dc:date":["2015-09-29T21:03:12Z","2017-09-30T09:15:38Z","2015-08","2015-07-15","2015-8"],"dc:description":["Sequencing technology has improved dramatically over the past few decades. Before the sequencing of complete genomes was possible, the sequencing of a gene was directly linked to the biochemical characterization of its product [1], however biochemical and genetic characterization has not benefited from being scaled up in the same way as has sequencing. Thus, the scientific community is confronted with exponentially growing sequence databases in which roughly half of the entries are either annotated incorrectly or not at all. Therefore, in order to realize the true potential of the data being generated by sequencing projects, something must be done about the way the functions of those sequences are being discovered and identified. One approach to addressing the problem of the growing number of sequences without a known function is that set forth by the Enzyme Function Initiative (EFI). The goal of the EFI is to develop tools and strategies to characterize enzymes discovered in genome projects, and the EFI uses an interdisciplinary approach to address the problem. EFI labs include those with expertise in bioinformatics, computational biology, structural biology, enzymology, and biology, that work together to develop a systematic approach that starts with using bioinformatics to select enzyme candidates for structural elucidation, ligand docking to identify potential substrates, in vitro biochemistry to test those predictions, and microbiology to test for the physiological role of activities identified in vitro. The approach just described is the general approach taken, but other tools and approaches also have been tested and developed in each of the areas mentioned (e.g., bioinformatics, computational biology). Bioinformatics tools that have been further developed include sequence similarity networks (SSNs) and genomic context networks. SSNs have a long history and are useful in visualizing trends across groups of related protein sequences, namely function. Before this work, access to SSNs by experimentalists with little bioinformatics training was limited. To provide the ability for experimentalist to generate an SSN for any protein family (~16,000 now in Pfam), we developed a web tool to generate SSNs quickly and easily. The networks can be viewed in Cytoscape and contain an aggregate of annotation data pulled from different sources (e.g., UniProt, GenomesOnline). The first part of this work (Chapter 2) describes the web tool and provides an example in which members of the enolase superfamily from Agrobacterium tumefaciens strain C58 are mined in a shotgun approach to discover novel enzymatic activities. In the second part of this work, combined bioinformatics and experimental approaches are used to identify two novel enzymes in the oxidative pathway to degrade pectin, the abundant plant cell wall polysaccharide. In the first example (Chapter 3), genomic context and pathway reconstruction combined with in vitro biochemistry and gene expression analysis reveal a novel enzymatic activity of isomerizing the 6-member ring lactone of D-galacturonate (D-galA) to its 5-member ring lactone counterpart. An enzyme to catalyze this reaction had not been identified before this work. In the second example (Chapter 4), in a large scale screening of transporters we were lead to microbial gene neighborhoods containing many enzymes in the known D-galA oxidative pathway but noticed in a number of cases components of the known pathway were missing; in their place candidate enzymes were likely involved in an alternative pathway for metabolizing D-galA. This work lead us to the discovery of an enzyme that hydrolyzed the 6-member ring lactone of D-galA to its acyclic diacid counterpart, meso-galactarate.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2017-08-01","The student, Jason Bouvier, accepted the attached license on 2015-07-10 at 14:47.","The student, Jason Bouvier, submitted this Dissertation for approval on 2015-07-10 at 14:49.","This Dissertation was approved for publication on 2015-07-15 at 13:28.","DSpace SAF Submission Ingestion Package generated from Vireo submission #8409 on 2015-09-29 at 15:05:57","Made available in DSpace on 2015-09-29T21:03:12Z (GMT). No. of bitstreams: 2 BOUVIER-DISSERTATION-2015.pdf: 6029594 bytes, checksum: 50aa0f1af0cb8e364c942c67217b70c9 (MD5) LICENSE.txt: 4210 bytes, checksum: 605337e9b2237c2297a48648e54d301a (MD5) Previous issue date: 2015-07-15","Embargo set by: Seth Robbins for item 89555 Lift date: 2017-09-29T21:03:28Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 89555 Lift date: 2017-09-29T21:08:35Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 89555 on 2017-09-30T09:15:38Z."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/88275"],"dc:language":["en"],"dc:rights":["Copyright 2015 Jason Bouvier"],"dc:subject":["Hexuronate degradation","sequence similarity network","Enzyme Function Initiative"],"dc:title":["Functional discovery in the oxidative D-galacturonate assimilation pathway and development of the enzyme similarity web tool"],"dc:type":["text"],"thesis:degree_discipline":["Biochemistry"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:26:31Z"}