{"id":{"repo_id":"cambridge","oai_identifier":"oai:www.repository.cam.ac.uk:1810/386696"},"canonical_url":"https://search.dev.ndltd.org/etd/cambridge/oai:www.repository.cam.ac.uk:1810/386696","repository":{"repo_id":"cambridge","name":"Cambridge University","base_url":"https://api.repository.cam.ac.uk/server/oai/request"},"display":{"title":"Enabling the discovery of novel bioactives and enzymes from billions of protein sequences","abstract":"Microorganisms are the most abundant life form on earth and are found in every naturally occurring ecosystem. Most of these microorganisms have yet to be cultured. Metagenomics- based methods allow the culture-independent analysis of any biome through sequencing the whole DNA content of an environment sample and analysing it with computational methods. MGnify is one of the most widely used platforms for the analysis of metagenomic sequences. Since 2018, MGnify has assembled and analysed over 57,000 shotgun metagenomics datasets. From these assemblies, over 2.4 billion non-redundant protein sequences have been identified. This protein database outscales other protein repositories and contains a significant fraction of functionally unexplored sequences, making it a treasure trove for protein mining. However, due to the scale of the resource, even simply downloading the raw sequence data can be challenging. Subsequently, search results can contain hundreds of thousands of matches, making exploring the results difficult. With this thesis, I outline the approaches I undertook to expand the data in the MGnify protein database to facilitate filtering and exploration of the data. To enable this to be accessible to all, new technical solutions were explored to allow rapid querying of the data. Finally, to connect search results to the underlying database, I developed an interactive platform to facilitate intuitive data mining using a comprehensive toolbox. This platform enables the rapid identification of potentially novel bioactives and enzymes within MGnify. I demonstrate the platform’s utility through a series of use cases focused on enzymes capable of degrading plastics. Based on this experience and additional use cases looking at CRIPSR- Cas systems, I extended the platform to enable the integration of multiple search results. This extension facilitates both comparative analysis and genomic context. These additional capabilities make it a valuable tool for discovering gene clusters and understanding the distribution of proteins across different environments. Overall, this thesis presents the significant expansion of one of the largest protein databases available and the development of a versatile platform for efficient protein mining. The ability to rapidly mine vast metagenomic datasets paves the way for discovering novel enzymes and bioactive compounds that can be applied in industrial processes, bioremediation or biomedical contexts.","abstract_html":"Microorganisms are the most abundant life form on earth and are found in every naturally occurring ecosystem. Most of these microorganisms have yet to be cultured. Metagenomics- based methods allow the culture-independent analysis of any biome through sequencing the whole DNA content of an environment sample and analysing it with computational methods. MGnify is one of the most widely used platforms for the analysis of metagenomic sequences. Since 2018, MGnify has assembled and analysed over 57,000 shotgun metagenomics datasets. From these assemblies, over 2.4 billion non-redundant protein sequences have been identified. This protein database outscales other protein repositories and contains a significant fraction of functionally unexplored sequences, making it a treasure trove for protein mining. However, due to the scale of the resource, even simply downloading the raw sequence data can be challenging. Subsequently, search results can contain hundreds of thousands of matches, making exploring the results difficult. With this thesis, I outline the approaches I undertook to expand the data in the MGnify protein database to facilitate filtering and exploration of the data. To enable this to be accessible to all, new technical solutions were explored to allow rapid querying of the data. Finally, to connect search results to the underlying database, I developed an interactive platform to facilitate intuitive data mining using a comprehensive toolbox. This platform enables the rapid identification of potentially novel bioactives and enzymes within MGnify. I demonstrate the platform’s utility through a series of use cases focused on enzymes capable of degrading plastics. Based on this experience and additional use cases looking at CRIPSR- Cas systems, I extended the platform to enable the integration of multiple search results. This extension facilitates both comparative analysis and genomic context. These additional capabilities make it a valuable tool for discovering gene clusters and understanding the distribution of proteins across different environments. Overall, this thesis presents the significant expansion of one of the largest protein databases available and the development of a versatile platform for efficient protein mining. The ability to rapidly mine vast metagenomic datasets paves the way for discovering novel enzymes and bioactive compounds that can be applied in industrial processes, bioremediation or biomedical contexts.","abstract_has_math":false,"creators":["Langer, Felix"],"institution":"University of Cambridge","degree_name":"Doctor of Philosophy (PhD)","degree_level":"Doctoral","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Finn, Robert"],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-07-31","date_published":"2024-07-31","updated_at":"2026-07-22T22:24:32Z","subjects":["bioinformatics","computational biology","protein mining"],"languages":[],"rights":[],"rights_urls":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/e4d14855-032f-4d69-8aec-789fadf041e0/download","http://purl.org/NET/rdflicense/allrightsreserved"],"identifier_entries":[]},"links":{"outbound_url":"https://doi.org/10.17863/CAM.119801","outbound_label":"DOI","outbound_source":"dc:identifier.doi"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Finn, Robert"]},{"key":"dc:contributor.sponsor","label":"Sponsor","values":["This work was supported by EMBL and the EMBL International PhD Programme, BBSRC and Unilever"]},{"key":"dc:creator","label":"Author","values":["Langer, Felix"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2024-07-31"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["University of Cambridge"]},{"key":"dc:relation.isreferencedby.uri","label":"Dc Relation Isreferencedby URI","values":["https://www.repository.cam.ac.uk/handle/1810/386696"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["Doctoral"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["Doctor of Philosophy (PhD)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["bioinformatics","computational biology","protein mining"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/e4d14855-032f-4d69-8aec-789fadf041e0/download","http://purl.org/NET/rdflicense/allrightsreserved"]},{"key":"dc:rights.embargodate","label":"Dc Rights Embargodate","values":["2026-07-16"]},{"key":"dc:rights.embargotype","label":"Dc Rights Embargotype","values":["embargo"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.17863/CAM.119801"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/a9bdc504-5b60-4de8-8f80-f5fa8d53dcc9/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Microorganisms are the most abundant life form on earth and are found in every naturally occurring ecosystem. Most of these microorganisms have yet to be cultured. Metagenomics- based methods allow the culture-independent analysis of any biome through sequencing the whole DNA content of an environment sample and analysing it with computational methods. MGnify is one of the most widely used platforms for the analysis of metagenomic sequences. Since 2018, MGnify has assembled and analysed over 57,000 shotgun metagenomics datasets. From these assemblies, over 2.4 billion non-redundant protein sequences have been identified. This protein database outscales other protein repositories and contains a significant fraction of functionally unexplored sequences, making it a treasure trove for protein mining. However, due to the scale of the resource, even simply downloading the raw sequence data can be challenging. Subsequently, search results can contain hundreds of thousands of matches, making exploring the results difficult. With this thesis, I outline the approaches I undertook to expand the data in the MGnify protein database to facilitate filtering and exploration of the data. To enable this to be accessible to all, new technical solutions were explored to allow rapid querying of the data. Finally, to connect search results to the underlying database, I developed an interactive platform to facilitate intuitive data mining using a comprehensive toolbox. This platform enables the rapid identification of potentially novel bioactives and enzymes within MGnify. I demonstrate the platform’s utility through a series of use cases focused on enzymes capable of degrading plastics. Based on this experience and additional use cases looking at CRIPSR- Cas systems, I extended the platform to enable the integration of multiple search results. This extension facilitates both comparative analysis and genomic context. These additional capabilities make it a valuable tool for discovering gene clusters and understanding the distribution of proteins across different environments. Overall, this thesis presents the significant expansion of one of the largest protein databases available and the development of a versatile platform for efficient protein mining. The ability to rapidly mine vast metagenomic datasets paves the way for discovering novel enzymes and bioactive compounds that can be applied in industrial processes, bioremediation or biomedical contexts."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["204df9a09800229f1cb1941e34a64a74","87eda9de84448d1f82354d60eee3eb5f"]},{"key":"dc:title","label":"Title","values":["Enabling the discovery of novel bioactives and enzymes from billions of protein sequences"]}]}],"canonical_facts":{"dc:contributor.advisor":["Finn, Robert"],"dc:contributor.sponsor":["This work was supported by EMBL and the EMBL International PhD Programme, BBSRC and Unilever"],"dc:creator":["Langer, Felix"],"dc:date.issued":["2024-07-31"],"dc:description.abstract":["Microorganisms are the most abundant life form on earth and are found in every naturally occurring ecosystem. Most of these microorganisms have yet to be cultured. Metagenomics- based methods allow the culture-independent analysis of any biome through sequencing the whole DNA content of an environment sample and analysing it with computational methods. MGnify is one of the most widely used platforms for the analysis of metagenomic sequences. Since 2018, MGnify has assembled and analysed over 57,000 shotgun metagenomics datasets. From these assemblies, over 2.4 billion non-redundant protein sequences have been identified. This protein database outscales other protein repositories and contains a significant fraction of functionally unexplored sequences, making it a treasure trove for protein mining. However, due to the scale of the resource, even simply downloading the raw sequence data can be challenging. Subsequently, search results can contain hundreds of thousands of matches, making exploring the results difficult. With this thesis, I outline the approaches I undertook to expand the data in the MGnify protein database to facilitate filtering and exploration of the data. To enable this to be accessible to all, new technical solutions were explored to allow rapid querying of the data. Finally, to connect search results to the underlying database, I developed an interactive platform to facilitate intuitive data mining using a comprehensive toolbox. This platform enables the rapid identification of potentially novel bioactives and enzymes within MGnify. I demonstrate the platform’s utility through a series of use cases focused on enzymes capable of degrading plastics. Based on this experience and additional use cases looking at CRIPSR- Cas systems, I extended the platform to enable the integration of multiple search results. This extension facilitates both comparative analysis and genomic context. These additional capabilities make it a valuable tool for discovering gene clusters and understanding the distribution of proteins across different environments. Overall, this thesis presents the significant expansion of one of the largest protein databases available and the development of a versatile platform for efficient protein mining. The ability to rapidly mine vast metagenomic datasets paves the way for discovering novel enzymes and bioactive compounds that can be applied in industrial processes, bioremediation or biomedical contexts."],"dc:format.checksum.md5":["204df9a09800229f1cb1941e34a64a74","87eda9de84448d1f82354d60eee3eb5f"],"dc:identifier.doi":["https://doi.org/10.17863/CAM.119801"],"dc:identifier.uri":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/a9bdc504-5b60-4de8-8f80-f5fa8d53dcc9/download"],"dc:publisher.institution":["University of Cambridge"],"dc:relation.isreferencedby.uri":["https://www.repository.cam.ac.uk/handle/1810/386696"],"dc:rights":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/e4d14855-032f-4d69-8aec-789fadf041e0/download","http://purl.org/NET/rdflicense/allrightsreserved"],"dc:rights.embargodate":["2026-07-16"],"dc:rights.embargotype":["embargo"],"dc:subject":["bioinformatics","computational biology","protein mining"],"dc:title":["Enabling the discovery of novel bioactives and enzymes from billions of protein sequences"],"dc:type":["Thesis"],"dc:type.qualificationlevel":["Doctoral"],"dc:type.qualificationname":["Doctor of Philosophy (PhD)"]},"updated_at":"2026-07-22T22:24:32Z"}