{"id":{"repo_id":"vt","oai_identifier":"oai:vtechworks.lib.vt.edu:10919/112828"},"canonical_url":"https://search.dev.ndltd.org/etd/vt/oai:vtechworks.lib.vt.edu:10919/112828","repository":{"repo_id":"vt","name":"Virginia Tech","base_url":"https://vtechworks.lib.vt.edu/oai/request"},"display":{"title":"DeepARG+ - A Computational Pipeline for the Prediction of Antibiotic Resistance","abstract":"The global spread of antibiotic resistance warrants concerted surveillance in the clinic and in the environment. The widespread use of metagenomics for various studies has led to the generation of a large amount of sequencing data. Next-generation sequencing of microbial communities provides an opportunity for proactive detection of emerging antibiotic resistance genes (ARGs) from such data, but there are a limited number of pipelines that enable the identification of novel ARGs belonging to diverse antibiotic classes at present. Therefore, there is a need for the development of computational pipelines that can identify these putative novel ARGs. Such pipelines should be scalable, accessible and have good performance. To address this problem we develop a new method for predicting novel ARGs from genomic or metagenomic sequences, leveraging known ARGs of different resistance categories. Our method takes into account the physio-chemical properties that are intrinsic to different ARG families. Traditionally, new ARGs are predicted by making sequence alignment and calculating sequence similarity to existing ARG reference databases, which can be very time consuming. Here we introduce an alignment free and deep learning prediction method that incorporates both the primary protein sequences of ARGs and their physio-chemical properties. We compare our method with existing pipelines including hidden Markov model based Resfams and fARGene, sequence alignment and machine learning-based DeepARG-LS, and homology modelling based Pairwise Comparative Modelling. We also use our model to detect novel ARGs from various environments including human-gut, soil, activated sludge and the influent samples collected from a waste water treatment plant. Results show that our method achieves greater accuracy compared to existing models for the prediction of ARGs and enables the detection of putative novel ARGs, providing promising targets for experimental characterization to the scientific community.","abstract_html":"The global spread of antibiotic resistance warrants concerted surveillance in the clinic and in the environment. The widespread use of metagenomics for various studies has led to the generation of a large amount of sequencing data. Next-generation sequencing of microbial communities provides an opportunity for proactive detection of emerging antibiotic resistance genes (ARGs) from such data, but there are a limited number of pipelines that enable the identification of novel ARGs belonging to diverse antibiotic classes at present. Therefore, there is a need for the development of computational pipelines that can identify these putative novel ARGs. Such pipelines should be scalable, accessible and have good performance. To address this problem we develop a new method for predicting novel ARGs from genomic or metagenomic sequences, leveraging known ARGs of different resistance categories. Our method takes into account the physio-chemical properties that are intrinsic to different ARG families. Traditionally, new ARGs are predicted by making sequence alignment and calculating sequence similarity to existing ARG reference databases, which can be very time consuming. Here we introduce an alignment free and deep learning prediction method that incorporates both the primary protein sequences of ARGs and their physio-chemical properties. We compare our method with existing pipelines including hidden Markov model based Resfams and fARGene, sequence alignment and machine learning-based DeepARG-LS, and homology modelling based Pairwise Comparative Modelling. We also use our model to detect novel ARGs from various environments including human-gut, soil, activated sludge and the influent samples collected from a waste water treatment plant. Results show that our method achieves greater accuracy compared to existing models for the prediction of ARGs and enables the detection of putative novel ARGs, providing promising targets for experimental characterization to the scientific community.","abstract_has_math":false,"creators":["Kulkarni, Rutwik Shashank"],"institution":"Virginia Tech","degree_name":"Master of Science","degree_level":"masters","degree_discipline":"Computer Science and Applications","degree_department":"Computer Science","school":null,"contributors":[],"advisors":[],"committee_chairs":["Zhang, Liqing"],"committee_members":["Pruden, Amy","Karpatne, Anuj"],"year":2021,"date_issued":"2021-06-16","date_published":"2021-06-16","updated_at":"2026-07-22T22:20:34Z","subjects":["Antibiotic Resistance","Deep Learning","Machine Learning","Protein Structure"],"languages":[],"rights":["In Copyright"],"rights_urls":["http://rightsstatements.org/vocab/InC/1.0/"],"identifier_entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:31271"],"render_values":[{"text":"vt_gsexam:31271","href":null,"code":true}]}]},"links":{"outbound_url":"http://hdl.handle.net/10919/112828","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.committeechair","label":"Committee Chair","values":["Zhang, Liqing"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Pruden, Amy","Karpatne, Anuj"]},{"key":"dc:contributor.department","label":"Department","values":["Computer Science"]},{"key":"dc:creator","label":"Author","values":["Kulkarni, Rutwik Shashank"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2022-12-09T07:00:39Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2022-12-09T07:00:39Z"]},{"key":"dc:date.issued","label":"Date","values":["2021-06-16"]},{"key":"dc:publisher","label":"Institution","values":["Virginia Tech"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science and Applications"]},{"key":"thesis:degree_level","label":"Degree Level","values":["masters"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Virginia Polytechnic Institute and State University"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Antibiotic Resistance","Deep Learning","Machine Learning","Protein Structure"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["In Copyright"]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://rightsstatements.org/vocab/InC/1.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:31271"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["http://hdl.handle.net/10919/112828"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["The global spread of antibiotic resistance warrants concerted surveillance in the clinic and in the environment. The widespread use of metagenomics for various studies has led to the generation of a large amount of sequencing data. Next-generation sequencing of microbial communities provides an opportunity for proactive detection of emerging antibiotic resistance genes (ARGs) from such data, but there are a limited number of pipelines that enable the identification of novel ARGs belonging to diverse antibiotic classes at present. Therefore, there is a need for the development of computational pipelines that can identify these putative novel ARGs. Such pipelines should be scalable, accessible and have good performance. To address this problem we develop a new method for predicting novel ARGs from genomic or metagenomic sequences, leveraging known ARGs of different resistance categories. Our method takes into account the physio-chemical properties that are intrinsic to different ARG families. Traditionally, new ARGs are predicted by making sequence alignment and calculating sequence similarity to existing ARG reference databases, which can be very time consuming. Here we introduce an alignment free and deep learning prediction method that incorporates both the primary protein sequences of ARGs and their physio-chemical properties. We compare our method with existing pipelines including hidden Markov model based Resfams and fARGene, sequence alignment and machine learning-based DeepARG-LS, and homology modelling based Pairwise Comparative Modelling. We also use our model to detect novel ARGs from various environments including human-gut, soil, activated sludge and the influent samples collected from a waste water treatment plant. Results show that our method achieves greater accuracy compared to existing models for the prediction of ARGs and enables the detection of putative novel ARGs, providing promising targets for experimental characterization to the scientific community."]},{"key":"dc:description.abstractgeneral","label":"General Abstract","values":["Various bacteria contain genes that allow them to survive and grow even after the application of antibiotics. Such genes are called antibiotic resistance genes (ARGs). Each ARG has properties that make it resistant to a particular class of antibiotics. This class is called the resistance class/category of the gene. Antimicrobial resistance (AMR) is one of the biggest challenges to public health in recent times. It has been projected that a large number of deaths might occur due to AMR in the future. Therefore, there is a need for monitoring AMR in various environments. Currently, developed methods use the sequence's similarity with the existing database as a feature for ARG prediction. Some tools also use the 3D structure of proteins as a feature for ARG prediction. In this thesis, we develop a tool that incorporates both the sequence similarity and the structural information of proteins for ARG prediction. The structural information is encoded with physio-chemical properties (such as hydrophobicity, molecular weight etc.) of the amino acids. Our results show the efficacy of the pipeline in various environments. Results also show that our method achieves accuracy greater than existing models for the prediction of ARGs from metagenomic data. It also enables the detection of putative novel ARGs, providing promising targets for experimental characterization to the scientific community."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["Master of Science"]},{"key":"dc:format.medium","label":"Dc Format Medium","values":["ETD"]},{"key":"dc:title","label":"Title","values":["DeepARG+ - A Computational Pipeline for the Prediction of Antibiotic Resistance"]}]}],"canonical_facts":{"dc:contributor.committeechair":["Zhang, Liqing"],"dc:contributor.committeemember":["Pruden, Amy","Karpatne, Anuj"],"dc:contributor.department":["Computer Science"],"dc:creator":["Kulkarni, Rutwik Shashank"],"dc:date.accessioned":["2022-12-09T07:00:39Z"],"dc:date.available":["2022-12-09T07:00:39Z"],"dc:date.issued":["2021-06-16"],"dc:description.abstract":["The global spread of antibiotic resistance warrants concerted surveillance in the clinic and in the environment. The widespread use of metagenomics for various studies has led to the generation of a large amount of sequencing data. Next-generation sequencing of microbial communities provides an opportunity for proactive detection of emerging antibiotic resistance genes (ARGs) from such data, but there are a limited number of pipelines that enable the identification of novel ARGs belonging to diverse antibiotic classes at present. Therefore, there is a need for the development of computational pipelines that can identify these putative novel ARGs. Such pipelines should be scalable, accessible and have good performance. To address this problem we develop a new method for predicting novel ARGs from genomic or metagenomic sequences, leveraging known ARGs of different resistance categories. Our method takes into account the physio-chemical properties that are intrinsic to different ARG families. Traditionally, new ARGs are predicted by making sequence alignment and calculating sequence similarity to existing ARG reference databases, which can be very time consuming. Here we introduce an alignment free and deep learning prediction method that incorporates both the primary protein sequences of ARGs and their physio-chemical properties. We compare our method with existing pipelines including hidden Markov model based Resfams and fARGene, sequence alignment and machine learning-based DeepARG-LS, and homology modelling based Pairwise Comparative Modelling. We also use our model to detect novel ARGs from various environments including human-gut, soil, activated sludge and the influent samples collected from a waste water treatment plant. Results show that our method achieves greater accuracy compared to existing models for the prediction of ARGs and enables the detection of putative novel ARGs, providing promising targets for experimental characterization to the scientific community."],"dc:description.abstractgeneral":["Various bacteria contain genes that allow them to survive and grow even after the application of antibiotics. Such genes are called antibiotic resistance genes (ARGs). Each ARG has properties that make it resistant to a particular class of antibiotics. This class is called the resistance class/category of the gene. Antimicrobial resistance (AMR) is one of the biggest challenges to public health in recent times. It has been projected that a large number of deaths might occur due to AMR in the future. Therefore, there is a need for monitoring AMR in various environments. Currently, developed methods use the sequence's similarity with the existing database as a feature for ARG prediction. Some tools also use the 3D structure of proteins as a feature for ARG prediction. In this thesis, we develop a tool that incorporates both the sequence similarity and the structural information of proteins for ARG prediction. The structural information is encoded with physio-chemical properties (such as hydrophobicity, molecular weight etc.) of the amino acids. Our results show the efficacy of the pipeline in various environments. Results also show that our method achieves accuracy greater than existing models for the prediction of ARGs from metagenomic data. It also enables the detection of putative novel ARGs, providing promising targets for experimental characterization to the scientific community."],"dc:description.degree":["Master of Science"],"dc:format.medium":["ETD"],"dc:identifier.other":["vt_gsexam:31271"],"dc:identifier.uri":["http://hdl.handle.net/10919/112828"],"dc:publisher":["Virginia Tech"],"dc:rights":["In Copyright"],"dc:rights.uri":["http://rightsstatements.org/vocab/InC/1.0/"],"dc:subject":["Antibiotic Resistance","Deep Learning","Machine Learning","Protein Structure"],"dc:title":["DeepARG+ - A Computational Pipeline for the Prediction of Antibiotic Resistance"],"dc:type":["Thesis"],"thesis:degree_discipline":["Computer Science and Applications"],"thesis:degree_level":["masters"],"thesis:degree_name":["Master of Science"],"thesis:institution_name":["Virginia Polytechnic Institute and State University"]},"updated_at":"2026-07-22T22:20:34Z"}