{"id":{"repo_id":"claremont","oai_identifier":"oai:scholarship.claremont.edu:cgu_etd-1300"},"canonical_url":"https://search.dev.ndltd.org/etd/claremont/oai:scholarship.claremont.edu:cgu_etd-1300","repository":{"repo_id":"claremont","name":"Claremont Graduate University","base_url":"https://scholarship.claremont.edu/do/oai/"},"display":{"title":"Machine Learning Methods for the Analysis of Metagenomes","abstract":"<p>As of October 2020, there are 18.6 × 1015 DNA base pairs publicly available in the Sequence Read Archive and this number is growing at an exponential rate. As DNA sequencing prices continue to drop, many research groups around the world have incorporated high throughput sequencing in their research, giving us access to sequences from many distinct ecosystems. This has revolutionized the field of metagenomics, which aims to fully characterize all organisms and their interactions in a particular system. Nevertheless, the plethora of available data has made its analysis difficult as traditional techniques such as genome assembly or sequence alignment are bound to fail due to the high noise of metagenomes, or take an impractically long time due to their size. Through this thesis, we explore those challenges and develop techniques to meet them. Chapter 1 serves as an introduction to the fields of metagenomics and machine learning and the applications where the two meet. Chapter 2 examines the different kinds of noises in sequencing datasets and presents PRINSEQ++, a C++ multi-threaded software for quality control of sequencing datasets. Chapter 3 describes the analysis of 63 metagenomic samples from children with ”nodding syndrome” using Random Forest to give insights into the etiology of the disease. Chapter 4 explores the use of artificial neutral networks to classify phage structural proteins derived from metagenomes.</p>","abstract_html":"&lt;p&gt;As of October 2020, there are 18.6 × 1015 DNA base pairs publicly available in the Sequence Read Archive and this number is growing at an exponential rate. As DNA sequencing prices continue to drop, many research groups around the world have incorporated high throughput sequencing in their research, giving us access to sequences from many distinct ecosystems. This has revolutionized the field of metagenomics, which aims to fully characterize all organisms and their interactions in a particular system. Nevertheless, the plethora of available data has made its analysis difficult as traditional techniques such as genome assembly or sequence alignment are bound to fail due to the high noise of metagenomes, or take an impractically long time due to their size. Through this thesis, we explore those challenges and develop techniques to meet them. Chapter 1 serves as an introduction to the fields of metagenomics and machine learning and the applications where the two meet. Chapter 2 examines the different kinds of noises in sequencing datasets and presents PRINSEQ++, a C++ multi-threaded software for quality control of sequencing datasets. Chapter 3 describes the analysis of 63 metagenomic samples from children with ”nodding syndrome” using Random Forest to give insights into the etiology of the disease. Chapter 4 explores the use of artificial neutral networks to classify phage structural proteins derived from metagenomes.&lt;/p&gt;","abstract_has_math":false,"creators":["Cantu Alessio Robles, Vito Adrian"],"institution":null,"degree_name":"Computational Science Joint PhD with San Diego State University, PhD","degree_level":"Open Access Dissertation","degree_discipline":"Institute of Mathematical Sciences","degree_department":null,"school":null,"contributors":["Claudia Rangel","Anca Segall","Allon Percus"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2020,"date_issued":"2020-01-01T08:00:00Z","date_published":"2020-01-01T08:00:00Z","updated_at":"2026-07-24T01:39:50Z","subjects":["Artificial Intelligence and Robotics","Bioinformatics","Genetics"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://scholarship.claremont.edu/cgu_etd/276","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Claudia Rangel","Anca Segall","Allon Percus"]},{"key":"dc:creator","label":"Author","values":["Cantu Alessio Robles, Vito Adrian"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2022-02-28T08:00:00Z"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Institute of Mathematical Sciences"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Open Access Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Computational Science Joint PhD with San Diego State University, PhD"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Artificial Intelligence and Robotics","Bioinformatics","Genetics"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://scholarship.claremont.edu/cgu_etd/276"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>As of October 2020, there are 18.6 × 1015 DNA base pairs publicly available in the Sequence Read Archive and this number is growing at an exponential rate. As DNA sequencing prices continue to drop, many research groups around the world have incorporated high throughput sequencing in their research, giving us access to sequences from many distinct ecosystems. This has revolutionized the field of metagenomics, which aims to fully characterize all organisms and their interactions in a particular system. Nevertheless, the plethora of available data has made its analysis difficult as traditional techniques such as genome assembly or sequence alignment are bound to fail due to the high noise of metagenomes, or take an impractically long time due to their size. Through this thesis, we explore those challenges and develop techniques to meet them. Chapter 1 serves as an introduction to the fields of metagenomics and machine learning and the applications where the two meet. Chapter 2 examines the different kinds of noises in sequencing datasets and presents PRINSEQ++, a C++ multi-threaded software for quality control of sequencing datasets. Chapter 3 describes the analysis of 63 metagenomic samples from children with ”nodding syndrome” using Random Forest to give insights into the etiology of the disease. Chapter 4 explores the use of artificial neutral networks to classify phage structural proteins derived from metagenomes.</p>"]},{"key":"dc:title","label":"Title","values":["Machine Learning Methods for the Analysis of Metagenomes"]}]}],"canonical_facts":{"dc:contributor":["Claudia Rangel","Anca Segall","Allon Percus"],"dc:creator":["Cantu Alessio Robles, Vito Adrian"],"dc:date.available":["2022-02-28T08:00:00Z"],"dc:description.abstract":["<p>As of October 2020, there are 18.6 × 1015 DNA base pairs publicly available in the Sequence Read Archive and this number is growing at an exponential rate. As DNA sequencing prices continue to drop, many research groups around the world have incorporated high throughput sequencing in their research, giving us access to sequences from many distinct ecosystems. This has revolutionized the field of metagenomics, which aims to fully characterize all organisms and their interactions in a particular system. Nevertheless, the plethora of available data has made its analysis difficult as traditional techniques such as genome assembly or sequence alignment are bound to fail due to the high noise of metagenomes, or take an impractically long time due to their size. Through this thesis, we explore those challenges and develop techniques to meet them. Chapter 1 serves as an introduction to the fields of metagenomics and machine learning and the applications where the two meet. Chapter 2 examines the different kinds of noises in sequencing datasets and presents PRINSEQ++, a C++ multi-threaded software for quality control of sequencing datasets. Chapter 3 describes the analysis of 63 metagenomic samples from children with ”nodding syndrome” using Random Forest to give insights into the etiology of the disease. Chapter 4 explores the use of artificial neutral networks to classify phage structural proteins derived from metagenomes.</p>"],"dc:identifier":["https://scholarship.claremont.edu/cgu_etd/276"],"dc:subject":["Artificial Intelligence and Robotics","Bioinformatics","Genetics"],"dc:title":["Machine Learning Methods for the Analysis of Metagenomes"],"thesis:degree_discipline":["Institute of Mathematical Sciences"],"thesis:degree_level":["Open Access Dissertation"],"thesis:degree_name":["Computational Science Joint PhD with San Diego State University, PhD"]},"updated_at":"2026-07-24T01:39:50Z"}