{"id":{"repo_id":"rutgers","oai_identifier":"oai:example.org:rutgers-lib:45449"},"canonical_url":"https://search.dev.ndltd.org/etd/rutgers/oai:example.org:rutgers-lib:45449","repository":{"repo_id":"rutgers","name":"Rutgers University","base_url":"https://rucore.libraries.rutgers.edu/api/oai-pmh/"},"display":{"title":"Improving genome assembly by identifying reliable sequencing data","abstract":"De novo Genome assembly and k-mer frequency counting are two of the classical prob- lems of Bioinformatics. k-mer counting helps to identify genomic k-mers from sequenced reads which may then inform read correction or genome assembly. Genome assembly has two major subproblems: contig construction and scaffolding. A contig is a continu- ous sub-sequence of the genome assembled from sequencing reads. Scaffolding attempts to construct a linear sequence of contigs (with possible gaps in between) using paired reads (two reads whose distance on the genome is approximately known). In this the- sis I will present a new computationally efficient tool for identifying frequent k-mers which are more likely to be genomic, and a set of linear inequalities which can improve scaffolding (which is known to be NP-hard) by identifying reliable paired reads. Identifying reliable k-mers from Whole Genome Amplification (WGA) data is more challenging compared to multi-cell data due to the coverage variation introduced by the amplification step (MDA, MALBEC, etc.), which implies that applying a simple k- mer frequency cutoff is unreasonable. We observed that with sufficient coverage, using partial reads (read prefix of a certain length) of length approximately twice or less than that of the k-mer length recovers a large proportion of genomic k-mers while keeping the proportion of erroneous k-mers low. We show that using partial reads for assembly ii and gene prediction recovers a significant proportion of genes and propose to use this approach for rapid pathogen detection in combination with Single Cell Genomics (SCG). Thanks to SCG, it is now possible to isolate one single cell from environmental sam- ple, extract its DNA and perform genetic sequencing without any need for culturing the cell in the lab. We show that current bioinformatic tools are capable of charac- terizing a novel organism by producing a draft genome assembly and gene annotation from single cell data of a MAST-4 stramenopile. This demonstrates the potential of SCG for genetic study of the vast majority of environmental organisms that has so far eluded scientists as they cannot be brought into culture, typically a necessity for future studies.","abstract_html":"De novo Genome assembly and k-mer frequency counting are two of the classical prob- lems of Bioinformatics. k-mer counting helps to identify genomic k-mers from sequenced reads which may then inform read correction or genome assembly. Genome assembly has two major subproblems: contig construction and scaffolding. A contig is a continu- ous sub-sequence of the genome assembled from sequencing reads. Scaffolding attempts to construct a linear sequence of contigs (with possible gaps in between) using paired reads (two reads whose distance on the genome is approximately known). In this the- sis I will present a new computationally efficient tool for identifying frequent k-mers which are more likely to be genomic, and a set of linear inequalities which can improve scaffolding (which is known to be NP-hard) by identifying reliable paired reads. Identifying reliable k-mers from Whole Genome Amplification (WGA) data is more challenging compared to multi-cell data due to the coverage variation introduced by the amplification step (MDA, MALBEC, etc.), which implies that applying a simple k- mer frequency cutoff is unreasonable. We observed that with sufficient coverage, using partial reads (read prefix of a certain length) of length approximately twice or less than that of the k-mer length recovers a large proportion of genomic k-mers while keeping the proportion of erroneous k-mers low. We show that using partial reads for assembly ii and gene prediction recovers a significant proportion of genes and propose to use this approach for rapid pathogen detection in combination with Single Cell Genomics (SCG). Thanks to SCG, it is now possible to isolate one single cell from environmental sam- ple, extract its DNA and perform genetic sequencing without any need for culturing the cell in the lab. We show that current bioinformatic tools are capable of charac- terizing a novel organism by producing a draft genome assembly and gene annotation from single cell data of a MAST-4 stramenopile. This demonstrates the potential of SCG for genetic study of the vast majority of environmental organisms that has so far eluded scientists as they cannot be brought into culture, typically a necessity for future studies.","abstract_has_math":false,"creators":[],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":["Roy, Rajat Shuvro, 1983- (author)","Schliep, Alexander (chair)","Bhattacharya, Debashish (co-chair)","Chen, Kevin (internal member)","Farach-Colton, Martin (internal member)","Grigoriev, Andrey (outside member)","Rutgers University","Graduate School - New Brunswick"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2014,"date_issued":"2014","date_published":"2014","updated_at":"2026-07-27T20:49:29Z","subjects":["Computer Science","Genomes--Analysis","Gene amplification"],"languages":["eng"],"rights":["The author owns the copyright to this work."],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["rutgers-lib:45449","ETD_5745"],"render_values":[{"text":"rutgers-lib:45449","href":null,"code":true},{"text":"ETD_5745","href":null,"code":true}]}]},"links":{"outbound_url":null,"outbound_label":null,"outbound_source":null},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Roy, Rajat Shuvro, 1983- (author)","Schliep, Alexander (chair)","Bhattacharya, Debashish (co-chair)","Chen, Kevin (internal member)","Farach-Colton, Martin (internal member)","Grigoriev, Andrey (outside member)","Rutgers University","Graduate School - New Brunswick"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2014","2014-10"]},{"key":"dc:relation","label":"Dc Relation","values":["Rutgers University Electronic Theses and Dissertations","ETD","Graduate School - New Brunswick Electronic Theses and Dissertations","rucore19991600001"]},{"key":"dc:type","label":"Dc Type","values":["Text","theses"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer Science","Genomes--Analysis","Gene amplification"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["The author owns the copyright to this work."]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["rutgers-lib:45449","ETD_5745"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["De novo Genome assembly and k-mer frequency counting are two of the classical prob- lems of Bioinformatics. k-mer counting helps to identify genomic k-mers from sequenced reads which may then inform read correction or genome assembly. Genome assembly has two major subproblems: contig construction and scaffolding. A contig is a continu- ous sub-sequence of the genome assembled from sequencing reads. Scaffolding attempts to construct a linear sequence of contigs (with possible gaps in between) using paired reads (two reads whose distance on the genome is approximately known). In this the- sis I will present a new computationally efficient tool for identifying frequent k-mers which are more likely to be genomic, and a set of linear inequalities which can improve scaffolding (which is known to be NP-hard) by identifying reliable paired reads. Identifying reliable k-mers from Whole Genome Amplification (WGA) data is more challenging compared to multi-cell data due to the coverage variation introduced by the amplification step (MDA, MALBEC, etc.), which implies that applying a simple k- mer frequency cutoff is unreasonable. We observed that with sufficient coverage, using partial reads (read prefix of a certain length) of length approximately twice or less than that of the k-mer length recovers a large proportion of genomic k-mers while keeping the proportion of erroneous k-mers low. We show that using partial reads for assembly ii and gene prediction recovers a significant proportion of genes and propose to use this approach for rapid pathogen detection in combination with Single Cell Genomics (SCG). Thanks to SCG, it is now possible to isolate one single cell from environmental sam- ple, extract its DNA and perform genetic sequencing without any need for culturing the cell in the lab. We show that current bioinformatic tools are capable of charac- terizing a novel organism by producing a draft genome assembly and gene annotation from single cell data of a MAST-4 stramenopile. This demonstrates the potential of SCG for genetic study of the vast majority of environmental organisms that has so far eluded scientists as they cannot be brought into culture, typically a necessity for future studies.","Ph.D.","Includes bibliographical references","by Rajat Shuvro Roy"]},{"key":"dc:format","label":"Dc Format","values":["1 online resource (xi, 120 p. : ill.)","electronic resource","application/pdf"]},{"key":"dc:title","label":"Title","values":["Improving genome assembly by identifying reliable sequencing data"]}]}],"canonical_facts":{"dc:contributor":["Roy, Rajat Shuvro, 1983- (author)","Schliep, Alexander (chair)","Bhattacharya, Debashish (co-chair)","Chen, Kevin (internal member)","Farach-Colton, Martin (internal member)","Grigoriev, Andrey (outside member)","Rutgers University","Graduate School - New Brunswick"],"dc:date":["2014","2014-10"],"dc:description":["De novo Genome assembly and k-mer frequency counting are two of the classical prob- lems of Bioinformatics. k-mer counting helps to identify genomic k-mers from sequenced reads which may then inform read correction or genome assembly. Genome assembly has two major subproblems: contig construction and scaffolding. A contig is a continu- ous sub-sequence of the genome assembled from sequencing reads. Scaffolding attempts to construct a linear sequence of contigs (with possible gaps in between) using paired reads (two reads whose distance on the genome is approximately known). In this the- sis I will present a new computationally efficient tool for identifying frequent k-mers which are more likely to be genomic, and a set of linear inequalities which can improve scaffolding (which is known to be NP-hard) by identifying reliable paired reads. Identifying reliable k-mers from Whole Genome Amplification (WGA) data is more challenging compared to multi-cell data due to the coverage variation introduced by the amplification step (MDA, MALBEC, etc.), which implies that applying a simple k- mer frequency cutoff is unreasonable. We observed that with sufficient coverage, using partial reads (read prefix of a certain length) of length approximately twice or less than that of the k-mer length recovers a large proportion of genomic k-mers while keeping the proportion of erroneous k-mers low. We show that using partial reads for assembly ii and gene prediction recovers a significant proportion of genes and propose to use this approach for rapid pathogen detection in combination with Single Cell Genomics (SCG). Thanks to SCG, it is now possible to isolate one single cell from environmental sam- ple, extract its DNA and perform genetic sequencing without any need for culturing the cell in the lab. We show that current bioinformatic tools are capable of charac- terizing a novel organism by producing a draft genome assembly and gene annotation from single cell data of a MAST-4 stramenopile. This demonstrates the potential of SCG for genetic study of the vast majority of environmental organisms that has so far eluded scientists as they cannot be brought into culture, typically a necessity for future studies.","Ph.D.","Includes bibliographical references","by Rajat Shuvro Roy"],"dc:format":["1 online resource (xi, 120 p. : ill.)","electronic resource","application/pdf"],"dc:identifier":["rutgers-lib:45449","ETD_5745"],"dc:language":["eng"],"dc:relation":["Rutgers University Electronic Theses and Dissertations","ETD","Graduate School - New Brunswick Electronic Theses and Dissertations","rucore19991600001"],"dc:rights":["The author owns the copyright to this work."],"dc:subject":["Computer Science","Genomes--Analysis","Gene amplification"],"dc:title":["Improving genome assembly by identifying reliable sequencing data"],"dc:type":["Text","theses"]},"updated_at":"2026-07-27T20:49:29Z"}