{"id":{"repo_id":"buffalo","oai_identifier":"oai:ubir.buffalo.edu:10477/79919"},"canonical_url":"https://search.dev.ndltd.org/etd/buffalo/oai:ubir.buffalo.edu:10477/79919","repository":{"repo_id":"buffalo","name":"Buffalo","base_url":"https://ubir.buffalo.edu/oai/request"},"display":{"title":"Efficient Sequence Clustering and Embedding Algorithms for Large-Scale Metagenomics Data","abstract":"Ph.D.","abstract_html":"Ph.D.","abstract_has_math":false,"creators":["Zheng, Wei"],"institution":"State University of New York at Buffalo","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":["Sun, Yijun","Computer Science and Engineering"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019-07-30T15:11:06Z","date_published":"2019-07-30T15:11:06Z","updated_at":"2026-07-27T19:05:21Z","subjects":["computer science"],"languages":["eng"],"rights":["Users of works found in University at Buffalo Institutional Repository (UBIR) are responsible for identifying and contacting the copyright owner for permission to reuse. University at Buffalo Libraries do not manage rights for copyright-protected works and cannot assist with permissions.","Copyright retained by author."],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/10477/79919","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Sun, Yijun","Computer Science and Engineering"]},{"key":"dc:creator","label":"Author","values":["Zheng, Wei"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2019-07-30T15:11:06Z","2019","2019-05-13 09:21:28"]},{"key":"dc:publisher","label":"Institution","values":["State University of New York at Buffalo"]},{"key":"dc:type","label":"Dc Type","values":["Text","Dissertation"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["computer science"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Users of works found in University at Buffalo Institutional Repository (UBIR) are responsible for identifying and contacting the copyright owner for permission to reuse. University at Buffalo Libraries do not manage rights for copyright-protected works and cannot assist with permissions.","Copyright retained by author."]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/10477/79919"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Ph.D.","While the rapid accumulation of genomic information represents a valuable source to significantly expand biological knowledge, the flood of data poses a formidable challenge in data analysis, demanding new computational algorithms for efficient data processing. Usually, the first major step in processing 16S rRNA sequence data is to bin sequences into taxonomic or genotypic units, which forms the basis for performing ecological statistics and comparative studies. Hierarchical clustering is one of the most widely used approaches for sequence binning. However, due to the quadratic computation complexity of both hierarchical clustering and sequence distance calculation, It is very limited when dealing with large-scale sequence data.To address the computational challenge, we propose two efficient algorithms for large-scale metegenomics data, namely SLAD and SENSE. SLAD is a generic computational framework that can be used to parallelize various de novo OTU picking methods and comes with theoretical guarantees on both accuracy and efficiency. SENSE is a siamese neural network for efficient and accurate alignment-free sequence comparison by projecting sequences into an embedding space where the distance calculation is linear to embedding dimension. In the meanwhile, the mean square error between alignment distances and pairwise distances defined in the embedding space is minimized."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Efficient Sequence Clustering and Embedding Algorithms for Large-Scale Metagenomics Data"]}]}],"canonical_facts":{"dc:contributor":["Sun, Yijun","Computer Science and Engineering"],"dc:creator":["Zheng, Wei"],"dc:date":["2019-07-30T15:11:06Z","2019","2019-05-13 09:21:28"],"dc:description":["Ph.D.","While the rapid accumulation of genomic information represents a valuable source to significantly expand biological knowledge, the flood of data poses a formidable challenge in data analysis, demanding new computational algorithms for efficient data processing. Usually, the first major step in processing 16S rRNA sequence data is to bin sequences into taxonomic or genotypic units, which forms the basis for performing ecological statistics and comparative studies. Hierarchical clustering is one of the most widely used approaches for sequence binning. However, due to the quadratic computation complexity of both hierarchical clustering and sequence distance calculation, It is very limited when dealing with large-scale sequence data.To address the computational challenge, we propose two efficient algorithms for large-scale metegenomics data, namely SLAD and SENSE. SLAD is a generic computational framework that can be used to parallelize various de novo OTU picking methods and comes with theoretical guarantees on both accuracy and efficiency. SENSE is a siamese neural network for efficient and accurate alignment-free sequence comparison by projecting sequences into an embedding space where the distance calculation is linear to embedding dimension. In the meanwhile, the mean square error between alignment distances and pairwise distances defined in the embedding space is minimized."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/10477/79919"],"dc:language":["eng"],"dc:publisher":["State University of New York at Buffalo"],"dc:rights":["Users of works found in University at Buffalo Institutional Repository (UBIR) are responsible for identifying and contacting the copyright owner for permission to reuse. University at Buffalo Libraries do not manage rights for copyright-protected works and cannot assist with permissions.","Copyright retained by author."],"dc:subject":["computer science"],"dc:title":["Efficient Sequence Clustering and Embedding Algorithms for Large-Scale Metagenomics Data"],"dc:type":["Text","Dissertation"]},"updated_at":"2026-07-27T19:05:21Z"}