{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/129918"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/129918","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Multiple sequence alignment with sequence length heterogeneity and its applications","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-20 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2025-10-20 without embargo terms","abstract_has_math":false,"creators":["Shen, Chengze"],"institution":"University of Illinois Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Warnow, Tandy","El-Kebir, Mohammed","Gropp, William D","Williams, Kelly P."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-07-07","date_published":"2025-07-07","updated_at":"2026-07-22T22:25:06Z","subjects":["Multiple Sequence Alignment","Metagenomic Analysis"],"languages":["en","eng"],"rights":["Copyright 2025 Chengze Shen"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/129918","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Warnow, Tandy","El-Kebir, Mohammed","Gropp, William D","Williams, Kelly P."]},{"key":"dc:creator","label":"Author","values":["Shen, Chengze"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-07-07","2025-08"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Multiple Sequence Alignment","Metagenomic Analysis"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Chengze Shen"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/129918"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-20 without embargo terms","The student, Chengze Shen, accepted the attached license on 2025-07-04 at 08:50.","The student, Chengze Shen, submitted this Dissertation for approval on 2025-07-04 at 08:55.","This Dissertation was approved for publication on 2025-07-07 at 15:13.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22409 on 2025-10-20 at 20:14:56","Two major challenges in the field of bioinformatics are multiple sequence alignment (MSA) and abundance profiling. Both are hard problems that can become computationally prohibitive with a large amount of input data, which most existing methods can fail to handle. In addition, only a few approaches can adequately align input sequences with various lengths (i.e., sequence length heterogeneity), which are often present in metagenomic data used for profiling species abundances. With the increasing availability and size of sequence data that may have sequence length heterogeneity, new approaches are required to perform scalable and accurate analyses. In this thesis, we show our progress in developing new alignment methods that deal with sequence length heterogeneity and can scale to large data using divide-and-conquer. We also apply the new alignment methods to perform abundance profiling and demonstrate improved accuracy. We hope that the newer methods can help scientists conduct more accurate and effective analyses on larger and larger datasets."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Multiple sequence alignment with sequence length heterogeneity and its applications"]}]}],"canonical_facts":{"dc:contributor":["Warnow, Tandy","El-Kebir, Mohammed","Gropp, William D","Williams, Kelly P."],"dc:creator":["Shen, Chengze"],"dc:date":["2025-07-07","2025-08"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-20 without embargo terms","The student, Chengze Shen, accepted the attached license on 2025-07-04 at 08:50.","The student, Chengze Shen, submitted this Dissertation for approval on 2025-07-04 at 08:55.","This Dissertation was approved for publication on 2025-07-07 at 15:13.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22409 on 2025-10-20 at 20:14:56","Two major challenges in the field of bioinformatics are multiple sequence alignment (MSA) and abundance profiling. Both are hard problems that can become computationally prohibitive with a large amount of input data, which most existing methods can fail to handle. In addition, only a few approaches can adequately align input sequences with various lengths (i.e., sequence length heterogeneity), which are often present in metagenomic data used for profiling species abundances. With the increasing availability and size of sequence data that may have sequence length heterogeneity, new approaches are required to perform scalable and accurate analyses. In this thesis, we show our progress in developing new alignment methods that deal with sequence length heterogeneity and can scale to large data using divide-and-conquer. We also apply the new alignment methods to perform abundance profiling and demonstrate improved accuracy. We hope that the newer methods can help scientists conduct more accurate and effective analyses on larger and larger datasets."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/129918"],"dc:language":["en","eng"],"dc:rights":["Copyright 2025 Chengze Shen"],"dc:subject":["Multiple Sequence Alignment","Metagenomic Analysis"],"dc:title":["Multiple sequence alignment with sequence length heterogeneity and its applications"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:06Z"}