University of Illinois Urbana-Champaign
Multiple sequence alignment with sequence length heterogeneity and its applications
Abstract
dc:descriptionTwo major challenges in the field of bioinformatics are multiple sequence alignment (MSA) and abundance profiling. Both are hard problems that can become computationally prohibitive with a large amount of input data, which most existing methods can fail to handle. In addition, only a few approaches can adequately align input sequences with various lengths (i.e., sequence length heterogeneity), which are often present in metagenomic data used for profiling species abundances. With the increasing availability and size of sequence data that may have sequence length heterogeneity, new approaches are required to perform scalable and accurate analyses. In this thesis, we show our progress in developing new alignment methods that deal with sequence length heterogeneity and can scale to large data using divide-and-conquer. We also apply the new alignment methods to perform abundance profiling and demonstrate improved accuracy. We hope that the newer methods can help scientists conduct more accurate and effective analyses on larger and larger datasets.
Degree
thesis:*- Name thesis:degree_name
- Ph.D.
- Level thesis:degree_level
- Dissertation
- Discipline thesis:degree_discipline
- Computer Science
- Grantor
- University of Illinois Urbana-Champaign
- Year dc:date
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Shen, Chengze
- Contributors dc:contributor
-
- Warnow, Tandy
- El-Kebir, Mohammed
- Gropp, William D
- Williams, Kelly P.
Subjects
dc:subject × 2Rights
dc:rights- Statement dc:rights
-
- Copyright 2025 Chengze Shen
- Language dc:language
- en, eng
Identifiers
dc:identifier.*- Handle dc:identifier
- https://hdl.handle.net/2142/129918