{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/108034"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/108034","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Methods to summarize and reduce the solution space of tumor phylogeny inference","abstract":"Cancer phylogenies are key to studying tumorigenesis and have clinical implications. Due to the heterogeneous nature of cancer and limitations in current sequencing technology, current cancer phylogeny inference methods identify a large solution space of plausible phylogenies. To facilitate further downstream analysis, one can either summarize the set of cancer phylogenies or use additional data to eliminate trees and further reduce the solution space. Current summary methods are limited to a single consensus tree or graph and may miss important topological features that are present in different subsets of candidate trees. On the other hand, while single-cell sequencing (SCS) provides the data that we need to reduce solution space, it may become prohibitively costly as the number of cells to sequence increases. In this thesis, we first introduce the Multiple Consensus Tree (MCT) problem to simultaneously cluster the trees in the solution space and infer a consensus tree for each cluster. We show that MCT is NP-hard, and present an exact algorithm based on mixed integer linear programming (MILP) and a heuristic algorithm that efficiently identifies high-quality consensus trees. We demonstrate the applicability of our methods on both simulated and real data, showing that our approach selects the number of clusters depending on the complexity of the solution space. Next, we introduce PhyDOSE, a method that uses bulk sequencing data to strategically optimize the design of follow-up single-cell sequencing experiments. We incorporate distinguishing features - features that uniquely identify a tree - into a probabilistic model that infers the number of cells to sequence so as to confidently reconstruct the phylogeny of the tumor. We validate PhyDOSE using simulations and a retrospective analysis of a childhood leukemia patient, concluding that PhyDOSE's computed number of cells resolves tree ambiguity even in the presence of typical single-cell sequencing errors. We also conduct a retrospective analysis on an acute myeloid leukemia cohort, demonstrating the potential of significant reduction in the number of cells to sequence. In a prospective analysis, we demonstrate that only a small number of cells suffice to disambiguate the solution space of trees in a recent lung cancer cohort. Finally, we provide an R package and web interface for the ease of use of PhyDOSE.","abstract_html":"Cancer phylogenies are key to studying tumorigenesis and have clinical implications. Due to the heterogeneous nature of cancer and limitations in current sequencing technology, current cancer phylogeny inference methods identify a large solution space of plausible phylogenies. To facilitate further downstream analysis, one can either summarize the set of cancer phylogenies or use additional data to eliminate trees and further reduce the solution space. Current summary methods are limited to a single consensus tree or graph and may miss important topological features that are present in different subsets of candidate trees. On the other hand, while single-cell sequencing (SCS) provides the data that we need to reduce solution space, it may become prohibitively costly as the number of cells to sequence increases. In this thesis, we first introduce the Multiple Consensus Tree (MCT) problem to simultaneously cluster the trees in the solution space and infer a consensus tree for each cluster. We show that MCT is NP-hard, and present an exact algorithm based on mixed integer linear programming (MILP) and a heuristic algorithm that efficiently identifies high-quality consensus trees. We demonstrate the applicability of our methods on both simulated and real data, showing that our approach selects the number of clusters depending on the complexity of the solution space. Next, we introduce PhyDOSE, a method that uses bulk sequencing data to strategically optimize the design of follow-up single-cell sequencing experiments. We incorporate distinguishing features - features that uniquely identify a tree - into a probabilistic model that infers the number of cells to sequence so as to confidently reconstruct the phylogeny of the tumor. We validate PhyDOSE using simulations and a retrospective analysis of a childhood leukemia patient, concluding that PhyDOSE&#x27;s computed number of cells resolves tree ambiguity even in the presence of typical single-cell sequencing errors. We also conduct a retrospective analysis on an acute myeloid leukemia cohort, demonstrating the potential of significant reduction in the number of cells to sequence. In a prospective analysis, we demonstrate that only a small number of cells suffice to disambiguate the solution space of trees in a recent lung cancer cohort. Finally, we provide an R package and web interface for the ease of use of PhyDOSE.","abstract_has_math":false,"creators":["Aguse, Nuraini Binti"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["El-Kebir, Mohammed"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2020,"date_issued":"2020-08-26T21:58:02Z","date_published":"2020-08-26T21:58:02Z","updated_at":"2026-07-22T22:24:47Z","subjects":["Tumor Phylogeny","Summary","Single-cell sequencing"],"languages":["en"],"rights":["Copyright 2020 Nuraini Aguse"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/108034","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["El-Kebir, Mohammed"]},{"key":"dc:creator","label":"Author","values":["Aguse, Nuraini Binti"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2020-08-26T21:58:02Z","2020-05-12","2020-05"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Tumor Phylogeny","Summary","Single-cell sequencing"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2020 Nuraini Aguse"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/108034"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Cancer phylogenies are key to studying tumorigenesis and have clinical implications. Due to the heterogeneous nature of cancer and limitations in current sequencing technology, current cancer phylogeny inference methods identify a large solution space of plausible phylogenies. To facilitate further downstream analysis, one can either summarize the set of cancer phylogenies or use additional data to eliminate trees and further reduce the solution space. Current summary methods are limited to a single consensus tree or graph and may miss important topological features that are present in different subsets of candidate trees. On the other hand, while single-cell sequencing (SCS) provides the data that we need to reduce solution space, it may become prohibitively costly as the number of cells to sequence increases. In this thesis, we first introduce the Multiple Consensus Tree (MCT) problem to simultaneously cluster the trees in the solution space and infer a consensus tree for each cluster. We show that MCT is NP-hard, and present an exact algorithm based on mixed integer linear programming (MILP) and a heuristic algorithm that efficiently identifies high-quality consensus trees. We demonstrate the applicability of our methods on both simulated and real data, showing that our approach selects the number of clusters depending on the complexity of the solution space. Next, we introduce PhyDOSE, a method that uses bulk sequencing data to strategically optimize the design of follow-up single-cell sequencing experiments. We incorporate distinguishing features - features that uniquely identify a tree - into a probabilistic model that infers the number of cells to sequence so as to confidently reconstruct the phylogeny of the tumor. We validate PhyDOSE using simulations and a retrospective analysis of a childhood leukemia patient, concluding that PhyDOSE's computed number of cells resolves tree ambiguity even in the presence of typical single-cell sequencing errors. We also conduct a retrospective analysis on an acute myeloid leukemia cohort, demonstrating the potential of significant reduction in the number of cells to sequence. In a prospective analysis, we demonstrate that only a small number of cells suffice to disambiguate the solution space of trees in a recent lung cancer cohort. Finally, we provide an R package and web interface for the ease of use of PhyDOSE.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2020-08-25 without embargo terms","The student, Nuraini Aguse, accepted the attached license on 2020-05-11 at 13:45.","The student, Nuraini Aguse, submitted this Thesis for approval on 2020-05-11 at 14:00.","This Thesis was approved for publication on 2020-05-12 at 15:29.","DSpace SAF Submission Ingestion Package generated from Vireo submission #15325 on 2020-08-25 at 17:13:56","Made available in DSpace on 2020-08-26T21:58:02Z (GMT). No. of bitstreams: 2 AGUSE-THESIS-2020.pdf: 2951988 bytes, checksum: dbfb5edc9e24d22809d6747aabaebbcf (MD5) LICENSE.txt: 4210 bytes, checksum: 2eeb9267e4580d6924330809aa4a7dc5 (MD5) Previous issue date: 2020-05-12"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Methods to summarize and reduce the solution space of tumor phylogeny inference"]}]}],"canonical_facts":{"dc:contributor":["El-Kebir, Mohammed"],"dc:creator":["Aguse, Nuraini Binti"],"dc:date":["2020-08-26T21:58:02Z","2020-05-12","2020-05"],"dc:description":["Cancer phylogenies are key to studying tumorigenesis and have clinical implications. Due to the heterogeneous nature of cancer and limitations in current sequencing technology, current cancer phylogeny inference methods identify a large solution space of plausible phylogenies. To facilitate further downstream analysis, one can either summarize the set of cancer phylogenies or use additional data to eliminate trees and further reduce the solution space. Current summary methods are limited to a single consensus tree or graph and may miss important topological features that are present in different subsets of candidate trees. On the other hand, while single-cell sequencing (SCS) provides the data that we need to reduce solution space, it may become prohibitively costly as the number of cells to sequence increases. In this thesis, we first introduce the Multiple Consensus Tree (MCT) problem to simultaneously cluster the trees in the solution space and infer a consensus tree for each cluster. We show that MCT is NP-hard, and present an exact algorithm based on mixed integer linear programming (MILP) and a heuristic algorithm that efficiently identifies high-quality consensus trees. We demonstrate the applicability of our methods on both simulated and real data, showing that our approach selects the number of clusters depending on the complexity of the solution space. Next, we introduce PhyDOSE, a method that uses bulk sequencing data to strategically optimize the design of follow-up single-cell sequencing experiments. We incorporate distinguishing features - features that uniquely identify a tree - into a probabilistic model that infers the number of cells to sequence so as to confidently reconstruct the phylogeny of the tumor. We validate PhyDOSE using simulations and a retrospective analysis of a childhood leukemia patient, concluding that PhyDOSE's computed number of cells resolves tree ambiguity even in the presence of typical single-cell sequencing errors. We also conduct a retrospective analysis on an acute myeloid leukemia cohort, demonstrating the potential of significant reduction in the number of cells to sequence. In a prospective analysis, we demonstrate that only a small number of cells suffice to disambiguate the solution space of trees in a recent lung cancer cohort. Finally, we provide an R package and web interface for the ease of use of PhyDOSE.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2020-08-25 without embargo terms","The student, Nuraini Aguse, accepted the attached license on 2020-05-11 at 13:45.","The student, Nuraini Aguse, submitted this Thesis for approval on 2020-05-11 at 14:00.","This Thesis was approved for publication on 2020-05-12 at 15:29.","DSpace SAF Submission Ingestion Package generated from Vireo submission #15325 on 2020-08-25 at 17:13:56","Made available in DSpace on 2020-08-26T21:58:02Z (GMT). No. of bitstreams: 2 AGUSE-THESIS-2020.pdf: 2951988 bytes, checksum: dbfb5edc9e24d22809d6747aabaebbcf (MD5) LICENSE.txt: 4210 bytes, checksum: 2eeb9267e4580d6924330809aa4a7dc5 (MD5) Previous issue date: 2020-05-12"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/108034"],"dc:language":["en"],"dc:rights":["Copyright 2020 Nuraini Aguse"],"dc:subject":["Tumor Phylogeny","Summary","Single-cell sequencing"],"dc:title":["Methods to summarize and reduce the solution space of tumor phylogeny inference"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:47Z"}