{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/113071"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/113071","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Algorithms for infection and cancer genomics","abstract":"The student, Palash Sashittal, submitted this Thesis for approval on 2021-07-19 at 11:33.","abstract_html":"The student, Palash Sashittal, submitted this Thesis for approval on 2021-07-19 at 11:33.","abstract_has_math":false,"creators":["Sashittal, Palash"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["El-Kebir, Mohammed"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-01-12T21:46:55Z","date_published":"2022-01-12T21:46:55Z","updated_at":"2026-07-22T22:24:53Z","subjects":["Infection genomics","Cancer genomics","Combinatorial optimization","Transcript assembly","Doublet detection"],"languages":["en"],"rights":["Copyright 2021 Palash Sashittal"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/113071","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["El-Kebir, Mohammed"]},{"key":"dc:creator","label":"Author","values":["Sashittal, Palash"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022-01-12T21:46:55Z","2021-07-19","2021-08"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Infection genomics","Cancer genomics","Combinatorial optimization","Transcript assembly","Doublet detection"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2021 Palash Sashittal"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/113071"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The student, Palash Sashittal, submitted this Thesis for approval on 2021-07-19 at 11:33.","Continuous innovations and advances in sequencing technologies have led to the birth and development of several fields of research. In this thesis we propose four methods to address open problems in two such fields, infection genomics and cancer genomics. The first problem we address is reconstruction of transmission history of an outbreak using genomic and epidemiological data collected from infected hosts. It is challenging to account for all the relevant biological processes that occur during evolution and transmission of the pathogens in the outbreak while also addressing the uncertainty in the most likely solution. Our method, TiTUS, overcomes these challenges by first uniformly sampling from the set of all possible feasible transmission histories of the outbreak under a realistic model of evolution and transmission. Then, a consensus-based solution is generated that summarizes the candidate solutions in a biologically meaningful way. We show that TiTUS efficiently samples the solution space enabling accurate reconstruction of transmission history of an outbreak. The second method we introduce, Jumper, reconstructs viral transcripts using RNA-sequencing data from infected cells. In this study, we focus our attention on viruses in the Coronaviridae family, such as SARS-CoV-2, that express genes by a process of discontinuous transcription mediated by the viral RNA-dependent RNA polymerase. The viral transcriptome provides valuable information with clinical implications such as differential expression of viral genes, the host cell response to viral infection and the viral life cycle. We show that Jumper accurately infers the viral transcripts, outperforming existing transcript assembly methods, and facilitates the study of coronavirus transcriptomes under varying conditions. The third problem we address is doublet detection in single-cell DNA-sequencing data. Our method, doubletD, is the first stand-alone doublet detection method for single-cell DNA-sequencing data. We use a simple probabilistic model allowing a closed-form maximum likelihood solution that efficiently and accurately detects doublets by identifying characteristic signal in the variant allele frequency (VAF) distribution in the data. On simulations and multiple real datasets, we show that doublet identification and removal using doubletD improves downstream analysis such as genotype calling and phylogeny reconstruction. Finally, we present a new method, PACTION, which proposes a solution to the tumor phylogeny inference problem in cancer. Due to technological and methodological limitations, existing methods are restricted to identifying tumor clones and phylogenies only based on either small-scale mutations, such as single nucleotide variations (SNVs), or large-scale mutations, such as copy number aberrations (CNAs), preventing a comprehensive characterization of a tumor’s clonal composition. To overcome these challenges, we formulate the identification of clones in terms of both SNVs and CNAs as a reconciliation problem. We show that PACTION reliably identifies tumor clones and their evolutionary relationships even in the presence of noise or error in input SNVs and CNAs.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-01-12 without embargo terms","The student, Palash Sashittal, accepted the attached license on 2021-07-19 at 11:25.","This Thesis was approved for publication on 2021-07-19 at 15:17.","DSpace SAF Submission Ingestion Package generated from Vireo submission #17007 on 2022-01-12 at 12:46:15","Made available in DSpace on 2022-01-12T21:46:55Z (GMT). No. of bitstreams: 2 SASHITTAL-THESIS-2021.pdf: 9144068 bytes, checksum: 0adc91169cdfd824434a8f55193c2997 (MD5) LICENSE.txt: 4213 bytes, checksum: 9dbc004765206a45fd3294a17a302e9b (MD5) Previous issue date: 2021-07-19"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Algorithms for infection and cancer genomics"]}]}],"canonical_facts":{"dc:contributor":["El-Kebir, Mohammed"],"dc:creator":["Sashittal, Palash"],"dc:date":["2022-01-12T21:46:55Z","2021-07-19","2021-08"],"dc:description":["The student, Palash Sashittal, submitted this Thesis for approval on 2021-07-19 at 11:33.","Continuous innovations and advances in sequencing technologies have led to the birth and development of several fields of research. In this thesis we propose four methods to address open problems in two such fields, infection genomics and cancer genomics. The first problem we address is reconstruction of transmission history of an outbreak using genomic and epidemiological data collected from infected hosts. It is challenging to account for all the relevant biological processes that occur during evolution and transmission of the pathogens in the outbreak while also addressing the uncertainty in the most likely solution. Our method, TiTUS, overcomes these challenges by first uniformly sampling from the set of all possible feasible transmission histories of the outbreak under a realistic model of evolution and transmission. Then, a consensus-based solution is generated that summarizes the candidate solutions in a biologically meaningful way. We show that TiTUS efficiently samples the solution space enabling accurate reconstruction of transmission history of an outbreak. The second method we introduce, Jumper, reconstructs viral transcripts using RNA-sequencing data from infected cells. In this study, we focus our attention on viruses in the Coronaviridae family, such as SARS-CoV-2, that express genes by a process of discontinuous transcription mediated by the viral RNA-dependent RNA polymerase. The viral transcriptome provides valuable information with clinical implications such as differential expression of viral genes, the host cell response to viral infection and the viral life cycle. We show that Jumper accurately infers the viral transcripts, outperforming existing transcript assembly methods, and facilitates the study of coronavirus transcriptomes under varying conditions. The third problem we address is doublet detection in single-cell DNA-sequencing data. Our method, doubletD, is the first stand-alone doublet detection method for single-cell DNA-sequencing data. We use a simple probabilistic model allowing a closed-form maximum likelihood solution that efficiently and accurately detects doublets by identifying characteristic signal in the variant allele frequency (VAF) distribution in the data. On simulations and multiple real datasets, we show that doublet identification and removal using doubletD improves downstream analysis such as genotype calling and phylogeny reconstruction. Finally, we present a new method, PACTION, which proposes a solution to the tumor phylogeny inference problem in cancer. Due to technological and methodological limitations, existing methods are restricted to identifying tumor clones and phylogenies only based on either small-scale mutations, such as single nucleotide variations (SNVs), or large-scale mutations, such as copy number aberrations (CNAs), preventing a comprehensive characterization of a tumor’s clonal composition. To overcome these challenges, we formulate the identification of clones in terms of both SNVs and CNAs as a reconciliation problem. We show that PACTION reliably identifies tumor clones and their evolutionary relationships even in the presence of noise or error in input SNVs and CNAs.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-01-12 without embargo terms","The student, Palash Sashittal, accepted the attached license on 2021-07-19 at 11:25.","This Thesis was approved for publication on 2021-07-19 at 15:17.","DSpace SAF Submission Ingestion Package generated from Vireo submission #17007 on 2022-01-12 at 12:46:15","Made available in DSpace on 2022-01-12T21:46:55Z (GMT). No. of bitstreams: 2 SASHITTAL-THESIS-2021.pdf: 9144068 bytes, checksum: 0adc91169cdfd824434a8f55193c2997 (MD5) LICENSE.txt: 4213 bytes, checksum: 9dbc004765206a45fd3294a17a302e9b (MD5) Previous issue date: 2021-07-19"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/113071"],"dc:language":["en"],"dc:rights":["Copyright 2021 Palash Sashittal"],"dc:subject":["Infection genomics","Cancer genomics","Combinatorial optimization","Transcript assembly","Doublet detection"],"dc:title":["Algorithms for infection and cancer genomics"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:53Z"}