{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/97339"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/97339","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Information theoretic and machine learning techniques for emerging genomic data analysis","abstract":"\"The completion of the Human Genome Project in 2003 opened a new era for scientists. Through advanced high-throughput sequencing technologies, we now have access to a large amount of genomic data and we can use it to answer key biological questions, such as the factors contributing to the development of cancer. Large data sets and rapidly advancing sequencing technology pose challenges for processing and storing large volumes of genomic data. Moreover, the analysis of datasets may be both computationally and theoretically challenging because statistical methods have not been developed for new emerging data. In this work, I address some of these problems using tools from information theory and machine learning. First, I focus on the data processing and storage aspect of metagenomics, the study of microbial communities in environmental samples and human organs. In particular, I introduce MetaCRAM, the first software suite specialized for metagenomic sequencing data processing and compression, and demonstrate that MetaCRAM compresses data to 2-13 percent of the original file size. Second, I analyze a biological dataset assaying the propensity of a DNA sequence to form a four-stranded structure called \"\"G-quadruplex\"\" (GQ). GQ structures have been proposed to regulate diverse key biological processes including transcription, replication, and translation. I present main factors that lead to GQ formation, and propose highly accurate linear regression and Gaussian process regression models to predict the ability of a DNA sequence to fold into GQ. Third, I study data structures to analyze and store three-dimensional chromatin conformation data generated from high-throughput sequencing technologies. In particular, I examine statistical properties of Hi-C contact maps and propose a few suitable formats to encode pairwise interactions between genome locations.\"","abstract_html":"&quot;The completion of the Human Genome Project in 2003 opened a new era for scientists. Through advanced high-throughput sequencing technologies, we now have access to a large amount of genomic data and we can use it to answer key biological questions, such as the factors contributing to the development of cancer. Large data sets and rapidly advancing sequencing technology pose challenges for processing and storing large volumes of genomic data. Moreover, the analysis of datasets may be both computationally and theoretically challenging because statistical methods have not been developed for new emerging data. In this work, I address some of these problems using tools from information theory and machine learning. First, I focus on the data processing and storage aspect of metagenomics, the study of microbial communities in environmental samples and human organs. In particular, I introduce MetaCRAM, the first software suite specialized for metagenomic sequencing data processing and compression, and demonstrate that MetaCRAM compresses data to 2-13 percent of the original file size. Second, I analyze a biological dataset assaying the propensity of a DNA sequence to form a four-stranded structure called &quot;&quot;G-quadruplex&quot;&quot; (GQ). GQ structures have been proposed to regulate diverse key biological processes including transcription, replication, and translation. I present main factors that lead to GQ formation, and propose highly accurate linear regression and Gaussian process regression models to predict the ability of a DNA sequence to fold into GQ. Third, I study data structures to analyze and store three-dimensional chromatin conformation data generated from high-throughput sequencing technologies. In particular, I examine statistical properties of Hi-C contact maps and propose a few suitable formats to encode pairwise interactions between genome locations.&quot;","abstract_has_math":false,"creators":["Kim, Minji"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Milenkovic, Olgica","Song, Jun S.","Veeravalli, Venugopal V.","Sinha, Saurabh","Peng, Jian"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2017,"date_issued":"2017-08-10T19:14:56Z","date_published":"2017-08-10T19:14:56Z","updated_at":"2026-07-22T22:24:32Z","subjects":["Genomic compression","DNA folding"],"languages":["en"],"rights":["Copyright 2017 Minji Kim"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/97339","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Milenkovic, Olgica","Song, Jun S.","Veeravalli, Venugopal V.","Sinha, Saurabh","Peng, Jian"]},{"key":"dc:creator","label":"Author","values":["Kim, Minji"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2017-08-10T19:14:56Z","2017-04-13","2017-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Genomic compression","DNA folding"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2017 Minji Kim"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/97339"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["\"The completion of the Human Genome Project in 2003 opened a new era for scientists. Through advanced high-throughput sequencing technologies, we now have access to a large amount of genomic data and we can use it to answer key biological questions, such as the factors contributing to the development of cancer. Large data sets and rapidly advancing sequencing technology pose challenges for processing and storing large volumes of genomic data. Moreover, the analysis of datasets may be both computationally and theoretically challenging because statistical methods have not been developed for new emerging data. In this work, I address some of these problems using tools from information theory and machine learning. First, I focus on the data processing and storage aspect of metagenomics, the study of microbial communities in environmental samples and human organs. In particular, I introduce MetaCRAM, the first software suite specialized for metagenomic sequencing data processing and compression, and demonstrate that MetaCRAM compresses data to 2-13 percent of the original file size. Second, I analyze a biological dataset assaying the propensity of a DNA sequence to form a four-stranded structure called \"\"G-quadruplex\"\" (GQ). GQ structures have been proposed to regulate diverse key biological processes including transcription, replication, and translation. I present main factors that lead to GQ formation, and propose highly accurate linear regression and Gaussian process regression models to predict the ability of a DNA sequence to fold into GQ. Third, I study data structures to analyze and store three-dimensional chromatin conformation data generated from high-throughput sequencing technologies. In particular, I examine statistical properties of Hi-C contact maps and propose a few suitable formats to encode pairwise interactions between genome locations.\"","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2017-08-10 without embargo terms","The student, Minji Kim, accepted the attached license on 2017-04-12 at 13:41.","The student, Minji Kim, submitted this Dissertation for approval on 2017-04-12 at 13:49.","This Dissertation was approved for publication on 2017-04-13 at 10:38.","DSpace SAF Submission Ingestion Package generated from Vireo submission #10723 on 2017-08-10 at 13:39:26","Made available in DSpace on 2017-08-10T19:14:56Z (GMT). No. of bitstreams: 3 KIM-DISSERTATION-2017.pdf: 27437790 bytes, checksum: dc762ec5b3d8bc7a6b0fc1bf1a782037 (MD5) LICENSE.txt: 4206 bytes, checksum: 20f033fa06f2f4d0e71754553fe7d37d (MD5) PROQUEST_LICENSE.txt: 4552 bytes, checksum: d15f6ede6cc6f17e02550827a8e171f0 (MD5) Previous issue date: 2017-04-13"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Information theoretic and machine learning techniques for emerging genomic data analysis"]}]}],"canonical_facts":{"dc:contributor":["Milenkovic, Olgica","Song, Jun S.","Veeravalli, Venugopal V.","Sinha, Saurabh","Peng, Jian"],"dc:creator":["Kim, Minji"],"dc:date":["2017-08-10T19:14:56Z","2017-04-13","2017-05"],"dc:description":["\"The completion of the Human Genome Project in 2003 opened a new era for scientists. Through advanced high-throughput sequencing technologies, we now have access to a large amount of genomic data and we can use it to answer key biological questions, such as the factors contributing to the development of cancer. Large data sets and rapidly advancing sequencing technology pose challenges for processing and storing large volumes of genomic data. Moreover, the analysis of datasets may be both computationally and theoretically challenging because statistical methods have not been developed for new emerging data. In this work, I address some of these problems using tools from information theory and machine learning. First, I focus on the data processing and storage aspect of metagenomics, the study of microbial communities in environmental samples and human organs. In particular, I introduce MetaCRAM, the first software suite specialized for metagenomic sequencing data processing and compression, and demonstrate that MetaCRAM compresses data to 2-13 percent of the original file size. Second, I analyze a biological dataset assaying the propensity of a DNA sequence to form a four-stranded structure called \"\"G-quadruplex\"\" (GQ). GQ structures have been proposed to regulate diverse key biological processes including transcription, replication, and translation. I present main factors that lead to GQ formation, and propose highly accurate linear regression and Gaussian process regression models to predict the ability of a DNA sequence to fold into GQ. Third, I study data structures to analyze and store three-dimensional chromatin conformation data generated from high-throughput sequencing technologies. In particular, I examine statistical properties of Hi-C contact maps and propose a few suitable formats to encode pairwise interactions between genome locations.\"","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2017-08-10 without embargo terms","The student, Minji Kim, accepted the attached license on 2017-04-12 at 13:41.","The student, Minji Kim, submitted this Dissertation for approval on 2017-04-12 at 13:49.","This Dissertation was approved for publication on 2017-04-13 at 10:38.","DSpace SAF Submission Ingestion Package generated from Vireo submission #10723 on 2017-08-10 at 13:39:26","Made available in DSpace on 2017-08-10T19:14:56Z (GMT). No. of bitstreams: 3 KIM-DISSERTATION-2017.pdf: 27437790 bytes, checksum: dc762ec5b3d8bc7a6b0fc1bf1a782037 (MD5) LICENSE.txt: 4206 bytes, checksum: 20f033fa06f2f4d0e71754553fe7d37d (MD5) PROQUEST_LICENSE.txt: 4552 bytes, checksum: d15f6ede6cc6f17e02550827a8e171f0 (MD5) Previous issue date: 2017-04-13"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/97339"],"dc:language":["en"],"dc:rights":["Copyright 2017 Minji Kim"],"dc:subject":["Genomic compression","DNA folding"],"dc:title":["Information theoretic and machine learning techniques for emerging genomic data analysis"],"dc:type":["text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:32Z"}