{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/104838"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/104838","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Entropy-based machine learning algorithms applied to genomics and pattern recognition","abstract":"Transcription factors (TF) are proteins that interact with DNA to regulate the transcription of DNA to RNA and play key roles in both healthy and cancerous cells. Thus, gaining a deeper understanding of the biological factors underlying transcription factor (TF) binding specificity is important for understanding the mechanism of oncogenesis. As large, biological datasets become more readily available, machine learning (ML) algorithms have proven to make up an important and useful set of tools for cancer researchers. However, there remain many areas for potential improvements for these ML models, including a higher degree of model interpretability and overall accuracy. In this thesis, we present decision tree (DT) methods applied to DNA sequence analysis that result in highly interpretable and accurate predictions. We propose a boosted decision tree (BDT) model using the binary counts of important DNA motifs to predict the binding specificity of TFs belonging to the same protein family of binding similar DNA sequences. We then proceed to introduce a novel application of Convolutional Decision Trees (CDT) and demonstrate that this approach has distinct advantages over the BDT modeil while still accurately predicting the binding specificty of TFs. The CDT models are trained using the Cross Entropy (CE) optimization method, a Monte Carlo optimization method based on concepts from information theory related to statistical mechanics. We then further study the CDT model as a general pattern recognition and transfer learning technique and demonstrate that this approach can learn translationally invariant patterns that lead to high classification accuracy while remaining more interpretable and learning higher quality convolutional filters compared to convolutional neural networks (CNN).","abstract_html":"Transcription factors (TF) are proteins that interact with DNA to regulate the transcription of DNA to RNA and play key roles in both healthy and cancerous cells. Thus, gaining a deeper understanding of the biological factors underlying transcription factor (TF) binding specificity is important for understanding the mechanism of oncogenesis. As large, biological datasets become more readily available, machine learning (ML) algorithms have proven to make up an important and useful set of tools for cancer researchers. However, there remain many areas for potential improvements for these ML models, including a higher degree of model interpretability and overall accuracy. In this thesis, we present decision tree (DT) methods applied to DNA sequence analysis that result in highly interpretable and accurate predictions. We propose a boosted decision tree (BDT) model using the binary counts of important DNA motifs to predict the binding specificity of TFs belonging to the same protein family of binding similar DNA sequences. We then proceed to introduce a novel application of Convolutional Decision Trees (CDT) and demonstrate that this approach has distinct advantages over the BDT modeil while still accurately predicting the binding specificty of TFs. The CDT models are trained using the Cross Entropy (CE) optimization method, a Monte Carlo optimization method based on concepts from information theory related to statistical mechanics. We then further study the CDT model as a general pattern recognition and transfer learning technique and demonstrate that this approach can learn translationally invariant patterns that lead to high classification accuracy while remaining more interpretable and learning higher quality convolutional filters compared to convolutional neural networks (CNN).","abstract_has_math":false,"creators":["Moon, Wooyoung"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Physics","degree_department":null,"school":null,"contributors":["Song, Jun S.","Dahmen, Karin","Kuehn, Seppe","Draper, Patrick"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019-08-23T19:51:55Z","date_published":"2019-08-23T19:51:55Z","updated_at":"2026-07-22T22:24:42Z","subjects":["Machine Learning, Decision Trees, Convolutional Filters, Genomics, Cancer, Entropy"],"languages":["en"],"rights":["Copyright 2019 Wooyoung Moon"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/104838","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Song, Jun S.","Dahmen, Karin","Kuehn, Seppe","Draper, Patrick"]},{"key":"dc:creator","label":"Author","values":["Moon, Wooyoung"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2019-08-23T19:51:55Z","2019-04-16","2019-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Physics"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Machine Learning, Decision Trees, Convolutional Filters, Genomics, Cancer, Entropy"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2019 Wooyoung Moon"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/104838"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Transcription factors (TF) are proteins that interact with DNA to regulate the transcription of DNA to RNA and play key roles in both healthy and cancerous cells. Thus, gaining a deeper understanding of the biological factors underlying transcription factor (TF) binding specificity is important for understanding the mechanism of oncogenesis. As large, biological datasets become more readily available, machine learning (ML) algorithms have proven to make up an important and useful set of tools for cancer researchers. However, there remain many areas for potential improvements for these ML models, including a higher degree of model interpretability and overall accuracy. In this thesis, we present decision tree (DT) methods applied to DNA sequence analysis that result in highly interpretable and accurate predictions. We propose a boosted decision tree (BDT) model using the binary counts of important DNA motifs to predict the binding specificity of TFs belonging to the same protein family of binding similar DNA sequences. We then proceed to introduce a novel application of Convolutional Decision Trees (CDT) and demonstrate that this approach has distinct advantages over the BDT modeil while still accurately predicting the binding specificty of TFs. The CDT models are trained using the Cross Entropy (CE) optimization method, a Monte Carlo optimization method based on concepts from information theory related to statistical mechanics. We then further study the CDT model as a general pattern recognition and transfer learning technique and demonstrate that this approach can learn translationally invariant patterns that lead to high classification accuracy while remaining more interpretable and learning higher quality convolutional filters compared to convolutional neural networks (CNN).","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2019-08-22 without embargo terms","The student, Wooyoung Moon, accepted the attached license on 2019-04-15 at 18:54.","The student, Wooyoung Moon, submitted this Dissertation for approval on 2019-04-15 at 19:00.","This Dissertation was approved for publication on 2019-04-16 at 17:53.","DSpace SAF Submission Ingestion Package generated from Vireo submission #13655 on 2019-08-22 at 14:43:52","Made available in DSpace on 2019-08-23T19:51:55Z (GMT). No. of bitstreams: 3 MOON-DISSERTATION-2019.pdf: 4259693 bytes, checksum: 078218171943952e581a43c42bc8ba85 (MD5) LICENSE.txt: 4210 bytes, checksum: 730fb4bfce44456ab60430ca9f755c3a (MD5) PROQUEST_LICENSE.txt: 4556 bytes, checksum: 09f885b94e30e363398310eb50ad9628 (MD5) Previous issue date: 2019-04-16"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Entropy-based machine learning algorithms applied to genomics and pattern recognition"]}]}],"canonical_facts":{"dc:contributor":["Song, Jun S.","Dahmen, Karin","Kuehn, Seppe","Draper, Patrick"],"dc:creator":["Moon, Wooyoung"],"dc:date":["2019-08-23T19:51:55Z","2019-04-16","2019-05"],"dc:description":["Transcription factors (TF) are proteins that interact with DNA to regulate the transcription of DNA to RNA and play key roles in both healthy and cancerous cells. Thus, gaining a deeper understanding of the biological factors underlying transcription factor (TF) binding specificity is important for understanding the mechanism of oncogenesis. As large, biological datasets become more readily available, machine learning (ML) algorithms have proven to make up an important and useful set of tools for cancer researchers. However, there remain many areas for potential improvements for these ML models, including a higher degree of model interpretability and overall accuracy. In this thesis, we present decision tree (DT) methods applied to DNA sequence analysis that result in highly interpretable and accurate predictions. We propose a boosted decision tree (BDT) model using the binary counts of important DNA motifs to predict the binding specificity of TFs belonging to the same protein family of binding similar DNA sequences. We then proceed to introduce a novel application of Convolutional Decision Trees (CDT) and demonstrate that this approach has distinct advantages over the BDT modeil while still accurately predicting the binding specificty of TFs. The CDT models are trained using the Cross Entropy (CE) optimization method, a Monte Carlo optimization method based on concepts from information theory related to statistical mechanics. We then further study the CDT model as a general pattern recognition and transfer learning technique and demonstrate that this approach can learn translationally invariant patterns that lead to high classification accuracy while remaining more interpretable and learning higher quality convolutional filters compared to convolutional neural networks (CNN).","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2019-08-22 without embargo terms","The student, Wooyoung Moon, accepted the attached license on 2019-04-15 at 18:54.","The student, Wooyoung Moon, submitted this Dissertation for approval on 2019-04-15 at 19:00.","This Dissertation was approved for publication on 2019-04-16 at 17:53.","DSpace SAF Submission Ingestion Package generated from Vireo submission #13655 on 2019-08-22 at 14:43:52","Made available in DSpace on 2019-08-23T19:51:55Z (GMT). No. of bitstreams: 3 MOON-DISSERTATION-2019.pdf: 4259693 bytes, checksum: 078218171943952e581a43c42bc8ba85 (MD5) LICENSE.txt: 4210 bytes, checksum: 730fb4bfce44456ab60430ca9f755c3a (MD5) PROQUEST_LICENSE.txt: 4556 bytes, checksum: 09f885b94e30e363398310eb50ad9628 (MD5) Previous issue date: 2019-04-16"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/104838"],"dc:language":["en"],"dc:rights":["Copyright 2019 Wooyoung Moon"],"dc:subject":["Machine Learning, Decision Trees, Convolutional Filters, Genomics, Cancer, Entropy"],"dc:title":["Entropy-based machine learning algorithms applied to genomics and pattern recognition"],"dc:type":["text"],"thesis:degree_discipline":["Physics"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:42Z"}