Back to results

University of Illinois at Urbana-Champaign

Understanding information at the biomolecular level using statistics and machine learning

Abstract

dc:description

The Central Dogma of molecular biology states that DNA is transcribed into RNA, which is then translated into protein. The majority of cellular functions can trace their origins to the various stages of the Central Dogma, making it a central pillar of our understanding of biological systems. However, this description of molecular biology only touches the surface of understanding biological function; the information encoded in these biomolecules and the way this information is processed is also crucial to understanding biological phenomena. With the recent development of genome editing tools and next-generation sequencing methods, researchers finally possess the necessary means to measure and even control information on the genomic, transcriptomic, and proteomic levels. Furthermore, the accumulation of large datasets with new modalities of data have opened up opportunities to develop new methods of biological data analysis based on machine learning. This thesis documents our efforts to develop and utilize statistical and machine learning techniques to analyze genetic engineering techniques and leverage data from novel next-generation sequencing assays to shed new light on previously studied biological phenomena, including the origin of keratinocyte cancers and immune system recognition of pathogens. First, we investigated the use of genome editing tools to engineer transcription. We performed statistical analysis on RNA sequence data and genomic DNA sequence data to show successful exclusion of exons from RNA transcripts after modifying the splice signal using CRISPR base editors as well as quantify base editing rates of DNA on-target and off-target sites. Additionally, we performed genome-wide prediction of editability of exons and developed a web interface to facilitate use of the technology. Second, we investigated the cell of origin of two cancers, Squamous Cell Carcinoma and Basal Cell Carcinoma, by developing a similarity metric between bulk RNA expression levels of cancer and single-cell RNA sequence data of keratinocytes in various stages of their differentiation process. Third, we used a convolutional neural network model inspired by concepts from deep metric learning and multimodal learning to predict binding between T-Cell receptors (TCRs) and antigen epitopes with accuracy comparable to, or greater than, the state-of-the-art. We used a neural network interpretation method to identify positions in the TCR important for binding and utilized crystal structure data to show that proximity of the TCR to the epitope may not be a good proxy of importance in determining epitope specificity.

Degree

thesis:*
Name thesis:degree_name
Ph.D.
Level thesis:degree_level
Dissertation
Discipline thesis:degree_discipline
Physics
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2022

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Luu, Alan M.
Contributors dc:contributor
  • Song, Jun
  • Maslov, Sergei
  • Golding, Ido
  • Perez-Pinera, Pablo

Subjects

dc:subject × 29

Rights

dc:rights
Statement dc:rights
  • Copyright 2022 Alan Luu
Language dc:language
en, eng

Identifiers

dc:identifier.*
Handle dc:identifier
https://hdl.handle.net/2142/115452

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Luu, Alan M.. Understanding information at the biomolecular level using statistics and machine learning. Dissertation thesis, University of Illinois at Urbana-Champaign, 2022. https://hdl.handle.net/2142/115452