University of Toronto
Novel Machine Learning Models for Identification, Characterization and Prioritization of Phenotype-Genotype Associations
Abstract
dc:description.abstractThis thesis is centered on the broad problem of inferring, characterizing and prioritizing molecular and functional phenotype effects of genetic variants. I first investigate the problems caused by phenotype heterogeneity in association studies of complex traits. Many complex diseases, such as autism, diabetes and cancer, are very heterogeneous in terms of the realized phenotypes, i.e. individuals diagnosed with the same disease might experience substantially different symptoms. I introduce the JBASE: Joint Bayesian Analysis of Subphenotyping and Epistasis model, which simultaneously accounts for non-additive interactions and phenotype heterogeneity while searching for associations. Next, inspired by the hypothesis that some phenotypes can be better explained by additive (small) effects of a large number of variants, I investigate the applicability of sparse supervised learning models for association studies of complex traits. In particular, I investigate integrative machine learning models that can leverage additional biological knowledge about phenotypes and genetic variants while searching for genetically informed associations. Towards this goal I present AAALasso: Accounting for Antagonistic Associations in Lasso, a novel sparse additive multi-task learning model that takes network structured information between markers and phenotypes into account. While identifying variants associated with a phenotype is important, being able to validate them by establishing causal biomolecular links is equally important. As such, in the last part of this thesis, I focus on the problem of predicting stability effects of non-synonymous coding mutations as a tool for the prioritization of variants discovered in association mapping studies. I introduce ELASPIC: Ensemble Learning Approach for Stability Prediction of Interface and Core mutations, a supervised ensemble learning approach to accurately predict effects of non-synonymous single nucleotide polymorphisms (SNPs) on protein stability and, for the first time, on protein-protein binding affinity. Overall, this thesis proposes three new complementary machine learning models for inferring, characterizing and prioritizing the phenotype effects of genetic variants. The first two approaches propose novel tools for finding underlying variants of a given phenotype, whereas the third approach focuses on predicting the molecular effects of the discovered variants.
Degree
thesis:*- Department dc:contributor.department
- Computer Science
- Year dc:date.issued
- 2015
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Colak, Recep
- Advisors dc:contributor.advisor
-
- Kim, Philip M
- Brudno, Michael
Subjects
dc:subject × 6Identifiers
dc:identifier.*- Handle dc:identifier.uri
- http://hdl.handle.net/1807/70835
- OAI identifier oai:identifier
- oai:utoronto.scholaris.ca:1807/70835