{"id":{"repo_id":"uthsc","oai_identifier":"oai:digitalcommons.library.tmc.edu:utgsbs_dissertations-2262"},"canonical_url":"https://search.dev.ndltd.org/etd/uthsc/oai:digitalcommons.library.tmc.edu:utgsbs_dissertations-2262","repository":{"repo_id":"uthsc","name":"University of Texas Health Science Center at Houston","base_url":"https://digitalcommons.library.tmc.edu/do/oai/"},"display":{"title":"Development of Graphical Models and Statistical Physics Motivated Approaches to Genomic Investigations","abstract":"<p>Identifying genes involved in disease pathology has been a goal of genomic research since the early days of the field. However, as technology improves and the body of research grows, we are faced with more questions than answers. Among these is the pressing matter of our incomplete understanding of the genetic underpinnings of complex diseases. Many hypotheses offer explanations as to why direct and independent analyses of variants, as done in genome-wide association studies (GWAS), may not fully elucidate disease genetics. These range from pointing out flaws in statistical testing to invoking the complex dynamics of epigenetic processes. In the studies outlined here, however, we focus on the hypothesis that interactions between genes may be a potential culprit. To probe this hypothesis, we begin by developing an algorithm, GeneEMBED, to model the total effect of protein coding variants in various genes across a molecular network of genetic interactions. Given a population of disease and healthy individuals, GeneEMBED systematically evaluates the relative contribution of a gene to disease. The associations are quantified by examining the patterns of differential perturbations in the gene's interactions throughout a biological network. As a proof-of-concept, we applied GeneEMBED to two late-onset Alzheimer's disease (AD) cohorts of 5,169 exomes and 969 genomes. We identified 143 candidate disease-associated genes across the two cohorts and three biological networks. These candidate genes were differentially expressed in both bulk and single-cell RNA expression data from post-mortem AD brains. Knockouts of these candidates in mice were known to lead to abnormal neurological phenotypes. Lastly, in vivo drosophila assays of candidates showed they modified neurodegenerative phenotypes. Next, we focus on the discrepancies between the functional impact of mutations across different genes. While tools to predict the degree of functional impact a given coding mutation will have on the encoded protein are widely successful, they often make predictions relative to the given gene. To this effect, we extend principles of statistical mechanics to biology to measure any given gene's relative mutational intolerance. Importantly, these mutational intolerance scores can distinguish essential genes from non-essential genes in E.coli. In humans, they can segregate genes that cause autosomal dominant Mendelian diseases from non-disease genes. Similarly, highly mutationally intolerant genes were enriched in core and conserved biological processes across three different species. Conversely, mutationally tolerant genes were involved in adaptive processes, again across three different species. Most notably, we found that mutational intolerance scores highly correlated with experimentally measured fitness effects of gene knockdowns. Together, these efforts provide new tools with which to investigate disease-gene associations and provide insights into the biological dynamics of gene networks.</p>","abstract_html":"&lt;p&gt;Identifying genes involved in disease pathology has been a goal of genomic research since the early days of the field. However, as technology improves and the body of research grows, we are faced with more questions than answers. Among these is the pressing matter of our incomplete understanding of the genetic underpinnings of complex diseases. Many hypotheses offer explanations as to why direct and independent analyses of variants, as done in genome-wide association studies (GWAS), may not fully elucidate disease genetics. These range from pointing out flaws in statistical testing to invoking the complex dynamics of epigenetic processes. In the studies outlined here, however, we focus on the hypothesis that interactions between genes may be a potential culprit. To probe this hypothesis, we begin by developing an algorithm, GeneEMBED, to model the total effect of protein coding variants in various genes across a molecular network of genetic interactions. Given a population of disease and healthy individuals, GeneEMBED systematically evaluates the relative contribution of a gene to disease. The associations are quantified by examining the patterns of differential perturbations in the gene&#x27;s interactions throughout a biological network. As a proof-of-concept, we applied GeneEMBED to two late-onset Alzheimer&#x27;s disease (AD) cohorts of 5,169 exomes and 969 genomes. We identified 143 candidate disease-associated genes across the two cohorts and three biological networks. These candidate genes were differentially expressed in both bulk and single-cell RNA expression data from post-mortem AD brains. Knockouts of these candidates in mice were known to lead to abnormal neurological phenotypes. Lastly, in vivo drosophila assays of candidates showed they modified neurodegenerative phenotypes. Next, we focus on the discrepancies between the functional impact of mutations across different genes. While tools to predict the degree of functional impact a given coding mutation will have on the encoded protein are widely successful, they often make predictions relative to the given gene. To this effect, we extend principles of statistical mechanics to biology to measure any given gene&#x27;s relative mutational intolerance. Importantly, these mutational intolerance scores can distinguish essential genes from non-essential genes in E.coli. In humans, they can segregate genes that cause autosomal dominant Mendelian diseases from non-disease genes. Similarly, highly mutationally intolerant genes were enriched in core and conserved biological processes across three different species. Conversely, mutationally tolerant genes were involved in adaptive processes, again across three different species. Most notably, we found that mutational intolerance scores highly correlated with experimentally measured fitness effects of gene knockdowns. Together, these efforts provide new tools with which to investigate disease-gene associations and provide insights into the biological dynamics of gene networks.&lt;/p&gt;","abstract_has_math":false,"creators":["Lagisetty, Yashwanth","<p>0000-0001-7364-4630</p>"],"institution":null,"degree_name":"Doctor of Philosophy (PhD)","degree_level":"Dissertation (PhD)","degree_discipline":null,"degree_department":null,"school":null,"contributors":["Edgar T. Walters - Advisory Professor","Olivier Lichtarge - Advisory Professor","Prahlad Ram"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-08-01T07:00:00Z","date_published":"2022-08-01T07:00:00Z","updated_at":"2026-07-24T05:50:31Z","subjects":["Machine Learning","Statistical Mechanics","Thermodynamics","Genomics","Evolution","Alzheimer's Disease","Artificial Intelligence and Robotics","Computational Biology","Geometry and Topology","Nervous System Diseases","Statistical, Nonlinear, and Soft Matter Physics"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://digitalcommons.library.tmc.edu/utgsbs_dissertations/1205","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Edgar T. Walters - Advisory Professor","Olivier Lichtarge - Advisory Professor","Prahlad Ram"]},{"key":"dc:creator","label":"Author","values":["Lagisetty, Yashwanth","<p>0000-0001-7364-4630</p>"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2023-08-01T07:00:00Z"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation (PhD)"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctor of Philosophy (PhD)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Machine Learning","Statistical Mechanics","Thermodynamics","Genomics","Evolution","Alzheimer's Disease","Artificial Intelligence and Robotics","Computational Biology","Geometry and Topology","Nervous System Diseases","Statistical, Nonlinear, and Soft Matter Physics"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://digitalcommons.library.tmc.edu/utgsbs_dissertations/1205"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>Identifying genes involved in disease pathology has been a goal of genomic research since the early days of the field. However, as technology improves and the body of research grows, we are faced with more questions than answers. Among these is the pressing matter of our incomplete understanding of the genetic underpinnings of complex diseases. Many hypotheses offer explanations as to why direct and independent analyses of variants, as done in genome-wide association studies (GWAS), may not fully elucidate disease genetics. These range from pointing out flaws in statistical testing to invoking the complex dynamics of epigenetic processes. In the studies outlined here, however, we focus on the hypothesis that interactions between genes may be a potential culprit. To probe this hypothesis, we begin by developing an algorithm, GeneEMBED, to model the total effect of protein coding variants in various genes across a molecular network of genetic interactions. Given a population of disease and healthy individuals, GeneEMBED systematically evaluates the relative contribution of a gene to disease. The associations are quantified by examining the patterns of differential perturbations in the gene's interactions throughout a biological network. As a proof-of-concept, we applied GeneEMBED to two late-onset Alzheimer's disease (AD) cohorts of 5,169 exomes and 969 genomes. We identified 143 candidate disease-associated genes across the two cohorts and three biological networks. These candidate genes were differentially expressed in both bulk and single-cell RNA expression data from post-mortem AD brains. Knockouts of these candidates in mice were known to lead to abnormal neurological phenotypes. Lastly, in vivo drosophila assays of candidates showed they modified neurodegenerative phenotypes. Next, we focus on the discrepancies between the functional impact of mutations across different genes. While tools to predict the degree of functional impact a given coding mutation will have on the encoded protein are widely successful, they often make predictions relative to the given gene. To this effect, we extend principles of statistical mechanics to biology to measure any given gene's relative mutational intolerance. Importantly, these mutational intolerance scores can distinguish essential genes from non-essential genes in E.coli. In humans, they can segregate genes that cause autosomal dominant Mendelian diseases from non-disease genes. Similarly, highly mutationally intolerant genes were enriched in core and conserved biological processes across three different species. Conversely, mutationally tolerant genes were involved in adaptive processes, again across three different species. Most notably, we found that mutational intolerance scores highly correlated with experimentally measured fitness effects of gene knockdowns. Together, these efforts provide new tools with which to investigate disease-gene associations and provide insights into the biological dynamics of gene networks.</p>"]},{"key":"dc:title","label":"Title","values":["Development of Graphical Models and Statistical Physics Motivated Approaches to Genomic Investigations"]}]}],"canonical_facts":{"dc:contributor":["Edgar T. Walters - Advisory Professor","Olivier Lichtarge - Advisory Professor","Prahlad Ram"],"dc:creator":["Lagisetty, Yashwanth","<p>0000-0001-7364-4630</p>"],"dc:date.available":["2023-08-01T07:00:00Z"],"dc:description.abstract":["<p>Identifying genes involved in disease pathology has been a goal of genomic research since the early days of the field. However, as technology improves and the body of research grows, we are faced with more questions than answers. Among these is the pressing matter of our incomplete understanding of the genetic underpinnings of complex diseases. Many hypotheses offer explanations as to why direct and independent analyses of variants, as done in genome-wide association studies (GWAS), may not fully elucidate disease genetics. These range from pointing out flaws in statistical testing to invoking the complex dynamics of epigenetic processes. In the studies outlined here, however, we focus on the hypothesis that interactions between genes may be a potential culprit. To probe this hypothesis, we begin by developing an algorithm, GeneEMBED, to model the total effect of protein coding variants in various genes across a molecular network of genetic interactions. Given a population of disease and healthy individuals, GeneEMBED systematically evaluates the relative contribution of a gene to disease. The associations are quantified by examining the patterns of differential perturbations in the gene's interactions throughout a biological network. As a proof-of-concept, we applied GeneEMBED to two late-onset Alzheimer's disease (AD) cohorts of 5,169 exomes and 969 genomes. We identified 143 candidate disease-associated genes across the two cohorts and three biological networks. These candidate genes were differentially expressed in both bulk and single-cell RNA expression data from post-mortem AD brains. Knockouts of these candidates in mice were known to lead to abnormal neurological phenotypes. Lastly, in vivo drosophila assays of candidates showed they modified neurodegenerative phenotypes. Next, we focus on the discrepancies between the functional impact of mutations across different genes. While tools to predict the degree of functional impact a given coding mutation will have on the encoded protein are widely successful, they often make predictions relative to the given gene. To this effect, we extend principles of statistical mechanics to biology to measure any given gene's relative mutational intolerance. Importantly, these mutational intolerance scores can distinguish essential genes from non-essential genes in E.coli. In humans, they can segregate genes that cause autosomal dominant Mendelian diseases from non-disease genes. Similarly, highly mutationally intolerant genes were enriched in core and conserved biological processes across three different species. Conversely, mutationally tolerant genes were involved in adaptive processes, again across three different species. Most notably, we found that mutational intolerance scores highly correlated with experimentally measured fitness effects of gene knockdowns. Together, these efforts provide new tools with which to investigate disease-gene associations and provide insights into the biological dynamics of gene networks.</p>"],"dc:identifier":["https://digitalcommons.library.tmc.edu/utgsbs_dissertations/1205"],"dc:subject":["Machine Learning","Statistical Mechanics","Thermodynamics","Genomics","Evolution","Alzheimer's Disease","Artificial Intelligence and Robotics","Computational Biology","Geometry and Topology","Nervous System Diseases","Statistical, Nonlinear, and Soft Matter Physics"],"dc:title":["Development of Graphical Models and Statistical Physics Motivated Approaches to Genomic Investigations"],"thesis:degree_level":["Dissertation (PhD)"],"thesis:degree_name":["Doctor of Philosophy (PhD)"]},"updated_at":"2026-07-24T05:50:31Z"}