{"id":{"repo_id":"ohiolink","oai_identifier":"oai:etd.ohiolink.edu:case1343769483"},"canonical_url":"https://search.dev.ndltd.org/etd/ohiolink/oai:etd.ohiolink.edu:case1343769483","repository":{"repo_id":"ohiolink","name":"OhioLINK","base_url":"https://etd.ohiolink.edu/acprod/odb_etd/ws/oai/oai"},"display":{"title":"Algorithms for discovering disease genes by integrating 'omics data","abstract":"<p>Systems-level characterization of complex human diseases remains as one of the biggest challenges in the post-genomic era. Information useful for mechanistic understanding of diseases comes from different types of “-omic&rdquo; data, including genomic sequences, gene expression, and molecular interactions. Genome Wide Association Studies (GWAS) compare genomic sequences from healthy and affected populations to identify genetic variants that are potentially associated with diseases. Monitoring of gene expression, on the other hand, enables identification of genes that are dysregulated in the development and progression of diseases. However, since complex diseases arise from the interplay among multiple interacting factors, analyses of individual variants in isolation provide limited insights. To this end, data on molecular networks, including protein-protein interactions (PPI), provide a useful resource for uncovering the disease association of multiple molecules in the context of their biological function and interactions. In this thesis, we develop algorithms that integrate different -omic data types to provide systems-level insights into complex diseases.</p><p>We first address the problem of identifying genetic interactions among multiple functionally related variants. For this purpose, we develop algorithms to identify groups of single nucleotide polymorphisms (SNPs) that are (i) associated with the same gene and (ii) exhibit more significant association with the disease when considered together. In order to achieve this, we represent the “genotype&rdquo; of a gene as a combination of a subset of SNPs within its region of interest and develop algorithms to identify the subset of SNPs that best describes the genotypic variation in the patient population. Subsequently, we focus on the problem of disease gene prioritization. We propose a novel algorithm, VAVIEN, that utilizes the topological similarity of proteins in the human PPI network to prioritize candidate disease genes that reside in linkage intervals potentially associated with the disease. Finally, we incorporate mRNA expression data into our studies and propose a set-cover based algorithm, referred as COBALT, that identifies class-specific, coordinately dysregulated subnetworks of genes, associated with the phenotype of interest. We show with comprehensive experimental studies that the proposed algorithms are very effective in generating novel insights into the systems biology of complex diseases.</p>","abstract_html":"&lt;p&gt;Systems-level characterization of complex human diseases remains as one of the biggest challenges in the post-genomic era. Information useful for mechanistic understanding of diseases comes from different types of “-omic&amp;rdquo; data, including genomic sequences, gene expression, and molecular interactions. Genome Wide Association Studies (GWAS) compare genomic sequences from healthy and affected populations to identify genetic variants that are potentially associated with diseases. Monitoring of gene expression, on the other hand, enables identification of genes that are dysregulated in the development and progression of diseases. However, since complex diseases arise from the interplay among multiple interacting factors, analyses of individual variants in isolation provide limited insights. To this end, data on molecular networks, including protein-protein interactions (PPI), provide a useful resource for uncovering the disease association of multiple molecules in the context of their biological function and interactions. In this thesis, we develop algorithms that integrate different -omic data types to provide systems-level insights into complex diseases.&lt;/p&gt;&lt;p&gt;We first address the problem of identifying genetic interactions among multiple functionally related variants. For this purpose, we develop algorithms to identify groups of single nucleotide polymorphisms (SNPs) that are (i) associated with the same gene and (ii) exhibit more significant association with the disease when considered together. In order to achieve this, we represent the “genotype&amp;rdquo; of a gene as a combination of a subset of SNPs within its region of interest and develop algorithms to identify the subset of SNPs that best describes the genotypic variation in the patient population. Subsequently, we focus on the problem of disease gene prioritization. We propose a novel algorithm, VAVIEN, that utilizes the topological similarity of proteins in the human PPI network to prioritize candidate disease genes that reside in linkage intervals potentially associated with the disease. Finally, we incorporate mRNA expression data into our studies and propose a set-cover based algorithm, referred as COBALT, that identifies class-specific, coordinately dysregulated subnetworks of genes, associated with the phenotype of interest. We show with comprehensive experimental studies that the proposed algorithms are very effective in generating novel insights into the systems biology of complex diseases.&lt;/p&gt;","abstract_has_math":false,"creators":["Erten, Mehmet Sinan"],"institution":"Case Western Reserve University School of Graduate Studies","degree_name":"Doctor of Philosophy","degree_level":"doctoral","degree_discipline":"EECS - Computer and Information Sciences","degree_department":null,"school":null,"contributors":["Koyuturk, Mehmet"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2013,"date_issued":"2013-03-07","date_published":"2013-03-07","updated_at":"2026-07-24T03:35:52Z","subjects":["Bioinformatics","Computer Science","GWAS","summary statistics","case-control studies","network-based disease gene prioritization","coordinately dysregulated subnetworks"],"languages":["English"],"rights":["unrestricted","This thesis or dissertation is protected by copyright: all rights reserved. It may not be copied or redistributed beyond the terms of applicable copyright laws."],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://rave.ohiolink.edu/etdc/view?acc_num=case1343769483","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Koyuturk, Mehmet"]},{"key":"dc:creator","label":"Author","values":["Erten, Mehmet Sinan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2013-03-07"]},{"key":"dc:publisher","label":"Institution","values":["Case Western Reserve University School of Graduate Studies / OhioLINK"]},{"key":"dc:type","label":"Dc Type","values":["Electronic Thesis or Dissertation"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["EECS - Computer and Information Sciences"]},{"key":"thesis:degree_level","label":"Degree Level","values":["doctoral"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctor of Philosophy"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Case Western Reserve University School of Graduate Studies"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Bioinformatics","Computer Science","GWAS","summary statistics","case-control studies","network-based disease gene prioritization","coordinately dysregulated subnetworks"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["English"]},{"key":"dc:rights","label":"Dc Rights","values":["unrestricted","This thesis or dissertation is protected by copyright: all rights reserved. It may not be copied or redistributed beyond the terms of applicable copyright laws."]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://rave.ohiolink.edu/etdc/view?acc_num=case1343769483"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["<p>Systems-level characterization of complex human diseases remains as one of the biggest challenges in the post-genomic era. Information useful for mechanistic understanding of diseases comes from different types of “-omic&rdquo; data, including genomic sequences, gene expression, and molecular interactions. Genome Wide Association Studies (GWAS) compare genomic sequences from healthy and affected populations to identify genetic variants that are potentially associated with diseases. Monitoring of gene expression, on the other hand, enables identification of genes that are dysregulated in the development and progression of diseases. However, since complex diseases arise from the interplay among multiple interacting factors, analyses of individual variants in isolation provide limited insights. To this end, data on molecular networks, including protein-protein interactions (PPI), provide a useful resource for uncovering the disease association of multiple molecules in the context of their biological function and interactions. In this thesis, we develop algorithms that integrate different -omic data types to provide systems-level insights into complex diseases.</p><p>We first address the problem of identifying genetic interactions among multiple functionally related variants. For this purpose, we develop algorithms to identify groups of single nucleotide polymorphisms (SNPs) that are (i) associated with the same gene and (ii) exhibit more significant association with the disease when considered together. In order to achieve this, we represent the “genotype&rdquo; of a gene as a combination of a subset of SNPs within its region of interest and develop algorithms to identify the subset of SNPs that best describes the genotypic variation in the patient population. Subsequently, we focus on the problem of disease gene prioritization. We propose a novel algorithm, VAVIEN, that utilizes the topological similarity of proteins in the human PPI network to prioritize candidate disease genes that reside in linkage intervals potentially associated with the disease. Finally, we incorporate mRNA expression data into our studies and propose a set-cover based algorithm, referred as COBALT, that identifies class-specific, coordinately dysregulated subnetworks of genes, associated with the phenotype of interest. We show with comprehensive experimental studies that the proposed algorithms are very effective in generating novel insights into the systems biology of complex diseases.</p>"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf","p.108","1023.9 KB"]},{"key":"dc:title","label":"Title","values":["Algorithms for discovering disease genes by integrating 'omics data"]}]}],"canonical_facts":{"dc:contributor":["Koyuturk, Mehmet"],"dc:creator":["Erten, Mehmet Sinan"],"dc:date":["2013-03-07"],"dc:description":["<p>Systems-level characterization of complex human diseases remains as one of the biggest challenges in the post-genomic era. Information useful for mechanistic understanding of diseases comes from different types of “-omic&rdquo; data, including genomic sequences, gene expression, and molecular interactions. Genome Wide Association Studies (GWAS) compare genomic sequences from healthy and affected populations to identify genetic variants that are potentially associated with diseases. Monitoring of gene expression, on the other hand, enables identification of genes that are dysregulated in the development and progression of diseases. However, since complex diseases arise from the interplay among multiple interacting factors, analyses of individual variants in isolation provide limited insights. To this end, data on molecular networks, including protein-protein interactions (PPI), provide a useful resource for uncovering the disease association of multiple molecules in the context of their biological function and interactions. In this thesis, we develop algorithms that integrate different -omic data types to provide systems-level insights into complex diseases.</p><p>We first address the problem of identifying genetic interactions among multiple functionally related variants. For this purpose, we develop algorithms to identify groups of single nucleotide polymorphisms (SNPs) that are (i) associated with the same gene and (ii) exhibit more significant association with the disease when considered together. In order to achieve this, we represent the “genotype&rdquo; of a gene as a combination of a subset of SNPs within its region of interest and develop algorithms to identify the subset of SNPs that best describes the genotypic variation in the patient population. Subsequently, we focus on the problem of disease gene prioritization. We propose a novel algorithm, VAVIEN, that utilizes the topological similarity of proteins in the human PPI network to prioritize candidate disease genes that reside in linkage intervals potentially associated with the disease. Finally, we incorporate mRNA expression data into our studies and propose a set-cover based algorithm, referred as COBALT, that identifies class-specific, coordinately dysregulated subnetworks of genes, associated with the phenotype of interest. We show with comprehensive experimental studies that the proposed algorithms are very effective in generating novel insights into the systems biology of complex diseases.</p>"],"dc:format":["application/pdf","p.108","1023.9 KB"],"dc:identifier":["http://rave.ohiolink.edu/etdc/view?acc_num=case1343769483"],"dc:language":["English"],"dc:publisher":["Case Western Reserve University School of Graduate Studies / OhioLINK"],"dc:rights":["unrestricted","This thesis or dissertation is protected by copyright: all rights reserved. It may not be copied or redistributed beyond the terms of applicable copyright laws."],"dc:subject":["Bioinformatics","Computer Science","GWAS","summary statistics","case-control studies","network-based disease gene prioritization","coordinately dysregulated subnetworks"],"dc:title":["Algorithms for discovering disease genes by integrating 'omics data"],"dc:type":["Electronic Thesis or Dissertation"],"thesis:degree_discipline":["EECS - Computer and Information Sciences"],"thesis:degree_level":["doctoral"],"thesis:degree_name":["Doctor of Philosophy"],"thesis:institution_name":["Case Western Reserve University School of Graduate Studies"]},"updated_at":"2026-07-24T03:35:52Z"}