{"id":{"repo_id":"mit","oai_identifier":"oai:dspace.mit.edu:1721.1/82392"},"canonical_url":"https://search.dev.ndltd.org/etd/mit/oai:dspace.mit.edu:1721.1/82392","repository":{"repo_id":"mit","name":"MIT","base_url":"https://dspace.mit.edu/oai/request"},"display":{"title":"Characterizing non-coding hits in genome-wide association studies using epigenetic data","abstract":"Understanding the molecular basis of human disease is one of the greatest challenges of our time, and recent explosion in genetic and genomic datasets are finally putting it within reach. In the last ten years, genome-wide association studies have identified thousands of genetic variants associated with disease. However, the majority of these variants fall outside genes making interpreting their role in disease difficult. In parallel, the ENCODE and Roadmap Epigenomics consortia have produced high resolution annotations of the genome which identify large portions with potential regulatory function. We develop methods to interpret genome-wide association studies using these annotations to generate hypotheses about how associated variants contribute to disease mechanism. In particular, we go beyond the usual stringent p-value threshold to investigate variants with small individual effect sizes which current methods do not have power to detect. Evaluating our methods on the Wellcome Trust Case Control Consortium 7 Disease studies, we find associated variants are enriched in a variety of functional categories even after controlling for various biases. We also find an unprecedented number of variants contribute to this enrichment, supporting our hypothesis that the architecture of these diseases involves combinatorial interaction of many variants with small individual effect sizes.","abstract_html":"Understanding the molecular basis of human disease is one of the greatest challenges of our time, and recent explosion in genetic and genomic datasets are finally putting it within reach. In the last ten years, genome-wide association studies have identified thousands of genetic variants associated with disease. However, the majority of these variants fall outside genes making interpreting their role in disease difficult. In parallel, the ENCODE and Roadmap Epigenomics consortia have produced high resolution annotations of the genome which identify large portions with potential regulatory function. We develop methods to interpret genome-wide association studies using these annotations to generate hypotheses about how associated variants contribute to disease mechanism. In particular, we go beyond the usual stringent p-value threshold to investigate variants with small individual effect sizes which current methods do not have power to detect. Evaluating our methods on the Wellcome Trust Case Control Consortium 7 Disease studies, we find associated variants are enriched in a variety of functional categories even after controlling for various biases. We also find an unprecedented number of variants contribute to this enrichment, supporting our hypothesis that the architecture of these diseases involves combinatorial interaction of many variants with small individual effect sizes.","abstract_has_math":false,"creators":["Sarkar, Abhishek Kulshreshtha"],"institution":"Massachusetts Institute of Technology","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":"Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science.","school":null,"contributors":[],"advisors":["Manolis Kellis."],"committee_chairs":[],"committee_members":[],"year":2013,"date_issued":"2013","date_published":"2013","updated_at":"2026-07-22T22:21:17Z","subjects":["Electrical Engineering and Computer Science."],"languages":["eng"],"rights":["M.I.T. theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. See provided URL for inquiries about permission."],"rights_urls":["http://dspace.mit.edu/handle/1721.1/7582"],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/1721.1/82392","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Manolis Kellis."]},{"key":"dc:contributor.department","label":"Department","values":["Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science."]},{"key":"dc:contributor.other","label":"Dc Contributor Other","values":["Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science."]},{"key":"dc:creator","label":"Author","values":["Sarkar, Abhishek Kulshreshtha"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2013-11-18T19:17:31Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2013-11-18T19:17:31Z"]},{"key":"dc:date.issued","label":"Date","values":["2013"]},{"key":"dc:publisher","label":"Institution","values":["Massachusetts Institute of Technology"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Electrical Engineering and Computer Science."]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["M.I.T. theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. See provided URL for inquiries about permission."]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://dspace.mit.edu/handle/1721.1/7582"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["http://hdl.handle.net/1721.1/82392"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Thesis (S.M.)--Massachusetts Institute of Technology, Dept. of Electrical Engineering and Computer Science, 2013.","Cataloged from PDF version of thesis.","Includes bibliographical references (p. 42-46)."]},{"key":"dc:description.abstract","label":"Abstract","values":["Understanding the molecular basis of human disease is one of the greatest challenges of our time, and recent explosion in genetic and genomic datasets are finally putting it within reach. In the last ten years, genome-wide association studies have identified thousands of genetic variants associated with disease. However, the majority of these variants fall outside genes making interpreting their role in disease difficult. In parallel, the ENCODE and Roadmap Epigenomics consortia have produced high resolution annotations of the genome which identify large portions with potential regulatory function. We develop methods to interpret genome-wide association studies using these annotations to generate hypotheses about how associated variants contribute to disease mechanism. In particular, we go beyond the usual stringent p-value threshold to investigate variants with small individual effect sizes which current methods do not have power to detect. Evaluating our methods on the Wellcome Trust Case Control Consortium 7 Disease studies, we find associated variants are enriched in a variety of functional categories even after controlling for various biases. We also find an unprecedented number of variants contribute to this enrichment, supporting our hypothesis that the architecture of these diseases involves combinatorial interaction of many variants with small individual effect sizes."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["S.M."]},{"key":"dc:title","label":"Title","values":["Characterizing non-coding hits in genome-wide association studies using epigenetic data"]}]}],"canonical_facts":{"dc:contributor.advisor":["Manolis Kellis."],"dc:contributor.department":["Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science."],"dc:contributor.other":["Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science."],"dc:creator":["Sarkar, Abhishek Kulshreshtha"],"dc:date.accessioned":["2013-11-18T19:17:31Z"],"dc:date.available":["2013-11-18T19:17:31Z"],"dc:date.issued":["2013"],"dc:description":["Thesis (S.M.)--Massachusetts Institute of Technology, Dept. of Electrical Engineering and Computer Science, 2013.","Cataloged from PDF version of thesis.","Includes bibliographical references (p. 42-46)."],"dc:description.abstract":["Understanding the molecular basis of human disease is one of the greatest challenges of our time, and recent explosion in genetic and genomic datasets are finally putting it within reach. In the last ten years, genome-wide association studies have identified thousands of genetic variants associated with disease. However, the majority of these variants fall outside genes making interpreting their role in disease difficult. In parallel, the ENCODE and Roadmap Epigenomics consortia have produced high resolution annotations of the genome which identify large portions with potential regulatory function. We develop methods to interpret genome-wide association studies using these annotations to generate hypotheses about how associated variants contribute to disease mechanism. In particular, we go beyond the usual stringent p-value threshold to investigate variants with small individual effect sizes which current methods do not have power to detect. Evaluating our methods on the Wellcome Trust Case Control Consortium 7 Disease studies, we find associated variants are enriched in a variety of functional categories even after controlling for various biases. We also find an unprecedented number of variants contribute to this enrichment, supporting our hypothesis that the architecture of these diseases involves combinatorial interaction of many variants with small individual effect sizes."],"dc:description.degree":["S.M."],"dc:identifier.uri":["http://hdl.handle.net/1721.1/82392"],"dc:language.iso":["eng"],"dc:publisher":["Massachusetts Institute of Technology"],"dc:rights":["M.I.T. theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. See provided URL for inquiries about permission."],"dc:rights.uri":["http://dspace.mit.edu/handle/1721.1/7582"],"dc:subject":["Electrical Engineering and Computer Science."],"dc:title":["Characterizing non-coding hits in genome-wide association studies using epigenetic data"],"dc:type":["Thesis"]},"updated_at":"2026-07-22T22:21:17Z"}