{"id":{"repo_id":"mit","oai_identifier":"oai:dspace.mit.edu:1721.1/85492"},"canonical_url":"https://search.dev.ndltd.org/etd/mit/oai:dspace.mit.edu:1721.1/85492","repository":{"repo_id":"mit","name":"MIT","base_url":"https://dspace.mit.edu/oai/request"},"display":{"title":"Characterizing and predicting enhancers in the human genome","abstract":"Characterizing the functions of sequences in the human genome is crucial for the study and treatment of human disease. Though it is known that approximately 5% of the human genome is conserved, about 40% of these sequences have yet to be characterized, many of which may be important players in human disease pathways (1). Experimental and computational techniques have been developed which use histone modifications to segment the human genome into 25 different chromatin states, including states corresponding to various functional sequences like promoters and enhancers (4). However, the availability of this data is very limited, as these assays have been performed on a limited number of cell types, and the distribution of chromatin states varies across different cell types. We therefore took a computational rather than experimental approach to discovering regulatory regions. We characterized the nucleotide contents, regulatory motif contents, conservation, gene distance, and human variation patterns of a subset of these regulatory sequences. By training a generalized linear classifier on this data, we created a predictor for enhancer sequences that achieved 70% accuracy.","abstract_html":"Characterizing the functions of sequences in the human genome is crucial for the study and treatment of human disease. Though it is known that approximately 5% of the human genome is conserved, about 40% of these sequences have yet to be characterized, many of which may be important players in human disease pathways (1). Experimental and computational techniques have been developed which use histone modifications to segment the human genome into 25 different chromatin states, including states corresponding to various functional sequences like promoters and enhancers (4). However, the availability of this data is very limited, as these assays have been performed on a limited number of cell types, and the distribution of chromatin states varies across different cell types. We therefore took a computational rather than experimental approach to discovering regulatory regions. We characterized the nucleotide contents, regulatory motif contents, conservation, gene distance, and human variation patterns of a subset of these regulatory sequences. By training a generalized linear classifier on this data, we created a predictor for enhancer sequences that achieved 70% accuracy.","abstract_has_math":false,"creators":["Roytman, Megan (Megan D.)"],"institution":"Massachusetts Institute of Technology","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":"Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science.","school":null,"contributors":[],"advisors":["Matthew L. Eaton."],"committee_chairs":[],"committee_members":[],"year":2013,"date_issued":"2013","date_published":"2013","updated_at":"2026-07-22T22:21:44Z","subjects":["Electrical Engineering and Computer Science."],"languages":["eng"],"rights":["M.I.T. theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. See provided URL for inquiries about permission."],"rights_urls":["http://dspace.mit.edu/handle/1721.1/7582"],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/1721.1/85492","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Matthew L. Eaton."]},{"key":"dc:contributor.department","label":"Department","values":["Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science."]},{"key":"dc:contributor.other","label":"Dc Contributor Other","values":["Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science."]},{"key":"dc:creator","label":"Author","values":["Roytman, Megan (Megan D.)"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2014-03-06T15:45:56Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2014-03-06T15:45:56Z"]},{"key":"dc:date.issued","label":"Date","values":["2013"]},{"key":"dc:publisher","label":"Institution","values":["Massachusetts Institute of Technology"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Electrical Engineering and Computer Science."]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["M.I.T. theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. See provided URL for inquiries about permission."]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://dspace.mit.edu/handle/1721.1/7582"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["http://hdl.handle.net/1721.1/85492"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Thesis: M. Eng., Massachusetts Institute of Technology, Department of Electrical Engineering and Computer Science, 2013.","Cataloged from PDF version of thesis.","Includes bibliographical references (page 31)."]},{"key":"dc:description.abstract","label":"Abstract","values":["Characterizing the functions of sequences in the human genome is crucial for the study and treatment of human disease. Though it is known that approximately 5% of the human genome is conserved, about 40% of these sequences have yet to be characterized, many of which may be important players in human disease pathways (1). Experimental and computational techniques have been developed which use histone modifications to segment the human genome into 25 different chromatin states, including states corresponding to various functional sequences like promoters and enhancers (4). However, the availability of this data is very limited, as these assays have been performed on a limited number of cell types, and the distribution of chromatin states varies across different cell types. We therefore took a computational rather than experimental approach to discovering regulatory regions. We characterized the nucleotide contents, regulatory motif contents, conservation, gene distance, and human variation patterns of a subset of these regulatory sequences. By training a generalized linear classifier on this data, we created a predictor for enhancer sequences that achieved 70% accuracy."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["M. Eng."]},{"key":"dc:title","label":"Title","values":["Characterizing and predicting enhancers in the human genome"]}]}],"canonical_facts":{"dc:contributor.advisor":["Matthew L. Eaton."],"dc:contributor.department":["Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science."],"dc:contributor.other":["Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science."],"dc:creator":["Roytman, Megan (Megan D.)"],"dc:date.accessioned":["2014-03-06T15:45:56Z"],"dc:date.available":["2014-03-06T15:45:56Z"],"dc:date.issued":["2013"],"dc:description":["Thesis: M. Eng., Massachusetts Institute of Technology, Department of Electrical Engineering and Computer Science, 2013.","Cataloged from PDF version of thesis.","Includes bibliographical references (page 31)."],"dc:description.abstract":["Characterizing the functions of sequences in the human genome is crucial for the study and treatment of human disease. Though it is known that approximately 5% of the human genome is conserved, about 40% of these sequences have yet to be characterized, many of which may be important players in human disease pathways (1). Experimental and computational techniques have been developed which use histone modifications to segment the human genome into 25 different chromatin states, including states corresponding to various functional sequences like promoters and enhancers (4). However, the availability of this data is very limited, as these assays have been performed on a limited number of cell types, and the distribution of chromatin states varies across different cell types. We therefore took a computational rather than experimental approach to discovering regulatory regions. We characterized the nucleotide contents, regulatory motif contents, conservation, gene distance, and human variation patterns of a subset of these regulatory sequences. By training a generalized linear classifier on this data, we created a predictor for enhancer sequences that achieved 70% accuracy."],"dc:description.degree":["M. Eng."],"dc:identifier.uri":["http://hdl.handle.net/1721.1/85492"],"dc:language.iso":["eng"],"dc:publisher":["Massachusetts Institute of Technology"],"dc:rights":["M.I.T. theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. See provided URL for inquiries about permission."],"dc:rights.uri":["http://dspace.mit.edu/handle/1721.1/7582"],"dc:subject":["Electrical Engineering and Computer Science."],"dc:title":["Characterizing and predicting enhancers in the human genome"],"dc:type":["Thesis"]},"updated_at":"2026-07-22T22:21:44Z"}