{"id":{"repo_id":"wustl","oai_identifier":"oai:openscholarship.wustl.edu:etd-1674"},"canonical_url":"https://search.dev.ndltd.org/etd/wustl/oai:openscholarship.wustl.edu:etd-1674","repository":{"repo_id":"wustl","name":"Washington University in St. Louis","base_url":"https://openscholarship.wustl.edu/do/oai/"},"display":{"title":"Quantitative Analysis Demonstrates Most Transcription Factors Require only Simple Models of Specificity","abstract":"Organisms must control their gene expression to properly respond to developmental, stress or other environmental cues. A key part of this process is transcriptional regulation, which is largely accomplished by a complex network of transcription factor proteins: TFs) interact with their specific binding sites in the genome. Understanding how TFs select correct binding sites out of the vast number of potential binding sites in the genome is a key challenge in molecular biology. Recently, unprecedented amount of quantitative binding data have become available as results of developments in high-throughput experimental techniques. However, interpretation of high-throughput binding data has proved to be controversial, largely due to the lack of physically principled data analysis methods. ii An important question in the analysis of binding data is the complexity of the specificity model needed. This has important implications for both the characterization of specificity and for the prediction of the consequences of mutations. Structurally, TF-DNA interactions are complex with a wide variety of interactions between the protein and DNA making a simple recognition code impossible. Energetically, however, the situation may be much simpler. Detailed studies of a handful of TFs have shown that individual base pairs often contribute independently to the total binding energy. This view of simplicity has been challenged by data from high-throughput binding experiments, although the extent to which the sample model breaks down is uncertain due to lack of rigorous analysis methods. The goal of this thesis is to assess the complexity of model required to accurately represent TF specificity. To this end, I have developed a new statistical analysis method BEEML: Binding Energy Estimation by Maximum Likelihood) that parameterizes models of TF specificity from high-throughput quantitative binding data, using a realistic biophysical model. Employing the BEEML method, I show that the energetics of most TF-DNA interactions are simple, with bases in the binding site contribute approximately independently to the total binding energy. Further, I show that interactions in the binding site occur mostly between adjacent positions.","abstract_html":"Organisms must control their gene expression to properly respond to developmental, stress or other environmental cues. A key part of this process is transcriptional regulation, which is largely accomplished by a complex network of transcription factor proteins: TFs) interact with their specific binding sites in the genome. Understanding how TFs select correct binding sites out of the vast number of potential binding sites in the genome is a key challenge in molecular biology. Recently, unprecedented amount of quantitative binding data have become available as results of developments in high-throughput experimental techniques. However, interpretation of high-throughput binding data has proved to be controversial, largely due to the lack of physically principled data analysis methods. ii An important question in the analysis of binding data is the complexity of the specificity model needed. This has important implications for both the characterization of specificity and for the prediction of the consequences of mutations. Structurally, TF-DNA interactions are complex with a wide variety of interactions between the protein and DNA making a simple recognition code impossible. Energetically, however, the situation may be much simpler. Detailed studies of a handful of TFs have shown that individual base pairs often contribute independently to the total binding energy. This view of simplicity has been challenged by data from high-throughput binding experiments, although the extent to which the sample model breaks down is uncertain due to lack of rigorous analysis methods. The goal of this thesis is to assess the complexity of model required to accurately represent TF specificity. To this end, I have developed a new statistical analysis method BEEML: Binding Energy Estimation by Maximum Likelihood) that parameterizes models of TF specificity from high-throughput quantitative binding data, using a realistic biophysical model. Employing the BEEML method, I show that the energetics of most TF-DNA interactions are simple, with bases in the binding site contribute approximately independently to the total binding energy. Further, I show that interactions in the binding site occur mostly between adjacent positions.","abstract_has_math":false,"creators":["Zhao, Yue"],"institution":null,"degree_name":"Doctor of Philosophy (PhD)","degree_level":"Dissertation","degree_discipline":"Biology and Biomedical Sciences: Computational and Systems Biology","degree_department":null,"school":null,"contributors":["Gary Stormo"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2011,"date_issued":"2011-01-01T08:00:00Z","date_published":"2011-01-01T08:00:00Z","updated_at":"2026-07-24T06:12:48Z","subjects":["Bioinformatics","Biochemistry","Quantitative Model","Transcription Factor Specificity"],"languages":["English (en)"],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.7936/K7707ZFG"],"render_values":[{"text":"https://doi.org/10.7936/K7707ZFG","href":"https://doi.org/10.7936/K7707ZFG","code":true}]}]},"links":{"outbound_url":"https://openscholarship.wustl.edu/etd/675","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Gary Stormo"]},{"key":"dc:creator","label":"Author","values":["Zhao, Yue"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2012-05-17T07:00:00Z"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Biology and Biomedical Sciences: Computational and Systems Biology"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctor of Philosophy (PhD)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Bioinformatics","Biochemistry","Quantitative Model","Transcription Factor Specificity"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["English (en)"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://openscholarship.wustl.edu/etd/675"]},{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.7936/K7707ZFG"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Organisms must control their gene expression to properly respond to developmental, stress or other environmental cues. A key part of this process is transcriptional regulation, which is largely accomplished by a complex network of transcription factor proteins: TFs) interact with their specific binding sites in the genome. Understanding how TFs select correct binding sites out of the vast number of potential binding sites in the genome is a key challenge in molecular biology. Recently, unprecedented amount of quantitative binding data have become available as results of developments in high-throughput experimental techniques. However, interpretation of high-throughput binding data has proved to be controversial, largely due to the lack of physically principled data analysis methods. ii An important question in the analysis of binding data is the complexity of the specificity model needed. This has important implications for both the characterization of specificity and for the prediction of the consequences of mutations. Structurally, TF-DNA interactions are complex with a wide variety of interactions between the protein and DNA making a simple recognition code impossible. Energetically, however, the situation may be much simpler. Detailed studies of a handful of TFs have shown that individual base pairs often contribute independently to the total binding energy. This view of simplicity has been challenged by data from high-throughput binding experiments, although the extent to which the sample model breaks down is uncertain due to lack of rigorous analysis methods. The goal of this thesis is to assess the complexity of model required to accurately represent TF specificity. To this end, I have developed a new statistical analysis method BEEML: Binding Energy Estimation by Maximum Likelihood) that parameterizes models of TF specificity from high-throughput quantitative binding data, using a realistic biophysical model. Employing the BEEML method, I show that the energetics of most TF-DNA interactions are simple, with bases in the binding site contribute approximately independently to the total binding energy. Further, I show that interactions in the binding site occur mostly between adjacent positions."]},{"key":"dc:title","label":"Title","values":["Quantitative Analysis Demonstrates Most Transcription Factors Require only Simple Models of Specificity"]}]}],"canonical_facts":{"dc:contributor":["Gary Stormo"],"dc:creator":["Zhao, Yue"],"dc:date.available":["2012-05-17T07:00:00Z"],"dc:description.abstract":["Organisms must control their gene expression to properly respond to developmental, stress or other environmental cues. A key part of this process is transcriptional regulation, which is largely accomplished by a complex network of transcription factor proteins: TFs) interact with their specific binding sites in the genome. Understanding how TFs select correct binding sites out of the vast number of potential binding sites in the genome is a key challenge in molecular biology. Recently, unprecedented amount of quantitative binding data have become available as results of developments in high-throughput experimental techniques. However, interpretation of high-throughput binding data has proved to be controversial, largely due to the lack of physically principled data analysis methods. ii An important question in the analysis of binding data is the complexity of the specificity model needed. This has important implications for both the characterization of specificity and for the prediction of the consequences of mutations. Structurally, TF-DNA interactions are complex with a wide variety of interactions between the protein and DNA making a simple recognition code impossible. Energetically, however, the situation may be much simpler. Detailed studies of a handful of TFs have shown that individual base pairs often contribute independently to the total binding energy. This view of simplicity has been challenged by data from high-throughput binding experiments, although the extent to which the sample model breaks down is uncertain due to lack of rigorous analysis methods. The goal of this thesis is to assess the complexity of model required to accurately represent TF specificity. To this end, I have developed a new statistical analysis method BEEML: Binding Energy Estimation by Maximum Likelihood) that parameterizes models of TF specificity from high-throughput quantitative binding data, using a realistic biophysical model. Employing the BEEML method, I show that the energetics of most TF-DNA interactions are simple, with bases in the binding site contribute approximately independently to the total binding energy. Further, I show that interactions in the binding site occur mostly between adjacent positions."],"dc:identifier":["https://openscholarship.wustl.edu/etd/675"],"dc:identifier.doi":["https://doi.org/10.7936/K7707ZFG"],"dc:language":["English (en)"],"dc:subject":["Bioinformatics","Biochemistry","Quantitative Model","Transcription Factor Specificity"],"dc:title":["Quantitative Analysis Demonstrates Most Transcription Factors Require only Simple Models of Specificity"],"thesis:degree_discipline":["Biology and Biomedical Sciences: Computational and Systems Biology"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Doctor of Philosophy (PhD)"]},"updated_at":"2026-07-24T06:12:48Z"}