{"id":{"repo_id":"cambridge","oai_identifier":"oai:www.repository.cam.ac.uk:1810/385512"},"canonical_url":"https://search.dev.ndltd.org/etd/cambridge/oai:www.repository.cam.ac.uk:1810/385512","repository":{"repo_id":"cambridge","name":"Cambridge University","base_url":"https://api.repository.cam.ac.uk/server/oai/request"},"display":{"title":"Sequence-driven gene regulation across cellular contexts","abstract":"Gene expression is a highly coordinated process that depends on the interplay of the DNA and transcription factors that recognise subsequences of the DNA. Transcription factors can induce chromatin accessibility and recruit co-factor proteins, leading to the indirect recruitment of RNA polymerase II to activate transcription. Observational screens and genetic data have given insights into which parts of the genome are likely involved in gene regulation but have not served as functional validation. Although massively parallel reporter assays have given some insight into how short sequences activate transcription, we do not know to what degree the determinants of expression and chromatin accessibility are disjoint. As trans-factors recognise DNA, a long-standing goal in computational biology has been to predict gene expression from the DNA sequence. In recent years, computational models have improved in the prediction of gene expression and chromatin accessibility from sequence, but it is currently hard to link predictive sequence patterns to the involved trans-regulators. Furthermore, we currently do not know how well models generalise as predictors of how human cells interpret any, including non-human, DNA. In this thesis, I address how gene expression is determined by the combination of DNA and cellular context, through a mix of dry- and wet-lab approaches. First, I developed a convolutional neural network for the prediction of pseudobulked single cell gene expression from promoter sequence alone. I additionally developed methods for linking the inferred sequence features to the involved trans-regulators. On average, these models were able to explain 21% of the variability in gene expression. Second, I developed a massively parallel assay for measuring the effects of short sequences on expression and chromatin accessibility in a single genomically integrated context. This revealed that only a subset of natively accessible sequences was able to drive high chromatin accessibility in isolation, and that the determinants of accessibility and expression were overlapping but somewhat distinct. Last, I analysed how well sequence models can predict how evolutionarily distant DNA is interpreted by human cells, using data generated with a chromosome transfer of chicken DNA in human cells. Mispredictions of multiple orders of magnitude were common, and prediction of distal chromatin accessibility was limited in accuracy.","abstract_html":"Gene expression is a highly coordinated process that depends on the interplay of the DNA and transcription factors that recognise subsequences of the DNA. Transcription factors can induce chromatin accessibility and recruit co-factor proteins, leading to the indirect recruitment of RNA polymerase II to activate transcription. Observational screens and genetic data have given insights into which parts of the genome are likely involved in gene regulation but have not served as functional validation. Although massively parallel reporter assays have given some insight into how short sequences activate transcription, we do not know to what degree the determinants of expression and chromatin accessibility are disjoint. As trans-factors recognise DNA, a long-standing goal in computational biology has been to predict gene expression from the DNA sequence. In recent years, computational models have improved in the prediction of gene expression and chromatin accessibility from sequence, but it is currently hard to link predictive sequence patterns to the involved trans-regulators. Furthermore, we currently do not know how well models generalise as predictors of how human cells interpret any, including non-human, DNA. In this thesis, I address how gene expression is determined by the combination of DNA and cellular context, through a mix of dry- and wet-lab approaches. First, I developed a convolutional neural network for the prediction of pseudobulked single cell gene expression from promoter sequence alone. I additionally developed methods for linking the inferred sequence features to the involved trans-regulators. On average, these models were able to explain 21% of the variability in gene expression. Second, I developed a massively parallel assay for measuring the effects of short sequences on expression and chromatin accessibility in a single genomically integrated context. This revealed that only a subset of natively accessible sequences was able to drive high chromatin accessibility in isolation, and that the determinants of accessibility and expression were overlapping but somewhat distinct. Last, I analysed how well sequence models can predict how evolutionarily distant DNA is interpreted by human cells, using data generated with a chromosome transfer of chicken DNA in human cells. Mispredictions of multiple orders of magnitude were common, and prediction of distal chromatin accessibility was limited in accuracy.","abstract_has_math":false,"creators":["Hepkema, Jacob"],"institution":"University of Cambridge","degree_name":"Doctor of Philosophy (PhD)","degree_level":"Doctoral","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Parts, Leopold"],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-12-28","date_published":"2024-12-28","updated_at":"2026-07-22T22:24:23Z","subjects":["Gene regulation","Genomics","Computational biology"],"languages":[],"rights":[],"rights_urls":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/e3bec0a7-77bd-4da5-a709-d2786010def2/download","http://purl.org/NET/rdflicense/allrightsreserved"],"identifier_entries":[]},"links":{"outbound_url":"https://doi.org/10.17863/CAM.119089","outbound_label":"DOI","outbound_source":"dc:identifier.doi"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Parts, Leopold"]},{"key":"dc:creator","label":"Author","values":["Hepkema, Jacob"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2024-12-28"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["University of Cambridge"]},{"key":"dc:relation.isreferencedby.uri","label":"Dc Relation Isreferencedby URI","values":["https://www.repository.cam.ac.uk/handle/1810/385512"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["Doctoral"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["Doctor of Philosophy (PhD)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Gene regulation","Genomics","Computational biology"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/e3bec0a7-77bd-4da5-a709-d2786010def2/download","http://purl.org/NET/rdflicense/allrightsreserved"]},{"key":"dc:rights.embargodate","label":"Dc Rights Embargodate","values":["2026-06-13"]},{"key":"dc:rights.embargotype","label":"Dc Rights Embargotype","values":["embargo"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.17863/CAM.119089"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/bd2d1b8f-7fb3-4db7-8d59-b3973ec30687/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Gene expression is a highly coordinated process that depends on the interplay of the DNA and transcription factors that recognise subsequences of the DNA. Transcription factors can induce chromatin accessibility and recruit co-factor proteins, leading to the indirect recruitment of RNA polymerase II to activate transcription. Observational screens and genetic data have given insights into which parts of the genome are likely involved in gene regulation but have not served as functional validation. Although massively parallel reporter assays have given some insight into how short sequences activate transcription, we do not know to what degree the determinants of expression and chromatin accessibility are disjoint. As trans-factors recognise DNA, a long-standing goal in computational biology has been to predict gene expression from the DNA sequence. In recent years, computational models have improved in the prediction of gene expression and chromatin accessibility from sequence, but it is currently hard to link predictive sequence patterns to the involved trans-regulators. Furthermore, we currently do not know how well models generalise as predictors of how human cells interpret any, including non-human, DNA. In this thesis, I address how gene expression is determined by the combination of DNA and cellular context, through a mix of dry- and wet-lab approaches. First, I developed a convolutional neural network for the prediction of pseudobulked single cell gene expression from promoter sequence alone. I additionally developed methods for linking the inferred sequence features to the involved trans-regulators. On average, these models were able to explain 21% of the variability in gene expression. Second, I developed a massively parallel assay for measuring the effects of short sequences on expression and chromatin accessibility in a single genomically integrated context. This revealed that only a subset of natively accessible sequences was able to drive high chromatin accessibility in isolation, and that the determinants of accessibility and expression were overlapping but somewhat distinct. Last, I analysed how well sequence models can predict how evolutionarily distant DNA is interpreted by human cells, using data generated with a chromosome transfer of chicken DNA in human cells. Mispredictions of multiple orders of magnitude were common, and prediction of distal chromatin accessibility was limited in accuracy."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["87eda9de84448d1f82354d60eee3eb5f","345edf75fa8a01287403f045c9799cdd"]},{"key":"dc:title","label":"Title","values":["Sequence-driven gene regulation across cellular contexts"]}]}],"canonical_facts":{"dc:contributor.advisor":["Parts, Leopold"],"dc:creator":["Hepkema, Jacob"],"dc:date.issued":["2024-12-28"],"dc:description.abstract":["Gene expression is a highly coordinated process that depends on the interplay of the DNA and transcription factors that recognise subsequences of the DNA. Transcription factors can induce chromatin accessibility and recruit co-factor proteins, leading to the indirect recruitment of RNA polymerase II to activate transcription. Observational screens and genetic data have given insights into which parts of the genome are likely involved in gene regulation but have not served as functional validation. Although massively parallel reporter assays have given some insight into how short sequences activate transcription, we do not know to what degree the determinants of expression and chromatin accessibility are disjoint. As trans-factors recognise DNA, a long-standing goal in computational biology has been to predict gene expression from the DNA sequence. In recent years, computational models have improved in the prediction of gene expression and chromatin accessibility from sequence, but it is currently hard to link predictive sequence patterns to the involved trans-regulators. Furthermore, we currently do not know how well models generalise as predictors of how human cells interpret any, including non-human, DNA. In this thesis, I address how gene expression is determined by the combination of DNA and cellular context, through a mix of dry- and wet-lab approaches. First, I developed a convolutional neural network for the prediction of pseudobulked single cell gene expression from promoter sequence alone. I additionally developed methods for linking the inferred sequence features to the involved trans-regulators. On average, these models were able to explain 21% of the variability in gene expression. Second, I developed a massively parallel assay for measuring the effects of short sequences on expression and chromatin accessibility in a single genomically integrated context. This revealed that only a subset of natively accessible sequences was able to drive high chromatin accessibility in isolation, and that the determinants of accessibility and expression were overlapping but somewhat distinct. Last, I analysed how well sequence models can predict how evolutionarily distant DNA is interpreted by human cells, using data generated with a chromosome transfer of chicken DNA in human cells. Mispredictions of multiple orders of magnitude were common, and prediction of distal chromatin accessibility was limited in accuracy."],"dc:format.checksum.md5":["87eda9de84448d1f82354d60eee3eb5f","345edf75fa8a01287403f045c9799cdd"],"dc:identifier.doi":["https://doi.org/10.17863/CAM.119089"],"dc:identifier.uri":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/bd2d1b8f-7fb3-4db7-8d59-b3973ec30687/download"],"dc:publisher.institution":["University of Cambridge"],"dc:relation.isreferencedby.uri":["https://www.repository.cam.ac.uk/handle/1810/385512"],"dc:rights":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/e3bec0a7-77bd-4da5-a709-d2786010def2/download","http://purl.org/NET/rdflicense/allrightsreserved"],"dc:rights.embargodate":["2026-06-13"],"dc:rights.embargotype":["embargo"],"dc:subject":["Gene regulation","Genomics","Computational biology"],"dc:title":["Sequence-driven gene regulation across cellular contexts"],"dc:type":["Thesis"],"dc:type.qualificationlevel":["Doctoral"],"dc:type.qualificationname":["Doctor of Philosophy (PhD)"]},"updated_at":"2026-07-22T22:24:23Z"}