Back to results

University of Cambridge

Sequence-driven gene regulation across cellular contexts

Abstract

dc:description.abstract

Gene expression is a highly coordinated process that depends on the interplay of the DNA and transcription factors that recognise subsequences of the DNA. Transcription factors can induce chromatin accessibility and recruit co-factor proteins, leading to the indirect recruitment of RNA polymerase II to activate transcription. Observational screens and genetic data have given insights into which parts of the genome are likely involved in gene regulation but have not served as functional validation. Although massively parallel reporter assays have given some insight into how short sequences activate transcription, we do not know to what degree the determinants of expression and chromatin accessibility are disjoint. As trans-factors recognise DNA, a long-standing goal in computational biology has been to predict gene expression from the DNA sequence. In recent years, computational models have improved in the prediction of gene expression and chromatin accessibility from sequence, but it is currently hard to link predictive sequence patterns to the involved trans-regulators. Furthermore, we currently do not know how well models generalise as predictors of how human cells interpret any, including non-human, DNA. In this thesis, I address how gene expression is determined by the combination of DNA and cellular context, through a mix of dry- and wet-lab approaches. First, I developed a convolutional neural network for the prediction of pseudobulked single cell gene expression from promoter sequence alone. I additionally developed methods for linking the inferred sequence features to the involved trans-regulators. On average, these models were able to explain 21% of the variability in gene expression. Second, I developed a massively parallel assay for measuring the effects of short sequences on expression and chromatin accessibility in a single genomically integrated context. This revealed that only a subset of natively accessible sequences was able to drive high chromatin accessibility in isolation, and that the determinants of accessibility and expression were overlapping but somewhat distinct. Last, I analysed how well sequence models can predict how evolutionarily distant DNA is interpreted by human cells, using data generated with a chromosome transfer of chicken DNA in human cells. Mispredictions of multiple orders of magnitude were common, and prediction of distal chromatin accessibility was limited in accuracy.

Degree

thesis:*
Name dc:type.qualificationname
Doctor of Philosophy (PhD)
Level dc:type.qualificationlevel
Doctoral
Grantor dc:publisher.institution
University of Cambridge
Year dc:date.issued
2024

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Hepkema, Jacob
Advisor dc:contributor.advisor
  • Parts, Leopold

Subjects

dc:subject × 3

Rights

dc:rights

Identifiers

dc:identifier.*
DOI dc:identifier.doi
https://doi.org/10.17863/CAM.119089
OAI identifier oai:identifier
oai:www.repository.cam.ac.uk:1810/385512

Chain of custody

source
Harvested from
Cambridge University
Base URL
api.repository.cam.ac.uk/server/oai/request
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Hepkema, Jacob. Sequence-driven gene regulation across cellular contexts. Doctoral thesis, University of Cambridge, 2024. https://doi.org/10.17863/CAM.119089