University of Cambridge
High-throughput protein engineering tools for efficient exploration and quantitative mapping of protein fitness landscapes
Abstract
dc:description.abstractOver the past few decades, directed evolution, a process involving iterative genotype diversification and isolation of desired phenotypes through screening or selection, has had a profound impact on society. It has led to the development of key pharmaceuticals, precision gene editing tools that target previously incurable diseases, and biocatalysts for the production of active pharmaceutical ingredients, fine chemicals, and biofuels that pave the way towards sustainable bioeconomy. Despite these inspiring achievements, engineering proteins to perform fit-for-purpose functions remains a complex challenge. Its difficulty arises from two main factors: (1) the nearly boundless search space of amino acid sequences resulting from the combinatorial explosion of possible mutations, and (2) the complex sequence-function relationships that make it challenging to rationally narrow down the vast sequence search space to a functionally relevant search space of experimentally tractable size. In Chapter 1, I discuss the available technologies for performing the two main pillars of directed evolution – diversifying the gene pool and selecting the right phenotype – along with their advantages and challenges. Based on this analysis, I argue that the limitations of current experimental approaches in directed evolution can be broadly divided into those that reduce sampleable sequence space (e.g. transformation efficiency, mutational bias that makes certain mutations inaccessible, low-throughput assays) and those that interfere with the accuracy of the measurements and thus the derived fitness function (e.g. stochastic noise in single-cell experiments, variability in expression levels in cell-based and lysate assays, use of proxy reaction substrates, discrepancies between assay reaction conditions and final application conditions). These limitations not only hinder functional screening - a desired phenotype might exist in the theoretical sequence space but be either inaccessible for experimental measurements or measured inaccurately - but also impede the generation of high-quality datasets for in silico machine learning-based protein engineering. Subsequent chapters introduce the two technology platforms that I developed to overcome these challenges. In Chapter 2, I describe the development, characterization, and successful implementation of a new method for the generation of DNA libraries in vivo, YeastIT mutagenesis (Yeast In vivo Targeted mutagenesis). By leveraging S. cerevisiae engineered to exhibit a high mutagenic activity specifically directed towards the gene of interest, YeastIT eliminates the transformation efficiency bottleneck, enables the creation of larger DNA libraries, and reduces the experimental timelines of iterative directed evolution cycles, thus permitting long adaptive walks to functional solutions. I outline the design and development of a mutagenesis platform, which relies on generating targeted DNA damage through nucleoside deamination, followed by the introduction of mutations by harnessing the process of error-prone DNA translesion synthesis. Furthermore, I characterize the mutational spectrum and frequency of the YeastIT-generated libraries and compare them with the existing methods (error-prone PCR, PACE, MutaT7, eMutaT7, OrthoRep, TRIDENT, EvolVR), demonstrating comparable rates and a significant reduction in the mutagenic bias relative to most of the existing alternatives. Finally, I demonstrate the effectiveness of the method by applying it for directed evolution of a DARPin to achieve a 15-fold improved affinity. In Chapter 3, I introduce the new platform for enzyme screening, BEAD-UnMESS (Bead-enabled Enzyme Activity Determination with Unbiased Microfluidic Emulsion Sorting and Sequencing), which enables quantitative enzymatic activity measurements in vitro at ultra-high throughput. By replacing single cells with monodisperse hydrogel beads displaying individual enzyme variants at precisely controlled concentrations and encoding their corresponding genetic sequences, the platform eliminates two limitations that are inherent to cell-based microfluidic screening: single-cell stochasticity and expression level variability between variants. Furthermore, the modularity of the platform enables the activity screening under defined reaction conditions rather than in crude cell lysates, thus avoiding the background signal from reactions catalysed by native host enzymes. These advantages eliminate the main noise components present in conventional droplet microfluidic assays and enable more accurate measurements of protein variant’s fitness. Using phosphotriesterase as an example, I describe the repurposing of previously developed polyacrylamide bead display strategy for enzyme screening and the development of individual platform modules: on-bead DNA library synthesis, emulsion-based in vitro transcription-translation, and microfluidic droplet sorting protocol that enables quantitative activity characterization and kinetic parameter inference of individual variants in enzyme libraries. Overall, the two high-throughput tools presented in this thesis aid protein engineering research by enabling efficient exploration and quantitative mapping of protein fitness landscapes. In the future, these workflows could be applied to accelerate the development of valuable enzyme catalysts for industrial bioprocesses, pharmaceutically relevant binders, and other proteins with important biotechnological applications. Furthermore, both methods can be harnessed in the fundamental research to unveil the complexity of Nature: YeastIT can be applied to study and model evolutionary processes, while BEAD-UnMESS can be applied to better understand epistatic interactions between protein residues, yielding mechanistic and biophysical insights.
Degree
thesis:*- Name dc:type.qualificationname
- Doctor of Philosophy (PhD)
- Level dc:type.qualificationlevel
- Doctoral
- Grantor dc:publisher.institution
- University of Cambridge
- Year dc:date.issued
- 2023
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Napiorkowska, Marta
- Advisor dc:contributor.advisor
-
- Hollfelder, Florian
Subjects
dc:subject × 6Rights
dc:rightsIdentifiers
dc:identifier.*- Author Identifier
- 0000-0001-7776-5392
- OAI identifier oai:identifier
- oai:www.repository.cam.ac.uk:1810/388339