Abstract
dc:descriptionDNA-based data storage systems have received significant attention in the synthetic biology, computer science and information theory communities due to their promise of ultrahigh storage density, recording durability, energy efficiency, environment friendliness and potential capability of integration with in-memory computing platforms. In such systems, user content is stored in synthetic DNA oligos comprised of natural DNA nucleotides (A, T, C, and G) and retrieved via next generation (e.g., Illumina) or third generation (e.g., Oxford Nanopores) sequencing technologies. Despite recent advances in DNA synthesis and sequencing methods, all known DNA-based data storage platforms suffer from high cost, read-write latency and significant error rates that render them noncompetitive with modern electronic storage devices. Here, we introduce new approaches for encoding and reading information in DNA molecules. We first demonstrate that one can use readily available native DNA extracted from living cells (rather than using synthetic DNA molecules) to store information, and we further show that information can be stored in the topology of the sugar-phosphate backbone in the form of single-bond breaks known as ‘nicks’ (rather than storing information only in the sequence content). We show that information written in nicks can also be retrieved via a commonly used sequencing platform such as Illumina MiSeq, which is similar to synthetic DNA-based data storage systems. We further demonstrate that nick-based and sequence content-based recording approaches can be combined to generate a two-dimensional data storage system, where the sequence is reserved for archival data and metadata is written in the backbone of the molecule. In a third project, we introduce a fundamentally new concept for a prototype of a DNA-based recorder that uses an extended DNA alphabet comprised of the four canonical DNA nucleotides in addition to seven chemically modified nucleotides. The DNA data storage platform with an extended alphabet holds the potential for a ~2-fold increase in the information storage density. We demonstrate that combinatorial patterns, generated from these additional nucleobases as well as the natural nucleotides, can be accurately discriminated using MspA and Oxford nanopores, making them suitable candidates for carrying digital information. Overall, the work presented in this thesis fundamentally advances the field of macromolecular data storage.
Degree
thesis:*- Name thesis:degree_name
- Ph.D.
- Level thesis:degree_level
- Dissertation
- Discipline thesis:degree_discipline
- Biophysics & Quant Biology
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2022
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Tabatabaei, Seyed Kasra
- Contributors dc:contributor
-
- Milenkovic, Olgica
- Schroeder, Charles
- Aksimentiev, Oleksii
- Lu, Yi
Subjects
dc:subject × 8Rights
dc:rights- Statement dc:rights
-
- © 2022 Seyed Kasra Tabatabaei. All rights reserved.
- Language dc:language
- en, eng
Identifiers
dc:identifier.*- Handle dc:identifier
- https://hdl.handle.net/2142/115298