DeSP: a systematic DNA storage error simulation pipeline.

BMC Bioinformatics

Ministry of Education Key Laboratory of Bioinformatics; Center for Synthetic and Systems Biology; Beijing National Research Center for Information Science and Technology; Department of Automation, Tsinghua University, Beijing, China, 100084.

Published: May 2022

Background: Using DNA as a storage medium is appealing due to the information density and longevity of DNA, especially in the era of data explosion. A significant challenge in the DNA data storage area is to deal with the noises introduced in the channel and control the trade-off between the redundancy of error correction codes and the information storage density. As running DNA data storage experiments in vitro is still expensive and time-consuming, a simulation model is needed to systematically optimize the redundancy to combat the channel's particular noise structure.

Results: Here, we present DeSP, a systematic DNA storage error Simulation Pipeline, which simulates the errors generated from all DNA storage stages and systematically guides the optimization of encoding redundancy. It covers both the sequence lost and the within-sequence errors in the particular context of the data storage channel. With this model, we explained how errors are generated and passed through different stages to form final sequencing results, analyzed the influence of error rate and sampling depth to final error rates, and demonstrated how to systemically optimize redundancy design in silico with the simulation model. These error simulation results are consistent with the in vitro experiments.

Conclusions: DeSP implemented in Python is freely available on Github ( https://github.com/WangLabTHU/DeSP ). It is a flexible framework for systematic error simulation in DNA storage and can be adapted to a wide range of experiment pipelines.

Download full-text PDF

Source
http://www.ncbi.nlm.nih.gov/pmc/articles/PMC9116035PMC
http://dx.doi.org/10.1186/s12859-022-04723-wDOI Listing

Publication Analysis

Top Keywords

dna storage
20
error simulation
16
data storage
12
storage
9
desp systematic
8
dna
8
systematic dna
8
storage error
8
simulation pipeline
8
dna data
8

Similar Publications

Want AI Summaries of new PubMed Abstracts delivered to your In-box?

Enter search terms and have AI summaries delivered each week - change queries or unsubscribe any time!