期刊
BMC BIOINFORMATICS
卷 23, 期 1, 页码 -出版社
BMC
DOI: 10.1186/s12859-022-04723-w
关键词
DeSP; DNA storage; Systematic error simulation; Encoding optimization; Web application
类别
资金
- National Key R&D Program of China [2020YFA0906900]
- National Natural Science Foundation of China [62050152, 61773230, 61721003]
This study presents DeSP, a simulation pipeline for DNA storage errors, which systematically optimizes encoding redundancy by simulating errors generated in different stages. The simulation results are consistent with the in vitro experiments.
Background: Using DNA as a storage medium is appealing due to the information density and longevity of DNA, especially in the era of data explosion. A significant challenge in the DNA data storage area is to deal with the noises introduced in the channel and control the trade-off between the redundancy of error correction codes and the information storage density. As running DNA data storage experiments in vitro is still expensive and time-consuming, a simulation model is needed to systematically optimize the redundancy to combat the channel's particular noise structure. Results: Here, we present DeSP, a systematic DNA storage error Simulation Pipeline, which simulates the errors generated from all DNA storage stages and systematically guides the optimization of encoding redundancy. It covers both the sequence lost and the within-sequence errors in the particular context of the data storage channel. With this model, we explained how errors are generated and passed through different stages to form final sequencing results, analyzed the influence of error rate and sampling depth to final error rates, and demonstrated how to systemically optimize redundancy design in silico with the simulation model. These error simulation results are consistent with the in vitro experiments. Conclusions: DeSP implemented in Python is freely available on Github (https://github.com/WangLabTHU/DeSP). It is a flexible framework for systematic error simulation in DNA storage and can be adapted to a wide range of experiment pipelines.
作者
我是这篇论文的作者
点击您的名字以认领此论文并将其添加到您的个人资料中。
推荐
暂无数据