CN113782102B - DNA data storage method, device, equipment and readable storage medium - Google Patents
DNA data storage method, device, equipment and readable storage medium Download PDFInfo
- Publication number
- CN113782102B CN113782102B CN202110929436.9A CN202110929436A CN113782102B CN 113782102 B CN113782102 B CN 113782102B CN 202110929436 A CN202110929436 A CN 202110929436A CN 113782102 B CN113782102 B CN 113782102B
- Authority
- CN
- China
- Prior art keywords
- sequence
- dna
- base
- fragments
- units
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Expired - Fee Related
Links
Images
Classifications
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B50/00—ICT programming tools or database systems specially adapted for bioinformatics
- G16B50/30—Data warehousing; Computing architectures
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- Health & Medical Sciences (AREA)
- Life Sciences & Earth Sciences (AREA)
- Bioethics (AREA)
- Biophysics (AREA)
- Databases & Information Systems (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Bioinformatics & Computational Biology (AREA)
- Biotechnology (AREA)
- Evolutionary Biology (AREA)
- General Health & Medical Sciences (AREA)
- Medical Informatics (AREA)
- Spectroscopy & Molecular Physics (AREA)
- Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
Abstract
Description
技术领域technical field
本申请属于数据存储技术领域,具体涉及一种DNA数据的存储方法、装置、设备及可读存储介质。The present application belongs to the technical field of data storage, and in particular relates to a DNA data storage method, device, equipment and readable storage medium.
背景技术Background technique
人工智能及大数据时代的发展对数据存储需求越来越高,迫切需要存储密度高、存储时间长、维护成本低的新型存储介质。脱氧核糖核酸(DeoxyriboNucleic Acid,DNA)作为一种近年来发展起来的信息存储介质,被认为是未来信息存储最有潜力的介质之一。The development of the era of artificial intelligence and big data has higher and higher demand for data storage, and there is an urgent need for new storage media with high storage density, long storage time, and low maintenance costs. Deoxyribonucleic acid (DeoxyriboNucleic Acid, DNA), as an information storage medium developed in recent years, is considered to be one of the most potential mediums for future information storage.
DNA分子具有四种碱基,它们分别是:腺嘌呤(Adenine,A)、胞嘧啶(Cytosine,C)、鸟嘌呤(Guanine,G)和胸腺嘧啶(Thymine,T)。基于DNA的数据存储技术是利用上述四种碱基序列来表示二进制“0”和“1”组成的数据系列。相比较于传统存储介质,DNA数据存储具有存储密度高,存储时间久,维护成本低,生物相容性好的特点。如:1g DNA能够存储超过百万部高清电影,其数据存储密度是目前传统硬盘等硅基存储介质7个数量级以上;同时,DNA能够稳定存储数据千年以上,是现有存储介质存储时间的百倍以上。此外,DNA维护成本低,存百年的维护费用仅是目前现有介质的万分之一。DNA molecules have four bases, which are: adenine (A), cytosine (Cytosine, C), guanine (Guanine, G) and thymine (Thymine, T). The DNA-based data storage technology uses the above four base sequences to represent the data series composed of binary "0" and "1". Compared with traditional storage media, DNA data storage has the characteristics of high storage density, long storage time, low maintenance cost, and good biocompatibility. For example: 1g of DNA can store more than one million high-definition movies, and its data storage density is more than 7 orders of magnitude higher than the current silicon-based storage media such as traditional hard disks; at the same time, DNA can store data stably for more than a thousand years, which is a hundred times the storage time of existing storage media above. In addition, the maintenance cost of DNA is low, and the maintenance cost of storing for a hundred years is only one ten-thousandth of the current existing media.
DNA数据存储流程通常包含以下步骤:(1)从图片、视频、文本等计算机信息中提取二进制信息;(2)根据二进制与碱基A、T、C、G之间的预设对应关系,将二进制序列信息转换为由碱基A、T、C、G编码形成的、存储有数据信息的A/T/C/G序列(即DNA序列);(3)采用DNA合成技术或其他技术将编码的A/T/C/G序列转换为DNA化学多聚物分子,并存储在合适的环境中。之后,当需要获取存储的数据时,则可以执行以下步骤:(4)利用DNA测序技术,将存储的DNA化学多聚物分子解读成A/T/C/G序列;(5)利用合适的解码方式将A/T/C/G序列转换为二进制信息;(6)将二进制信息转换为图片、视频、文本等计算机信息。The DNA data storage process usually includes the following steps: (1) Extract binary information from computer information such as pictures, videos, and texts; (2) According to the preset correspondence between binary and bases A, T, C, and G, the Binary sequence information is converted into A/T/C/G sequences (that is, DNA sequences) formed by base A, T, C, and G codes and stored with data information; (3) DNA synthesis technology or other technologies are used to convert the coded The A/T/C/G sequence is converted into a DNA chemical polymer molecule and stored in a suitable environment. Afterwards, when it is necessary to obtain the stored data, the following steps can be performed: (4) Utilize DNA sequencing technology to interpret the stored DNA chemical polymer molecules into A/T/C/G sequences; (5) Utilize appropriate The decoding method converts A/T/C/G sequences into binary information; (6) converts binary information into computer information such as pictures, videos, and texts.
其中,数据编码问题是目前的DNA数据存储方法中的核心问题。Among them, the data encoding problem is the core problem in the current DNA data storage methods.
发明内容Contents of the invention
本申请实施例的目的之一在于:提供一种DNA数据的存储方法、装置、设备及可读存储介质,旨在解决DNA数据存储技术中的数据编码问题。One of the purposes of the embodiments of the present application is to provide a DNA data storage method, device, device and readable storage medium, aiming at solving the data encoding problem in the DNA data storage technology.
本申请实施例采用的技术方案是:The technical scheme that the embodiment of the present application adopts is:
第一方面,提供了一种DNA数据的存储方法,包括:In the first aspect, a method for storing DNA data is provided, including:
获取与目标数据的二进制序列对应的碱基序列;Obtain the base sequence corresponding to the binary sequence of the target data;
对所述碱基序列进行分割,得到S个序列单元,且每个序列单元包括多个分割的序列片段,其中,S个所述序列单元共含有K个所述序列片段,且所述序列片段的长度为n,n、S和K均为大于或者等于2的整数;Segmenting the base sequence to obtain S sequence units, and each sequence unit includes a plurality of segmented sequence fragments, wherein, the S sequence units contain K sequence fragments in total, and the sequence fragments The length of is n, and n, S and K are all integers greater than or equal to 2;
利用预设的索引信息对K个所述序列片段和S个所述序列单元进行标记,得到K个标记序列片段和S个标记序列单元,其中,所述索引信息包括用于表示S个所述序列单元在所述碱基序列中的排列顺序的第一检索序列,和用于表示属于同一所述序列单元中的多个所述序列片段在所述序列单元中的排列顺序的第二检索序列,K个所述标记序列片段用于合成存储有所述目标数据的K个第一DNA分子。The K sequence fragments and the S sequence units are marked by using the preset index information to obtain K mark sequence fragments and S mark sequence units, wherein the index information includes information used to represent the S sequence units A first search sequence for the arrangement order of sequence units in the base sequence, and a second search sequence for indicating the arrangement order of multiple sequence fragments belonging to the same sequence unit in the sequence unit , K pieces of the marker sequence fragments are used to synthesize K first DNA molecules storing the target data.
在一个实施例中,利用所述第二检索序列标记属于同一所述序列单元中的多个所述序列片段的方式,包括:In one embodiment, using the second retrieval sequence to mark multiple sequence fragments belonging to the same sequence unit includes:
在所述序列片段的任一侧拼接第二检索序列,或splicing a second search sequence on either side of said sequence fragment, or
在所述序列片段的两侧同时拼接检索碱基组,两侧的所述检索碱基组形成所述第二检索序列。The search base groups are spliced simultaneously on both sides of the sequence fragment, and the search base groups on both sides form the second search sequence.
在一个实施例中,所述第一检索序列包括i条DNA序列片段,i为大于或等于1的整数,且每条所述DNA序列片段包括用作索引标志的第一碱基序列和用于标示所述序列单元编号的第二碱基序列。In one embodiment, the first search sequence includes i DNA sequence fragments, i is an integer greater than or equal to 1, and each of the DNA sequence fragments includes the first base sequence used as an index mark and the Mark the second base sequence of the sequence unit number.
在一个实施例中,第一检索序列和第二检索序列对应的DNA序列片段利用DNA合成技术获得。示例性的,DNA合成技术包括但不限于酶法合成、亚磷酰胺合成等。In one embodiment, the DNA sequence fragments corresponding to the first search sequence and the second search sequence are obtained using DNA synthesis technology. Exemplarily, DNA synthesis techniques include, but are not limited to, enzymatic synthesis, phosphoramidite synthesis, and the like.
在一个实施例中,第一检索序列和第二检索序列对应的DNA序列片段可以从预先合成的DNA通用分子库中扩增获得,比如PCR技术等。In one embodiment, the DNA sequence fragments corresponding to the first search sequence and the second search sequence can be amplified from a pre-synthesized DNA universal molecular library, such as PCR technology.
在一个实施例中,所述存储方法还包括:In one embodiment, the storage method also includes:
将K个所述标记序列片段合成存储有所述目标数据的K个第一DNA分子后,将K个所述第一DNA分子存储在S个第一物理空间,其中,同属于一个所述序列单元的所述标记序列片段对应的所述第一DNA分子存储在同一个所述第一物理空间,不属于同一个所述序列单元的所述标记序列片段对应的所述第一DNA分子存储在不同的所述第一物理空间。After synthesizing K first DNA molecules storing the target data from the K marker sequence fragments, storing the K first DNA molecules in S first physical spaces, wherein the same sequences belong to one The first DNA molecules corresponding to the marker sequence fragments of the unit are stored in the same first physical space, and the first DNA molecules corresponding to the marker sequence fragments that do not belong to the same sequence unit are stored in Different from said first physical space.
在一个实施例中,s个所述第一物理空间集成在一个DNA硬盘中。In one embodiment, the s first physical spaces are integrated into one DNA hard disk.
在一个实施例中,所述存储方法还包括:In one embodiment, the storage method also includes:
将第二DNA分子存储在与所述第一物理空间对应的第二物理空间,所述第二DNA分子存储有所述索引信息。storing a second DNA molecule in a second physical space corresponding to the first physical space, the second DNA molecule storing the index information.
在一个实施例中,K个所述第一DNA分子的解码方法,包括:In one embodiment, the decoding method for the K first DNA molecules includes:
对每个所述第一物理空间中存储的多个所述第一DNA分子进行测序,得到多个所述标记序列片段;根据所述第二检索序列对属于同一所述标记序列单元的每个所述标记序列片段对应的所述序列片段进行拼接,得到所述序列单元;Sequencing the plurality of first DNA molecules stored in each of the first physical spaces to obtain a plurality of marker sequence fragments; pairing each marker sequence unit belonging to the same marker sequence unit according to the second retrieval sequence splicing the sequence fragments corresponding to the marker sequence fragments to obtain the sequence unit;
根据所述第一检索序列将得到的S个所述序列单元进行拼接,得到所述碱基序列;splicing the obtained S sequence units according to the first search sequence to obtain the base sequence;
将所述碱基序列转换为所述目标数据。The base sequence is converted into the target data.
第二方面,提供了一种DNA数据存储装置,包括数据处理模块,In a second aspect, a DNA data storage device is provided, including a data processing module,
所述数据处理模块,用于获取与目标数据的二进制序列对应的碱基序列;对所述碱基序列进行分割,得到K个长度为n的序列片段,K个所述序列片段划分为S个序列单元,S和K均为大于或者等于2的整数;利用预设的索引信息对K个所述序列片段和S个所述序列单元进行标记,得到K个标记序列片段和S个标记序列单元,其中,所述索引信息包括用于表示S个所述序列单元在所述碱基序列中的排列顺序的第一检索序列,和用于表示属于同一所述序列单元中的多个所述序列片段在所述序列单元中的排列顺序的第二检索序列,K个所述标记序列片段用于合成存储有所述目标数据的K个第一DNA分子。The data processing module is used to obtain the base sequence corresponding to the binary sequence of the target data; segment the base sequence to obtain K sequence fragments with a length of n, and the K sequence fragments are divided into S Sequence units, S and K are both integers greater than or equal to 2; use preset index information to mark K said sequence fragments and S said sequence units to obtain K marked sequence fragments and S marked sequence units , wherein, the index information includes a first search sequence used to represent the arrangement order of the S sequence units in the base sequence, and a first search sequence used to represent multiple sequences belonging to the same sequence unit The second retrieval sequence of the arrangement order of the fragments in the sequence unit, the K fragments of the marker sequence are used to synthesize the K first DNA molecules storing the target data.
在一个实施例中,所述装置还包括:DNA合成模块,用于将K个所述标记序列片段合成存储有所述目标数据的K个第一DNA分子。In one embodiment, the device further includes: a DNA synthesis module, configured to synthesize K first DNA molecules storing the target data from the K marker sequence fragments.
在一个实施例中,所述装置还包括:DNA分子存储模块,用于将K个所述第一DNA分子存储在S个第一物理空间,其中,同属于一个所述序列单元的所述标记序列片段对应的所述第一DNA分子存储在同一个所述第一物理空间,不属于同一个所述序列单元的所述标记序列片段对应的所述第一DNA分子存储在不同的所述第一物理空间。In one embodiment, the device further includes: a DNA molecule storage module, configured to store K first DNA molecules in S first physical spaces, wherein the markers belonging to one sequence unit The first DNA molecules corresponding to the sequence fragments are stored in the same first physical space, and the first DNA molecules corresponding to the marker sequence fragments that do not belong to the same sequence unit are stored in different first physical spaces. a physical space.
在一个实施例中,所述DNA分子存储模块还用于将第二DNA分子存储在第二物理空间。In one embodiment, the DNA molecule storage module is also used to store the second DNA molecule in the second physical space.
在一个实施例中,还包括DNA分子测序模块,用于对每个所述第一物理空间中存储的多个所述第一DNA分子进行测序,得到多个所述标记序列片段;所述数据处理模块,还用于对根据所述第二检索序列对属于同一所述标记序列单元的每个所述标记序列片段对应的所述序列片段进行拼接,得到所述序列单元;In one embodiment, it also includes a DNA molecule sequencing module, configured to sequence a plurality of first DNA molecules stored in each of the first physical spaces to obtain a plurality of marker sequence fragments; the data The processing module is further configured to splice the sequence fragments corresponding to each of the tag sequence fragments belonging to the same tag sequence unit according to the second retrieval sequence, to obtain the sequence unit;
根据所述第一检索序列将得到的S个所述序列单元进行拼接,得到所述碱基序列;splicing the obtained S sequence units according to the first search sequence to obtain the base sequence;
将所述碱基序列转换为所述目标数据。The base sequence is converted into the target data.
第三方面,提供了一种DNA数据存储设备,包括一种终端设备,所述终端设备包括存储器、处理器以及存储在所述存储器中并可在所述处理器上运行的计算机程序,其特征在于,所述处理器执行所述计算机程序时实现如第一方面所述的DNA数据的存储方法。In a third aspect, a DNA data storage device is provided, including a terminal device, the terminal device includes a memory, a processor, and a computer program stored in the memory and operable on the processor, characterized in That is, when the processor executes the computer program, the method for storing DNA data as described in the first aspect is realized.
第四方面,提供了一种计算机可读存储介质,计算机可读存储介质存储有计算机程序,计算机程序被处理器执行时实现如第一方面的DNA数据的存储方法。In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the method for storing DNA data in the first aspect is implemented.
第五方面,提供了一种DNA硬盘,包括多个物理空间,所述物理空间由物理材料制成,每个所述物理空间用于存储DNA分子。In a fifth aspect, a DNA hard disk is provided, including a plurality of physical spaces, the physical spaces are made of physical materials, and each of the physical spaces is used to store DNA molecules.
在一个实施例中,所述物理材料为SiO2、金属氧化物、高分子聚合材料中的至少一种。这些物理材料形成将DNA分子包裹的物理空间,同时,将不同的DNA分子隔离。In one embodiment, the physical material is at least one of SiO 2 , metal oxides, and polymer materials. These physical materials form a physical space that wraps DNA molecules, and at the same time, isolates different DNA molecules.
本申请实施例提供的DNA数据的存储方法、装置、设备及可读存储介质的有益效果在于:本申请将与目标数据的二进制序列对应的碱基序列分割为S个序列,每个序列单元包括多个分割的序列片段,S个所述序列单元共含有K个所述序列片段,且所述序列片段的长度为n,n、S和K均为大于或者等于2的整数,利用预设的索引信息标记序列片段和序列单元的位置信息,将标记后的序列片段合成为DNA分子后分别存储。通过该方法,可以提升数据信息的存储量,实现大规模的数据信息在DNA中的存储。The beneficial effect of the DNA data storage method, device, device and readable storage medium provided by the embodiment of the present application is that: the present application divides the base sequence corresponding to the binary sequence of the target data into S sequences, and each sequence unit includes A plurality of segmented sequence fragments, the S sequence units contain K sequence fragments in total, and the length of the sequence fragments is n, n, S and K are all integers greater than or equal to 2, using the preset The index information marks the position information of sequence fragments and sequence units, and the marked sequence fragments are synthesized into DNA molecules and stored separately. Through this method, the storage capacity of data information can be increased, and the storage of large-scale data information in DNA can be realized.
在一种实施情形中,第一检索序列的长度和标记序列片段(带有第二检索序列的序列片段)的长度不同,则通过长度区别从序列单元中区分第一检索序列和标记序列片段。具体的,以m表示标记序列片段的碱基数,以q表示第二检索序列的碱基数,以i表示第一检索序列中DNA序列片段的条数,以p表示第一检索序列中第二碱基序列的碱基数,通过本申请提供的方法,可以实现含有D个碱基数的DNA数据的存储,其中,D的计算公式如下:In one implementation situation, the length of the first retrieval sequence is different from that of the marker sequence fragment (the sequence fragment with the second retrieval sequence), and the first retrieval sequence and the marker sequence fragment are distinguished from the sequence units by length difference. Specifically, m represents the base number of the marker sequence fragment, q represents the base number of the second retrieval sequence, i represents the number of DNA sequence fragments in the first retrieval sequence, and p represents the number of DNA sequence fragments in the first retrieval sequence The number of bases in the two-base sequence, through the method provided by this application, can realize the storage of DNA data containing D bases, wherein, the calculation formula of D is as follows:
D=4q×(m-q)×4i×p D = 4q×(mq)× 4i×p
在特定实施例中,m长度为8碱基的情况下,q=4,i=10,p=4,D=256×4×440=1.23×1027。1个碱基存储2bits信息时候,能够存储的信息量L=2.46×1027bits=3.075×1026bytes=3.075×105ZB,远大于目前的数据存储规模。In a specific embodiment, when the length of m is 8 bases, q=4, i=10, p=4, D=256×4×4 40 =1.23×10 27 . When 1 base stores 2 bits of information, the amount of information that can be stored is L=2.46×10 27 bits=3.075×10 26 bytes=3.075×10 5 ZB, which is much larger than the current data storage scale.
在另一种实施情形中,第一检索序列的长度和标记序列片段(带有第二检索序列的序列片段)的长度相同;第一检索序列中用作索引标志的第一碱基序列,可以是第二检索碱基序列的一部分,第一碱基序列的碱基数和第二检索碱基序列的碱基数相同。在这种情况下,以m表示标记序列片段的碱基数,以q表示第二检索序列的碱基数,以i表示第一检索序列中DNA序列片段的条数,以p表示第一检索序列中第二碱基序列的碱基数,通过本申请提供的方法的方法,可以实现含有D个碱基数的DNA数据的存储,其中,D的计算公式如下:In another implementation situation, the length of the first search sequence is the same as the length of the marker sequence fragment (sequence fragment with the second search sequence); the first base sequence used as an index mark in the first search sequence can be It is a part of the second search base sequence, and the base number of the first base sequence is the same as that of the second search base sequence. In this case, let m represent the base number of the marker sequence fragment, q represent the base number of the second search sequence, use i to represent the number of DNA sequence fragments in the first search sequence, and p represent the first search sequence The number of bases in the second base sequence in the sequence can be stored by the method provided by this application, and the DNA data containing D bases can be stored, wherein, the calculation formula of D is as follows:
D=(4q-i)×(m-q)×4i×p D=( 4q- i)×(mq)× 4i×p
在特定实施例中,m长度为8碱基的情况下,q=4,i=10,p=4,D=(256-10)×4×440=1.18×1027。1个碱基存储2bits信息时候,能够存储的信息量L=2.36×1027bits=2.95×1026bytes=2.95×105ZB,远大于目前的数据存储规模。In a specific embodiment, when the length of m is 8 bases, q=4, i=10, p=4, D=(256-10)×4×4 40 =1.18×10 27 . When 1 base stores 2 bits of information, the amount of information that can be stored is L=2.36×10 27 bits=2.95×10 26 bytes=2.95×10 5 ZB, which is much larger than the current data storage scale.
附图说明Description of drawings
为了更清楚地说明本申请实施例中的技术方案,下面将对实施例或示范性技术描述中所需要使用的附图作简单地介绍,显而易见地,下面描述中的附图仅仅是本申请的一些实施例,对于本领域普通技术人员来讲,在不付出创造性劳动的前提下,还可以根据这些附图获得其它的附图。In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings that need to be used in the embodiments or exemplary technical descriptions. Obviously, the accompanying drawings in the following descriptions are only for this application. For some embodiments, those skilled in the art can also obtain other drawings based on these drawings without creative efforts.
图1是本申请实施例提供的DNA存储装置的组成示意图;Figure 1 is a schematic diagram of the composition of the DNA storage device provided by the embodiment of the present application;
图2是本申请实施例提供的DNA数据的存储写入的工艺流程图;Fig. 2 is the process flow diagram of the storage and writing of DNA data provided by the embodiment of the present application;
图3是本申请实施例提供的S101中的碱基序列经分割后形成包含多个序列片段的序列单元的示意图;Fig. 3 is a schematic diagram of a sequence unit comprising multiple sequence fragments formed after the base sequence in S101 provided in the embodiment of the present application is segmented;
图4是本申请实施例提供的S103中的序列单元经过第一检索序列标记,且序列单元中的各序列片段经过第二检索序列标记后,得到分别包含多个信息序列片段的信息序列单元的示意图;Fig. 4 is the sequence unit in S103 provided by the embodiment of the present application after the first search sequence mark, and each sequence fragment in the sequence unit has been marked by the second search sequence, to obtain information sequence units respectively containing multiple information sequence fragments schematic diagram;
图5是本申请实施例提供的将K个第一DNA分子存储在S个不同的第一物理空间后的DNA存储的示意图;Fig. 5 is a schematic diagram of DNA storage after K first DNA molecules are stored in S different first physical spaces provided by the embodiment of the present application;
图6是本申请实施例提供的DNA数据的解读工艺流程图;Fig. 6 is the flow chart of the interpretation process of the DNA data provided by the embodiment of the present application;
图7本申请一实施例提供的终端设备的结构示意图。FIG. 7 is a schematic structural diagram of a terminal device provided by an embodiment of the present application.
具体实施方式detailed description
为了使本申请的目的、技术方案及优点更加清楚明白,以下结合附图及实施例,对本申请进行进一步详细说明。应当理解,此处所描述的具体实施例仅用以解释本发明,并不用于限定本申请。In order to make the purpose, technical solution and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described here are only used to explain the present invention, not to limit the present application.
以下描述中,为了说明而不是为了限定,提出了诸如特定系统结构、技术之类的具体细节,以便透彻理解本申请实施例。然而,本领域的技术人员应当清楚,在没有这些具体细节的其它实施例中也可以实现本申请。在其它情况中,省略对众所周知的系统、装置、电路以及方法的详细说明,以免不必要的细节妨碍本申请的描述。In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. It will be apparent, however, to one skilled in the art that the present application may be practiced in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, devices, circuits, and methods are omitted so as not to obscure the description of the present application with unnecessary detail.
应当理解,在本申请说明书和所附权利要求书中使用的术语“和/或”是指相关联列出的项中的一个或多个的任何组合以及所有可能组合,并且包括这些组合。另外,在本申请说明书和所附权利要求书的描述中,术语“第一”、“第二”、“第三”等仅用于区分描述,而不能理解为指示或暗示相对重要性。It should be understood that the term "and/or" used in the description of the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. In addition, in the description of the specification and appended claims of the present application, the terms "first", "second", "third" and so on are only used to distinguish descriptions, and should not be understood as indicating or implying relative importance.
还应当理解,在本申请说明书中描述的参考“一个实施例”或“一些实施例”等意味着在本申请的一个或多个实施例中包括结合该实施例描述的特定特征、结构或特点。由此,在本说明书中的不同之处出现的语句“在一个实施例中”、“在一些实施例中”、“在其他一些实施例中”、“在另外一些实施例中”等不是必然都参考相同的实施例,而是意味着“一个或多个但不是所有的实施例”,除非是以其他方式另外特别强调。术语“包括”、“包含”、“具有”及它们的变形都意味着“包括但不限于”,除非是以其他方式另外特别强调。It should also be understood that references to "one embodiment" or "some embodiments" or the like described in the specification of the present application mean that a particular feature, structure or characteristic described in connection with the embodiment is included in one or more embodiments of the present application . Thus, appearances of the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in other embodiments," etc. in various places in this specification are not necessarily All refer to the same embodiment, but mean "one or more but not all embodiments" unless specifically stated otherwise. The terms "including", "comprising", "having" and variations thereof mean "including but not limited to", unless specifically stated otherwise.
目前,数据编码问题是DNA数据存储技术中的核心技术问题,尤其针对大规模数据来说,缺乏有效的数据编码方法实现大规模数据存储。有鉴于此,本申请提供了一种DNA数据的存储方法,该方法在进行数据编码的过程中,先将待存储数据的二进制序列转换成对应的碱基序列之后,通过两次分割,将碱基序列分割为S个序列单元,且每个序列单元包括多个分割的序列片段,其中,S个所述序列单元共含有K个所述序列片段。然后利用预设的索引信息对K个所述序列片段和S个所述序列单元进行标记。最后将K个标记序列片段合成DNA分子后存储。基于本申请提供的数据编码方式,能够有效实现DNA数据存储。At present, data encoding is the core technical issue in DNA data storage technology, especially for large-scale data, there is a lack of effective data encoding methods to achieve large-scale data storage. In view of this, the present application provides a method for storing DNA data. In the process of data encoding, the method first converts the binary sequence of the data to be stored into the corresponding base sequence, and then divides the base sequence twice. The base sequence is divided into S sequence units, and each sequence unit includes a plurality of segmented sequence fragments, wherein the S sequence units contain K total of the sequence fragments. Then use the preset index information to mark the K sequence fragments and the S sequence units. Finally, the K marker sequence fragments are synthesized into DNA molecules and then stored. Based on the data encoding method provided by this application, DNA data storage can be effectively realized.
进一步的,本申请提供的数据编码方式,在数据存储量上具有明显的优势。Furthermore, the data encoding method provided by the present application has obvious advantages in terms of data storage capacity.
该存储方法通过图1所示的DNA数据存储装置实现。DNA数据存储装置包括数据处理模块、DNA分子合成模块、DNA分子存储模块、DNA分子测序模块。The storage method is realized by the DNA data storage device shown in FIG. 1 . The DNA data storage device includes a data processing module, a DNA molecule synthesis module, a DNA molecule storage module, and a DNA molecule sequencing module.
其中,数据处理模块用于实现数据编解码。例如,将待存储的数据转换为二进制信息,在按照预设的二进制数据与碱基的对应关系,把二进制信息转换成碱基序列。之后再按照预设的索引信息对碱基序列进行编码,得到最终用于生成DNA分子的碱基序列。Wherein, the data processing module is used to implement data encoding and decoding. For example, the data to be stored is converted into binary information, and the binary information is converted into a base sequence according to the preset correspondence between binary data and bases. Afterwards, the base sequence is encoded according to the preset index information to obtain the base sequence that is finally used to generate the DNA molecule.
DNA分子合成模块用于根据编码好的碱基序列合成DNA分子。DNA分子存储模块能够存放DNA分子。DNA分子测序模块用于将DNA分子翻译成碱基序列。相应的,数据处理模块也可以根据索引信息对DNA分子中测序得到的碱基序列进行解码,并通过数据转换得到DNA分子中存储的数据。The DNA molecule synthesis module is used to synthesize DNA molecules according to the encoded base sequence. The DNA molecule storage module can store DNA molecules. The DNA molecular sequencing module is used to translate DNA molecules into base sequences. Correspondingly, the data processing module can also decode the base sequence obtained by sequencing in the DNA molecule according to the index information, and obtain the data stored in the DNA molecule through data conversion.
在本申请实施例中,DNA数据存储装置可以是一个完整的DNA数据存储设备,即由多个功能模块集成的一个设备。该设备能够实现数据编码、DNA分子合成、DNA分子存储、DNA分子测序以及数据解码的完整流程。In the embodiment of the present application, the DNA data storage device may be a complete DNA data storage device, that is, a device integrated with multiple functional modules. The device can realize the complete process of data encoding, DNA molecule synthesis, DNA molecule storage, DNA molecule sequencing and data decoding.
在另一个实施例中,DNA数据存储装置也可以是由各个独立的设备构成系统。In another embodiment, the DNA data storage device may also be a system composed of individual devices.
例如,其中,数据处理模块可以是电脑、服务器、机器人等计算机设备。用于实现数据编解码。DNA分子合成模块可以DNA分子合成仪,用于根据编码好的碱基序列合成DNA分子。DNA分子存储模块可以是DNA硬盘,能够存放DNA分子。DNA分子测序模块可以是DNA分子测序仪,能够实现DNA分子测序功能。For example, the data processing module may be computer equipment such as computers, servers, and robots. Used to implement data encoding and decoding. The DNA molecular synthesis module can be a DNA molecular synthesizer, which is used to synthesize DNA molecules according to the coded base sequence. The DNA molecule storage module may be a DNA hard disk capable of storing DNA molecules. The DNA molecular sequencing module may be a DNA molecular sequencer, capable of realizing the function of DNA molecular sequencing.
图2为本申请实施例提供的一种DNA数据的存储方法的实现流程示意图,具体包括:Fig. 2 is a schematic diagram of the implementation flow of a DNA data storage method provided by the embodiment of the present application, specifically including:
S101.获取与待存储数据的二进制序列对应的碱基序列。S101. Obtain the base sequence corresponding to the binary sequence of the data to be stored.
该步骤中,获取与待存储数据的二进制序列对应的碱基序列是指,将待存储数据的二进制序列,转换为由A、T、C、G编码形成的、存储有数据信息的碱基序列。In this step, obtaining the base sequence corresponding to the binary sequence of the data to be stored refers to converting the binary sequence of the data to be stored into a base sequence formed by A, T, C, and G codes and storing data information .
在一些实施例中,获取与待存储数据的二进制序列对应的碱基序列包括:In some embodiments, obtaining the base sequence corresponding to the binary sequence of the data to be stored includes:
S111.提取待存储数据对应的二进制序列。S111. Extract the binary sequence corresponding to the data to be stored.
待存储数据为可以在终端设备中存在的任何数据信息,可以包括文字、图片、声音、视频、软件、程序等信息,但不限于此。The data to be stored is any data information that may exist in the terminal device, and may include text, pictures, sound, video, software, programs and other information, but is not limited thereto.
在提取待存储数据对应的二进制序列时,可以获取该待存储数据对应的编码信息,将对应的编码信息转换为二进制的编码信息,从而得到对应的二进制序列。比如,可以将文本信息中的文字转换为对应的ASCII(英文全称为American Standard Code forInformation Interchange,中文全称为美国标准信息交换码)编码,UNICODE(英文全称为Universal Character Set,中文全称为通用字符集)编码,然后将编码信息转换为二进制序列。When extracting the binary sequence corresponding to the data to be stored, the coding information corresponding to the data to be stored may be obtained, and the corresponding coding information is converted into binary coding information, thereby obtaining the corresponding binary sequence. For example, the text in the text information can be converted into the corresponding ASCII (English full name is American Standard Code for Information Interchange, Chinese full name is American Standard Code for Information Interchange) encoding, UNICODE (English full name is Universal Character Set, Chinese full name is Universal Character Set) ) encoding, and then convert the encoded information into a binary sequence.
示例性的,提取文本“春,已不再是想象之外的那只蝴蝶。”对应的二进制序列为“11100110 10011000 10100101 11101111 10111100 10001100 11100101 1011011110110010 11100100 10111000 10001101 11100101 10000110 10001101 1110011010011000 10101111 11100110 10000011 10110011 11101000 10110001 1010000111100100 10111001 10001011 11100101 10100100 10010110 11100111 1001101010000100 11101001 10000010 10100011 11100101 10001111 10101010 1110100010011101 10110100 11101000 10011101 10110110 11100011 10000000 10000010”。示例性的,提取文本“春,已不再是想象之外的那只蝴蝶。”对应的二进制序列为“11100110 10011000 10100101 11101111 10111100 10001100 11100101 1011011110110010 11100100 10111000 10001101 11100101 10000110 10001101 1110011010011000 10101111 11100110 10000011 10110011 11101000 10110001 1010000111100100 10111001 10001011 11100101 10100100110110 11100111 10011011010010010010010010000010 10100011 11100101 100011111010101010100100100100100100100100100110100110110110110110110110110110110110110110110110110110110 1110110 1110110 11100100011110 1110111 100100011110 111110 111100 111100 111100 111100 that
S121.根据预设的映射规则,将二进制序列转换为碱基序列。S121. Convert the binary sequence into a base sequence according to a preset mapping rule.
本申请实施例中,预设的映射规则是指预设的二进制与碱基之间的映射规则。根据二进制代码与碱基A、T、C、G之间预设的映射规则,将有0/1组成的二进制序列转换为由碱基A、T、C、G编码形成的、存储有数据信息的碱基序列。示例性的,二进制序列与碱基A、T、C、G之间的预设对应关系为:一个碱基A代表一个00,一个碱基T代表一个01,一个碱基C代表一个10,一个碱基G代表一个11。当二级制序列为00110110101100101011011011000011001001时,根据二进制与碱基A、T、C、G之间预设的映射规则,将二进制数据信息转换碱基序列为AGTCCGACCGTCGAAGACT的DNA序列。当然,二进制代码与碱基A、T、C、G之间的预设对应关系不限于上述示例,例如,也可以为:一个碱基T代表一个00,一个碱基A代表一个01,一个碱基G代表一个10,一个碱基C代表一个11,但不限于此。应当理解,二进制序列与碱基A、T、C、G之间预设的映射规则,只需要能够根据预设的映射规则将二进制序列转换碱基序列就行,并不限于上述示例。In the embodiment of the present application, the preset mapping rule refers to the preset mapping rule between binary and base. According to the preset mapping rules between the binary code and the bases A, T, C, G, the binary sequence consisting of 0/1 is converted into a code formed by the bases A, T, C, G, and stored with data information base sequence. Exemplarily, the preset correspondence between the binary sequence and bases A, T, C, and G is as follows: a base A represents a 00, a base T represents a 01, a base C represents a 10, and a base C represents a 10. The base G represents an 11. When the binary sequence is 00110110101100101011011011000011001001, according to the preset mapping rules between binary and bases A, T, C, G, the binary data information is converted into a base sequence into a DNA sequence of AGTCCGACCGTCGAAGACT. Of course, the preset correspondence between the binary code and the bases A, T, C, and G is not limited to the above example, for example, it can also be: a base T represents a 00, a base A represents a 01, and a base The base G represents a 10, and the base C represents an 11, but not limited thereto. It should be understood that the preset mapping rules between binary sequences and bases A, T, C, and G only need to be able to convert binary sequences into base sequences according to the preset mapping rules, and are not limited to the above examples.
示例性的,二进制与碱基之间的映射规则为:一个碱基A代表一个11,一个碱基T代表一个10,一个碱基C代表一个01,一个碱基G代表一个00。根据该映射规则,可以将文本“春,已不再是想象之外的那只蝴蝶。”对应的二进制序列“11100110 10011000 1010010111101111 10111100 10001100 11100101 10110111 10110010 11100100 1011100010001101 11100101 10000110 10001101 11100110 10011000 10101111 1110011010000011 10110011 11101000 10110001 10100001 11100100 10111001 1000101111100101 10100100 10010110 11100111 10011010 10000100 11101001 1000001010100011 11100101 10001111 10101010 11101000 10011101 10110100 1110100010011101 10110110 11100011 10000000 10000010”转换为碱基序列“ATCT TCTG TTCCATAA TAAG TGAG ATCC TACA TAGT ATCG TATG TGAC ATCC TGCT TGAC ATCT TCTG TTAAATCT TGGA TAGA ATTG TAGC TTGC ATCG TATC TGTA ATCC TTCG TCCT ATCA TCTT TGCGATTC TGGT TTGA ATCC TGAA TTTT ATTG TCAC TACG ATTG TCAC TACT ATGA TGGG TGGT”。Exemplarily, the mapping rule between binary and base is: a base A represents a 11, a base T represents a 10, a base C represents a 01, and a base G represents a 00.根据该映射规则,可以将文本“春,已不再是想象之外的那只蝴蝶。”对应的二进制序列“11100110 10011000 1010010111101111 10111100 10001100 11100101 10110111 10110010 11100100 1011100010001101 11100101 10000110 10001101 11100110 10011000 10101111 1110011010000011 10110011 11101000 10110001 10100001 11100100 10111001 1000101111100101 10100100 10010110 11100111 10011010 10000100 11101001 1000001010100011 11100101 10001111 10101010 11101000 10011101 10110100 1110100010011101 10110110 11100011 10000000 10000010”转换为碱基序列“ATCT TCTG TTCCATAA TAAG TGAG ATCC TACA TAGT ATCG TATG TGAC ATCC TGCT TGAC ATCT TCTG TTAAATCT TGGA TAGA ATTG TAGC TTGC ATCG TATC TGTA ATCC TTCG TCCT ATCA TCTT TGCGATTC TGGT TTGA ATCC TGAA TTTT ATTG TCAC TACG ATTG TCAC TACT ATGA TGGG TGGT".
S102.对碱基序列进行分割,得到S个序列单元,且每个序列单元包括多个分割的序列片段,其中,S个序列单元共含有K个序列片段,且序列片段的长度为n,n、S和K均为大于或者等于2的整数。S102. Segment the base sequence to obtain S sequence units, and each sequence unit includes a plurality of segmented sequence fragments, wherein, the S sequence units contain K sequence fragments in total, and the length of the sequence fragments is n, n , S and K are all integers greater than or equal to 2.
本申请实施例中,对碱基序列进行分割,经过分割后的碱基序列变成s个序列单元,且每个序列单元包括多个分割的序列片段,从而降低了序列长度,便于后续步骤对序列单元进行分别存储。应当理解的是,本申请实施例所指的长度,是指碱基长度,可以理解为碱基个数。In the embodiment of the present application, the base sequence is segmented, and the segmented base sequence becomes s sequence units, and each sequence unit includes a plurality of segmented sequence fragments, thereby reducing the sequence length and facilitating subsequent steps to Sequence units are stored separately. It should be understood that the length referred to in the embodiments of the present application refers to the base length, which can be understood as the number of bases.
碱基序列分割后得到的每个序列单元包括长度为n的多个序列片段,因此,序列单元的长度为序列片段的整数倍。在一些实施例中,经分割后形成的序列单元的长度相同,即碱基序列分割成s个长度相同的序列单元,且s个长度相同的序列单元同时含有相同数量、且长度为n的序列片段。在一些实施例中,经分割后形成的序列单元的长度不同,示例性的,按照预设的序列单元长度对碱基序列依次进行分割,得到S-1个序列单元,最后一次分割后剩余的碱基序列长度不足预设长度,此时,剩余碱基序列作为一个序列单元,其长度小于其他序列单元的长度。在一些实施例中,碱基序列可以按照其他预设的分割规则进行分割,得到序列单元长度不一致的s个序列单元。Each sequence unit obtained after base sequence segmentation includes multiple sequence fragments with a length of n, therefore, the length of the sequence unit is an integer multiple of the sequence fragments. In some embodiments, the sequence units formed after segmentation have the same length, that is, the base sequence is divided into s sequence units of the same length, and the s sequence units of the same length contain the same number of sequences of length n fragment. In some embodiments, the lengths of the sequence units formed after segmentation are different. Exemplarily, the base sequence is sequentially segmented according to the preset sequence unit length to obtain S-1 sequence units, and the remaining The length of the base sequence is less than the preset length. At this time, the remaining base sequence is regarded as a sequence unit, and its length is less than the length of other sequence units. In some embodiments, the base sequence can be segmented according to other preset segmentation rules to obtain s sequence units with different sequence unit lengths.
在一种可能的实现方式中,对碱基序列进行分割,得到s个序列单元,每个序列单元包括多个分割的序列片段,包括:In a possible implementation manner, the base sequence is segmented to obtain s sequence units, and each sequence unit includes a plurality of segmented sequence fragments, including:
将碱基序列分割为多个序列单元,且序列单元的长度为n的整数倍;Divide the base sequence into multiple sequence units, and the length of the sequence unit is an integer multiple of n;
将每个序列单元分割成多个长度为n的序列片段。Divide each sequence unit into multiple sequence fragments of length n.
在一些实施例中,将碱基序列分割为多个序列单元时,从碱基序列的一端开始,每隔一个预设的序列单元长度对碱基序列进行一次分割,得到s个序列单元,其中,预设长度为n的整数倍。当最后一次分割后剩余的碱基序列长度不足预设长度时,将剩余碱基序列作为一个序列单元。在其他实施例中,也可以从碱基序列的其他位点开始,对碱基序列进行分割。如,从碱基序列的中间位点开始,同时朝两端依次对碱基序列进行分割,得到长度为n的整数倍的序列单元。In some embodiments, when the base sequence is divided into multiple sequence units, starting from one end of the base sequence, the base sequence is divided every other preset sequence unit length to obtain s sequence units, wherein , the default length is an integer multiple of n. When the length of the remaining base sequence after the last division is less than the preset length, the remaining base sequence is regarded as a sequence unit. In other embodiments, the base sequence can also be segmented starting from other positions in the base sequence. For example, starting from the middle position of the base sequence, the base sequence is sequentially divided towards both ends at the same time, to obtain sequence units whose length is an integer multiple of n.
在一些实施例中,将每个序列单元分割成多个长度为n的序列片段,包括:In some embodiments, each sequence unit is divided into a plurality of sequence fragments of length n, including:
从序列单元的一端开始,每隔长度n对序列单元进行一次分割,得到多个序列片段。在其他实施例中,也可以从序列单元的其他位点开始,对序列单元进行分割。如,从序列单元的中间位点开始,同时朝两端依次对序列单元进行分割,得到长度为l的序列片段。Starting from one end of the sequence unit, the sequence unit is divided every length n to obtain multiple sequence fragments. In other embodiments, the sequence unit can also be segmented starting from other positions of the sequence unit. For example, starting from the middle position of the sequence unit, the sequence unit is sequentially segmented towards both ends at the same time to obtain a sequence fragment with a length of 1.
示例性的,如图3所示,S101中的碱基序列中,按照序列单元长度为60个碱基的标准,从碱基序列的一端开始对碱基序列依次进行分割,得到3组长度为60个碱基的序列单元,以及一组长度为12个碱基的序列单元;按照序列片段长度为4个碱基的标准,从序列单元的一端开始对序列单元依次进行分割,三组长度为60个碱基的序列单元分别分割成15组序列片段,长度为12个碱基的序列单元分割成3组序列片段。Exemplarily, as shown in Figure 3, in the base sequence in S101, according to the standard that the length of the sequence unit is 60 bases, the base sequence is sequentially segmented from one end of the base sequence, and three groups of lengths are obtained. A sequence unit of 60 bases, and a group of sequence units with a length of 12 bases; according to the standard of a sequence fragment length of 4 bases, the sequence units are sequentially divided from one end of the sequence unit, and the length of the three groups is The sequence unit of 60 bases was divided into 15 groups of sequence fragments, and the sequence unit of 12 bases was divided into 3 groups of sequence fragments.
在另一种可能的实现方式中,对碱基序列进行分割,得到s个序列单元,每个序列单元包括长度为n的多个序列片段,包括:In another possible implementation, the base sequence is segmented to obtain s sequence units, and each sequence unit includes multiple sequence fragments with a length of n, including:
将碱基序列分割为多个长度为n的序列片段;Divide the base sequence into multiple sequence fragments of length n;
按照预设的组合规则将序列片段进行组合,得到s个序列单元。The sequence fragments are combined according to a preset combination rule to obtain s sequence units.
在一些实施例中,将碱基序列分割为多个长度为n的序列片段,从碱基序列的一端开始,每隔长度n对碱基序列进行一次分割,得到多个序列片段。In some embodiments, the base sequence is divided into multiple sequence fragments with a length of n, starting from one end of the base sequence, the base sequence is divided every length n to obtain multiple sequence fragments.
在一些实施例中,按照预设的组合规则将序列片段进行组合时,预设的组合规则是指将多个序列片段归属为一个序列单元的规则,包括归属为一个序列单元的序列片段在碱基序列中的位置,序列片段的数量以及序列片段组合成序列单元时序列片段的排布顺序。组合成的序列单元的长度可以相同,也可以不同,但均为l的整数倍。示例性的,按序列片段在碱基序列中的顺序,依次将20个碱基数量为6的序列片段进行组合,形成一个序列单元。In some embodiments, when the sequence fragments are combined according to the preset combination rules, the preset combination rules refer to the rules for assigning multiple sequence fragments as a sequence unit, including the sequence fragments assigned to a sequence unit in the base The position in the base sequence, the number of sequence fragments and the arrangement order of the sequence fragments when the sequence fragments are combined into a sequence unit. The lengths of the combined sequence units can be the same or different, but they are all integer multiples of 1. Exemplarily, according to the order of the sequence fragments in the base sequence, 20 sequence fragments with a number of 6 bases are sequentially combined to form a sequence unit.
示例性的,序列片段的长度为4,序列单元包括15个序列片段。此时,步骤S101的碱基序列“ATCT TCTG TTCC ATAA TAAG TGAG ATCC TACA TAGT ATCG TATG TGAC ATCC TGCTTGAC ATCT TCTG TTAA ATCT TGGA TAGA ATTG TAGC TTGC ATCG TATC TGTA ATCC TTCGTCCT ATCA TCTT TGCG ATTC TGGT TTGA ATCC TGAA TTTT ATTG TCAC TACG ATTG TCACTACT ATGA TGGG TGGT”分割得到48个序列片段,分别为:ATCT、TCTG、TTCC、ATAA、TAAG、TGAG、ATCC、TACA、TAGT、ATCG、TATG、TGAC、ATCC、TGCT、TGAC、ATCT、TCTG、TTAA、ATCT、TGGA、TAGA、ATTG、TAGC、TTGC、ATCG、TATC、TGTA、ATCC、TTCG、TCCT、ATCA、TCTT、TGCG、ATTC、TGGT、TTGA、ATCC、TGAA、TTTT、ATTG、TCAC、TACG、ATTG、TCAC、TACT、ATGA、TGGG、TGGT;根据设定预设的组合规则和序列单元的长度,依次将各序列片段进行组合,且每15个序列片段组合成一个序列单元,得到四个序列单元,分别包含如下序列片段。第一个序列单元包括如下15个序列片段:ATCT、TCTG、TTCC、ATAA、TAAG、TGAG、ATCC、TACA、TAGT、ATCG、TATG、TGAC、ATCC、TGCT、TGAC;第二个序列单元包括如下15个序列片段:ATCT、TCTG、TTAA、ATCT、TGGA、TAGA、ATTG、TAGC、TTGC、ATCG、TATC、TGTA、ATCC、TTCG、TCCT;第三个序列单元包括如下15个序列片段:ATCA、TCTT、TGCG、ATTC、TGGT、TTGA、ATCC、TGAA、TTTT、ATTG、TCAC、TACG、ATTG、TCAC、TACT;第四个序列单元包括如下3个序列片段:ATGA、TGGG、TGGT。Exemplarily, the length of the sequence segment is 4, and the sequence unit includes 15 sequence segments. At this time, the base sequence of step S101 "ATCT TCTG TTCC ATAA TAAG TGAG ATCC TACA TAGT ATCG TATG TGAC ATCC TGCTTGAC ATCT TCTG TTAA ATCT TGGA TAGA ATTG TAGC TTGC ATCG TATC TGTA ATCC TTCGTCCT ATCA TCTT TGCG ATTC TGGT TTGACTTG ATCC TGAA TT TTC A ATTG TCACTACT ATGA TGGG TGGT” to obtain 48 sequence fragments, which are: ATCT, TCTG, TTCC, ATAA, TAAG, TGAG, ATCC, TACA, TAGT, ATCG, TATG, TGAC, ATCC, TGCT, TGAC, ATCT, TCTG, TTAA, ATCT, TGGA, TAGA, ATTG, TAGC, TTGC, ATCG, TATC, TGTA, ATCC, TTCG, TCCT, ATCA, TCTT, TGCG, ATTC, TGGT, TTGA, ATCC, TGAA, TTTT, ATTG, TCAC, TACG, ATTG, TCAC, TACT, ATGA, TGGG, TGGT; according to the preset combination rules and the length of the sequence unit, each sequence fragment is combined in turn, and every 15 sequence fragments are combined into a sequence unit to obtain four sequences The unit contains the following sequence fragments respectively. The first sequence unit includes the following 15 sequence fragments: ATCT, TCTG, TTCC, ATAA, TAAG, TGAG, ATCC, TACA, TAGT, ATCG, TATG, TGAC, ATCC, TGCT, TGAC; the second sequence unit includes the following 15 sequence fragments: ATCT, TCTG, TTAA, ATCT, TGGA, TAGA, ATTG, TAGC, TTGC, ATCG, TATC, TGTA, ATCC, TTCG, TCCT; the third sequence unit includes the following 15 sequence fragments: ATCA, TCTT, TGCG, ATTC, TGGT, TTGA, ATCC, TGAA, TTTT, ATTG, TCAC, TACG, ATTG, TCAC, TACT; the fourth sequence unit includes the following three sequence fragments: ATGA, TGGG, TGGT.
本申请实施例中,对碱基序列进行分割后,得到S个序列单元,S个序列单元共包含K个长度为n的序列片段。In the embodiment of the present application, after the base sequence is segmented, S sequence units are obtained, and the S sequence units contain K sequence fragments with a length of n in total.
S103.利用预设的索引信息对K个序列片段和S个序列单元进行标记,得到K个标记序列片段和S个标记序列单元,其中,索引信息包括用于表示S个序列单元在所述碱基序列中的排列顺序的第一检索序列,和用于表示属于同一序列单元中的多个序列片段在序列单元中的排列顺序的第二检索序列,K个标记序列片段用于合成存储有目标数据的K个第一DNA分子。S103. Use the preset index information to mark the K sequence fragments and S sequence units to obtain K sequence fragments and S sequence units, wherein the index information includes information used to indicate that the S sequence units are in the base The first search sequence of the arrangement order in the base sequence, and the second search sequence used to indicate the arrangement order of multiple sequence fragments belonging to the same sequence unit in the sequence unit, K marker sequence fragments are used to synthesize and store the target The K first DNA molecules of the data.
本申请实施例中,采用预设的索引信息对序列单元进行标记,以方便后续DNA存储数据的解码。采用预设的索引信息对序列单元进行标记,包括:利用第一检索序列对S个序列单元在所述碱基序列中的排列顺序进行标记,以及利用第二检索序列对属于同一序列单元中的多个序列片段在序列单元中的排列顺序进行标记,得到K个标记序列片段和S个标记序列单元。在这种情况下,碱基序列中的S个序列单元的顺序被记录,同时,属于同一序列单元中的多个序列片段的顺序也被记录下来。In the embodiment of the present application, the preset index information is used to mark the sequence unit, so as to facilitate the decoding of subsequent DNA storage data. Using the preset index information to mark the sequence units includes: using the first search sequence to mark the arrangement order of the S sequence units in the base sequence, and using the second search sequence to mark the sequence units belonging to the same sequence unit The order in which the multiple sequence fragments are arranged in the sequence unit is marked to obtain K marked sequence fragments and S marked sequence units. In this case, the sequence of S sequence units in the base sequence is recorded, and at the same time, the sequence of multiple sequence fragments belonging to the same sequence unit is also recorded.
本申请实施例中,第二检索序列为碱基形成的序列,可通过预先设定的规则来确定第二检索序列的碱基序列。示例性的,当序列单元中的序列片段的数量小于或等于4时,可以采用单碱基来表示第二检索序列。如:序号1对应双碱基A,序号2对应双碱基C,序号3对应双碱基G,序号4对应双碱基T,当然,序号和碱基之间不限于这种对应方式。例性的,当序列单元中的序列片段的数量小于或等于16时,可以采用双碱基来表示第二检索序列。如:序号1对应双碱基AA,序号2对应双碱基AC,序号3对应双碱基AG,序号4对应双碱基AT,序号5对应双碱基CA,序号6对应双碱基CC,序号7对应双碱基CG,序号8对应双碱基CT,序号9对应双碱基GA,序号10对应双碱基GC,序号11对应双碱基GG,序号12对应双碱基GT,序号13对应双碱基TA,序号14对应双碱基TC,序号15对应双碱基TG,序号16对应双碱基TT...。当然,基准碱基组中的碱基数量并不限于2,当序列单元中的序列片段的数量增加时,采用的第二检索序列中的碱基数量对应增加,如当序列单元中的序列片段的数量小于或等于64时,可以采用三碱基来表示第二检索序列。按照第二检索序列中4的碱基数量次方大于或等于序列单元中的序列片段的数量的规则,以此类推。In the embodiment of the present application, the second search sequence is a sequence formed of bases, and the base sequence of the second search sequence can be determined according to preset rules. Exemplarily, when the number of sequence fragments in the sequence unit is less than or equal to 4, a single base can be used to represent the second search sequence. For example: the sequence number 1 corresponds to the double base A, the sequence number 2 corresponds to the double base C, the sequence number 3 corresponds to the double base G, and the sequence number 4 corresponds to the double base T. Of course, the correspondence between the sequence number and the base is not limited to this method. Exemplarily, when the number of sequence fragments in the sequence unit is less than or equal to 16, two bases may be used to represent the second search sequence. For example: sequence number 1 corresponds to the two bases AA, sequence number 2 corresponds to the two bases AC, sequence number 3 corresponds to the two bases AG, sequence number 4 corresponds to the two bases AT, sequence number 5 corresponds to the two bases CA, and sequence number 6 corresponds to the two bases CC. The sequence number 7 corresponds to the two-base CG, the sequence number 8 corresponds to the two-base CT, the sequence number 9 corresponds to the two-base GA, the sequence number 10 corresponds to the two-base GC, the sequence number 11 corresponds to the two-base GG, the sequence number 12 corresponds to the two-base GT, and the sequence number 13 Corresponding to the two-base TA, the sequence number 14 corresponds to the two-base TC, the sequence number 15 corresponds to the two-base TG, and the sequence number 16 corresponds to the two-base TT.... Of course, the number of bases in the reference base group is not limited to 2. When the number of sequence fragments in the sequence unit increases, the number of bases in the second search sequence used increases correspondingly, such as when the sequence fragments in the sequence unit When the number of is less than or equal to 64, three bases can be used to represent the second search sequence. According to the rule that the power of 4 bases in the second search sequence is greater than or equal to the number of sequence fragments in the sequence unit, and so on.
本申请实施例中,采用第二检索序列标记序列片段时,可以在序列片段的特定位置拼接第二检索序列。在一些实施例中,利用第二检索序列标记属于同一序列单元中的多个序列片段的方式,包括:在序列片段的任一侧拼接第二检索序列。示例性的,在序列片段的起始端即左端拼接第二检索序列;或,在序列片段的终止端即右端拼接第二检索序列。在一个实施例中,利用第二检索序列标记属于同一序列单元中的多个序列片段的方式,包括:在序列片段的两侧同时拼接碱基组,两侧的碱基组形成第二检索序列。In the embodiment of the present application, when the second search sequence is used to mark the sequence fragment, the second search sequence can be spliced at a specific position of the sequence fragment. In some embodiments, using the second search sequence to mark multiple sequence fragments belonging to the same sequence unit includes: splicing the second search sequence on either side of the sequence fragment. Exemplarily, the second search sequence is spliced at the start end, that is, the left end of the sequence fragment; or, the second search sequence is spliced at the end end, that is, the right end of the sequence fragment. In one embodiment, the method of using the second search sequence to mark multiple sequence fragments belonging to the same sequence unit includes: simultaneously splicing base groups on both sides of the sequence fragment, and the base groups on both sides form the second search sequence .
本申请实施例中,K个序列片段经过第二检索序列标记,形成K个标记序列片段,又称为信息序列片段。示例性的,序列片段ATGC前标记第二检索序列AA后,形成AAATGC的标记序列片段。In the embodiment of the present application, K sequence fragments are marked by the second search sequence to form K marked sequence fragments, which are also called information sequence fragments. Exemplarily, after the second search sequence AA is marked before the sequence fragment ATGC, a marked sequence fragment of AAATGC is formed.
本申请实施例中,通过采用第一检索序列用来标示S个序列单元在碱基序列中的位置。在一些实施例中,第一检索序列包括i条DNA序列片段,i为大于或等于1的整数,即第一检索序列可以是一条DNA序列片段,也可以是多条DNA序列片段。其中,每条DNA序列片段包括用作索引标志的第一碱基序列和用于标示序列单元编号的第二碱基序列。其中,第一碱基序列可以根据预先设置置于第一检索序列的特定位置中,示例性的,第一碱基序列位于第一检索序列的起始段(左端);示例性的,第一碱基序列位于第一检索序列的终止段(右端);示例性的,第一碱基序列位于第一检索序列中特定的位置,如第一检索序列的第三位和第四位碱基,不限于此。In the embodiment of the present application, the position of the S sequence units in the base sequence is marked by using the first search sequence. In some embodiments, the first search sequence includes i DNA sequence fragments, where i is an integer greater than or equal to 1, that is, the first search sequence may be one DNA sequence fragment or multiple DNA sequence fragments. Wherein, each DNA sequence fragment includes a first base sequence used as an index mark and a second base sequence used to indicate the sequence unit number. Wherein, the first base sequence can be placed in a specific position of the first search sequence according to preset settings. Exemplarily, the first base sequence is located at the initial segment (left end) of the first search sequence; exemplary, the first The base sequence is located at the termination segment (right end) of the first search sequence; exemplary, the first base sequence is located at a specific position in the first search sequence, such as the third and fourth bases of the first search sequence, Not limited to this.
第一碱基序列可以预先设定。示例性的,TT作为第一碱基序列置于第一检索序列的起始段,用作索引标志,表示以序列单元中以TT开头的DNA序列片段第一检索序列。该示例中,当第一碱基序列位于第一检索序列的起始段时,第一碱基序列与第二检索序列的起始碱基序列不同,以避免识别过程中,误将序列片段识别为第一检索序列。The first base sequence can be set in advance. Exemplarily, TT is placed as the first base sequence at the initial segment of the first search sequence, and is used as an index mark, indicating the first search sequence of the DNA sequence fragment starting with TT in the sequence unit. In this example, when the first base sequence is located at the initial segment of the first search sequence, the first base sequence is different from the initial base sequence of the second search sequence, so as to avoid misidentifying sequence fragments during the identification process is the first search sequence.
同样的,第二碱基序列也可以预先设定。示例性的,序号1对应四碱基AAAA,序号2对应四碱基AAAC,序号3对应四碱基AAAG,序号4对应四碱基AAAT,当然,序号和碱基之间不限于这种对应方式,基准碱基组中的碱基数量也不限于4。本申请实施例中,序列单元经过第一检索序列标记,且序列单元中的各序列片段经过第二检索序列标记,得到标记后的序列单元,又称为形成信息序列单元。参考图4,将步骤S102中的序列单元(图4箭头左侧所示)经过第一检索序列标记,且序列单元中的各序列片段经过第二检索序列标记后,得到分别包含多个信息序列片段的四组标记序列单元(图4箭头右侧所示)。Similarly, the second base sequence can also be preset. Exemplarily, the sequence number 1 corresponds to the four-base AAAA, the sequence number 2 corresponds to the four-base AAAC, the sequence number 3 corresponds to the four-base AAAG, and the sequence number 4 corresponds to the four-base AAAT. Of course, the correspondence between the sequence number and the base is not limited to this , and the number of bases in the reference base set is not limited to four. In the embodiment of the present application, the sequence unit is marked by the first search sequence, and each sequence fragment in the sequence unit is marked by the second search sequence, and the marked sequence unit is obtained, which is also called the formation information sequence unit. Referring to Fig. 4, after the sequence unit in step S102 (shown on the left side of the arrow in Fig. 4) is marked with the first search sequence, and each sequence segment in the sequence unit is marked with the second search sequence, a sequence containing multiple information is obtained respectively. Four sets of marker sequence units for the fragment (shown to the right of the arrow in Figure 4).
在一些实施例中,第一检索序列和第二检索序列对应的DNA序列片段利用DNA合成技术获得。示例性的,DNA合成技术包括但不限于酶法合成、亚磷酰胺合成等。In some embodiments, the DNA sequence fragments corresponding to the first search sequence and the second search sequence are obtained using DNA synthesis technology. Exemplarily, DNA synthesis techniques include, but are not limited to, enzymatic synthesis, phosphoramidite synthesis, and the like.
在一些实施例中,第一检索序列和第二检索序列对应的DNA序列片段可以从预先合成的DNA通用分子库中扩增获得,比如PCR技术等。In some embodiments, the DNA sequence fragments corresponding to the first search sequence and the second search sequence can be amplified from a pre-synthesized DNA universal molecular library, such as PCR technology.
上述步骤S101至步骤S103通过图1所示装置中的处理模块实现。The above steps S101 to S103 are implemented by the processing module in the device shown in FIG. 1 .
在一些实施例中,存储方法还包括:In some embodiments, the storage method also includes:
S104.将K个标记序列片段合成存储有所述目标数据的K个第一DNA分子后,将K个第一DNA分子存储在S个第一物理空间,其中,同属于一个序列单元的标记序列片段对应的第一DNA分子存储在同一个第一物理空间,不属于同一个序列单元的标记序列片段对应的第一DNA分子存储在不同的第一物理空间。S104. After synthesizing K first DNA molecules storing the target data with K marker sequence fragments, store the K first DNA molecules in S first physical spaces, wherein the marker sequences belonging to the same sequence unit First DNA molecules corresponding to the fragments are stored in the same first physical space, and first DNA molecules corresponding to marker sequence fragments that do not belong to the same sequence unit are stored in different first physical spaces.
该步骤中,将K个标记序列片段分别合成存储有所述目标数据的K个第一DNA分子,通过图1所示装置中的DNA合成模块实现。本申请实施例可以通过现有的合成技术,将K个标记序列片段分别合成,得到K个第一DNA分子。In this step, the K marker sequence fragments are respectively synthesized into K first DNA molecules storing the target data, which is realized by the DNA synthesis module in the device shown in FIG. 1 . In the embodiment of the present application, the K marker sequence fragments can be synthesized respectively through the existing synthesis technology to obtain K first DNA molecules.
将K个第一DNA分子存储在S个不同的第一物理空间,并使得同属于一个序列单元的标记序列片段对应的第一DNA分子存储在同一个第一物理空间,不属于同一个序列单元的标记序列片段对应的第一DNA分子存储在不同的第一物理空间,该步骤通过图1所示装置中的DNA存储模块实现。将K个第一DNA分子存储在S个不同的第一物理空间后的DNA存储的示意图如图5所示,其中,每一个小方格表示一个第一物理空间。Store the K first DNA molecules in S different first physical spaces, and store the first DNA molecules corresponding to the marker sequence fragments belonging to one sequence unit in the same first physical space, not belonging to the same sequence unit The first DNA molecules corresponding to the marker sequence fragments are stored in different first physical spaces, and this step is realized by the DNA storage module in the device shown in FIG. 1 . A schematic diagram of DNA storage after storing K first DNA molecules in S different first physical spaces is shown in FIG. 5 , wherein each small square represents a first physical space.
该步骤中,将合成的K个第一DNA分子分别存储于不同的第一物理空间中,实现各信息序列单元的分别存储。进一步的,将S个所述第一物理空间集成在一个DNA硬盘中,实现存储有目标数据的碱基序列的存储。通过集成,S个第一物理空间中的K个第一DNA分子形成一个完整的整体得以保存,且保存过程中不容易出现遗漏而损失信息,从而有利于提高数据保存的完整性。对应的,在解码释放过程中,通过将集成的第一DNA分子解码释放,可以完整地恢复DNA碱基,保持数据的完整性。In this step, the synthesized K first DNA molecules are respectively stored in different first physical spaces, so as to realize the separate storage of each information sequence unit. Further, the S first physical spaces are integrated into one DNA hard disk to realize the storage of the base sequence storing the target data. Through the integration, the K first DNA molecules in the S first physical spaces form a complete whole to be preserved, and it is not easy to omit and lose information during the preservation process, which is conducive to improving the integrity of data preservation. Correspondingly, in the process of decoding and releasing, by decoding and releasing the integrated first DNA molecule, the DNA base can be completely recovered and the integrity of the data can be maintained.
在一些实施例中,存储方法还包括:将第二DNA分子存储在与第一物理空间对应的第二物理空间,第二DNA分子存储有索引信息。将存储有索引信息的第二DNA分子存储于与第一物理空间不同的第二物理空间,实现索引信息的保存。In some embodiments, the storage method further includes: storing the second DNA molecule in a second physical space corresponding to the first physical space, where index information is stored in the second DNA molecule. The second DNA molecule storing the index information is stored in a second physical space different from the first physical space to realize preservation of the index information.
本申请实施例提供的DNA数据的存储方法,将待存储数据对应的二进制序列转换为碱基序列后,与目标数据的二进制序列对应的碱基序列分割为S个序列,每个序列单元包括多个分割的序列片段,S个所述序列单元共含有K个所述序列片段,且所述序列片段的长度为n,n、S和K均为大于或者等于2的整数,利用预设的索引信息标记序列片段和序列单元的位置信息,将标记后的序列片段合成为DNA分子后分别存储。通过该方法,可以提升数据信息的存储量,实现大规模的数据信息在DNA中的存储。In the DNA data storage method provided in the embodiment of the present application, after converting the binary sequence corresponding to the data to be stored into a base sequence, the base sequence corresponding to the binary sequence of the target data is divided into S sequences, and each sequence unit includes multiple segmented sequence fragments, the S sequence units contain K sequence fragments in total, and the length of the sequence fragments is n, n, S and K are all integers greater than or equal to 2, using a preset index The information marks the position information of sequence fragments and sequence units, and the marked sequence fragments are synthesized into DNA molecules and stored separately. Through this method, the storage capacity of data information can be increased, and the storage of large-scale data information in DNA can be realized.
在某些具体实施例中,第一检索序列的长度和标记序列片段(带有第二检索序列的序列片段)的长度不同,则通过长度区别从序列单元中区分第一检索序列和标记序列片段。具体的,以m表示标记序列片段的碱基数,以q表示第二检索序列的碱基数,以i表示第一检索序列中DNA序列片段的条数,以p表示第一检索序列中第二碱基序列的碱基数,通过本申请提供的方法,可以实现含有D个碱基数的DNA数据的存储,其中,D的计算公式如下:In some specific embodiments, the length of the first retrieval sequence and the length of the marker sequence fragment (the sequence fragment with the second retrieval sequence) are different, and then the first retrieval sequence and the marker sequence fragment are distinguished from the sequence unit by length difference . Specifically, m represents the base number of the marker sequence fragment, q represents the base number of the second retrieval sequence, i represents the number of DNA sequence fragments in the first retrieval sequence, and p represents the number of DNA sequence fragments in the first retrieval sequence The number of bases in the two-base sequence, through the method provided by this application, can realize the storage of DNA data containing D bases, wherein, the calculation formula of D is as follows:
D=4q×(m-q)×4i×p D = 4q×(mq)× 4i×p
在特定实施例中,m长度为8碱基的情况下,q=4,i=10,p=4,D=256×4×440=1.23×1027。1个碱基存储2bits信息时候,能够存储的信息量L=2.46×1027bits=3.075×1026bytes=3.075×105ZB,远大于目前的数据存储规模。在某些具体实施例中,第一检索序列的长度和标记序列片段(带有第二检索序列的序列片段)的长度相同;第一检索序列中用作索引标志的第一碱基序列,可以是第二检索序列的一部分,和第二检索序列的碱基数相同。在这种情况下,以m表示标记序列片段的碱基数,以q表示第二检索序列的碱基数,以i表示第一检索序列中DNA序列片段的条数,以p表示第一检索序列中第二碱基序列的碱基数,通过本申请的方法,可以实现含有D个碱基数的DNA数据的存储,其中,D的计算公式如下:In a specific embodiment, when the length of m is 8 bases, q=4, i=10, p=4, D=256×4×4 40 =1.23×10 27 . When 1 base stores 2 bits of information, the amount of information that can be stored is L=2.46×10 27 bits=3.075×10 26 bytes=3.075×10 5 ZB, which is much larger than the current data storage scale. In some specific embodiments, the length of the first search sequence is the same as the length of the marker sequence fragment (sequence fragment with the second search sequence); the first base sequence used as an index mark in the first search sequence can be It is a part of the second search sequence and has the same number of bases as the second search sequence. In this case, let m represent the base number of the marker sequence fragment, q represent the base number of the second search sequence, use i to represent the number of DNA sequence fragments in the first search sequence, and p represent the first search sequence The number of bases in the second base sequence in the sequence, through the method of this application, can realize the storage of DNA data containing D bases, wherein, the calculation formula of D is as follows:
D=(4q-i)×(m-q)×4i×p D=( 4q- i)×(mq)× 4i×p
在特定实施例中,m长度为8碱基的情况下,q=4,i=10,p=4,D=(256-10)×4×440=1.18×1027。1个碱基存储2bits信息时候,能够存储的信息量L=2.36×1027bits=2.95×1026bytes=2.95×105ZB,远大于目前的数据存储规模。In a specific embodiment, when the length of m is 8 bases, q=4, i=10, p=4, D=(256-10)×4×4 40 =1.18×10 27 . When 1 base stores 2 bits of information, the amount of information that can be stored is L=2.36×10 27 bits=2.95×10 26 bytes=2.95×10 5 ZB, which is much larger than the current data storage scale.
即使序列片段中含有8个碱基,也能够实现远大于目前存储量的数据存储。此外,由于每个序列单元及序列片段都采用检索序列进行了标记,因此该方法能够成功实现从序列片段到序列单元,以及从序列单元到碱基序列的数据恢复。Even if the sequence fragment contains 8 bases, it can realize data storage much larger than the current storage capacity. In addition, since each sequence unit and sequence fragment is marked with the search sequence, the method can successfully realize data recovery from sequence fragment to sequence unit, and from sequence unit to base sequence.
此外,本申请实施例通过将第一DNA分子分别存储,并根据实际需要进行扩增保存。在这种情况下,将第一DNA分子纳入DNA分子库中后,可以根据实际需要来提取第一DNA分子的备份信息,从而免去每次重新从头合成的麻烦,可以大幅降低存储成本。In addition, in the embodiment of the present application, the first DNA molecules are stored separately, and amplified and stored according to actual needs. In this case, after the first DNA molecule is included in the DNA molecule library, the backup information of the first DNA molecule can be extracted according to actual needs, thereby avoiding the trouble of re-synthesis each time and greatly reducing storage costs.
在一个实施例中,如图6所示,K个所述第一DNA分子的解码方法,包括:In one embodiment, as shown in Figure 6, the decoding method of the first K DNA molecules includes:
S201.对每个第一物理空间中存储的多个第一DNA分子进行测序,得到多个标记序列片段;根据索引信息对属于同一标记序列单元的每个标记序列片段对应的序列片段进行拼接,得到序列单元。S201. Sequencing the multiple first DNA molecules stored in each first physical space to obtain multiple marker sequence fragments; splicing the sequence fragments corresponding to each marker sequence fragment belonging to the same marker sequence unit according to the index information, Get the sequence unit.
对K个第一DNA分子进行测序的方式包括任意可以读取DNA产物的方式,比如二代测序,三代测序等等,获取待解码的多个信息序列单元。The method of sequencing the K first DNA molecules includes any method that can read DNA products, such as second-generation sequencing, third-generation sequencing, etc., to obtain multiple information sequence units to be decoded.
在一个实施例中,对K个第一DNA分子进行测序,分别得到K个标记序列片段。在一些实施例中,当第二DNA分子存储在与第一物理空间对应的第二物理空间,第二DNA分子存储有索引信息时,测序还包括:对第二DNA分子进行测序,获取包括第一检索序列和第二检索序列的索引信息。In one embodiment, K first DNA molecules are sequenced to obtain K marker sequence fragments respectively. In some embodiments, when the second DNA molecule is stored in the second physical space corresponding to the first physical space, and the second DNA molecule stores index information, the sequencing further includes: sequencing the second DNA molecule, obtaining the A retrieval sequence and index information of a second retrieval sequence.
上述步骤S201,可以通过图1所示存储装置中的DNA测序模块实现。The above step S201 can be implemented by the DNA sequencing module in the storage device shown in FIG. 1 .
S202.根据第二检索序列对属于同一标记序列单元的每个标记序列片段对应的序列片段进行拼接,得到S个序列单元;根据第一检索序列将得到的S个序列单元进行拼接,得到碱基序列。S202. Splicing the sequence fragments corresponding to each marker sequence fragment belonging to the same marker sequence unit according to the second retrieval sequence to obtain S sequence units; splicing the obtained S sequence units according to the first retrieval sequence to obtain bases sequence.
该步骤通过检索信息将分属于S个序列单元中的K个检索片段拼接成碱基序列。本申请实施例根据第一检索序列获取标记序列单元在碱基序列中的位置信息;根据第二检索序列获取属于同一序列单元的多个标记序列片段的位置信息或编号信息。根据得到的多个序列片段的位置信息或编号信息,将多个序列片段拼接成为序列单元;根据得到的序列单元的位置信息或编号信息,S个序列单元拼接为碱基序列。In this step, the K retrieval fragments belonging to the S sequence units are spliced into base sequences by retrieval information. In the embodiment of the present application, the position information of the marker sequence unit in the base sequence is obtained according to the first search sequence; the position information or numbering information of multiple marker sequence fragments belonging to the same sequence unit is obtained according to the second search sequence. According to the obtained position information or numbering information of the multiple sequence fragments, the multiple sequence fragments are spliced into a sequence unit; according to the obtained position information or numbering information of the sequence units, S sequence units are spliced into a base sequence.
在一些实施例中,根据第二检索序列对属于同一标记序列单元的每个标记序列片段对应的序列片段进行拼接,得到S个序列单元,包括:根据第二检索序列获取属于同一标记序列单元的每个标记序列片段对应的序列片段,以及序列片段在序列单元中的位置;根据序列片段在序列单元中的位置,将序列片段拼接成序列单元。In some embodiments, according to the second search sequence, the sequence fragments corresponding to each marker sequence fragment belonging to the same marker sequence unit are spliced to obtain S sequence units. The sequence fragment corresponding to each marker sequence fragment, and the position of the sequence fragment in the sequence unit; according to the position of the sequence fragment in the sequence unit, the sequence fragments are spliced into a sequence unit.
在一个实施例中,根据第一检索序列将得到的S个序列单元进行拼接,得到碱基序列,包括:In one embodiment, the obtained S sequence units are spliced according to the first search sequence to obtain a base sequence, including:
根据第一检索序列获取序列单元在碱基序列中的位置;Obtaining the position of the sequence unit in the base sequence according to the first search sequence;
根据序列单元在碱基序列中的位置,将s个序列单元拼接成碱基序列。According to the positions of the sequence units in the base sequence, the s sequence units are spliced into a base sequence.
示例性的,根据第一检索序列和第二检索序列,解读出完整的DNA序列“ATCT TCTGTTCC ATAA TAAG TGAG ATCC TACA TAGT ATCG TATG TGAC ATCC TGCT TGAC ATCT TCTGTTAA ATCT TGGA TAGA ATTG TAGC TTGC ATCG TATC TGTA ATCC TTCG TCCT ATCA TCTTTGCG ATTC TGGT TTGA ATCC TGAA TTTT ATTG TCAC TACG ATTG TCAC TACT ATGA TGGGTGGT”。Exemplarily, according to the first search sequence and the second search sequence, read out the complete DNA sequence "ATCT TCTGTTCC ATAA TAAG TGAG ATCC TACA TAGT ATCG TATG TGAC ATCC TGCT TGAC ATCT TCTGTTAA ATCT TGGA TAGA ATTG TAGC TTGC ATCG TATC TGTA ATCC TTCG TCCT ATCA TCTTTGCG ATTC TGGT TTGA ATCC TGAA TTTT ATTG TCAC TACG ATTG TCAC TACT ATGA TGGGTGGT".
S203.将碱基序列转换为目标数据。S203. Convert the base sequence into target data.
可以预先设定的与S101的数据写入时匹配的映射关系,将碱基序列转换为二进制序列。比如,按照A对应11,T对应10,C对应01,G对应00,可以将S102得到的碱基序列转换为二进制序列为:“11100110 10011000 10100101 11101111 10111100 10001100 1110010110110111 10110010 11100100 10111000 10001101 11100101 10000110 1000110111100110 10011000 10101111 11100110 10000011 10110011 11101000 1011000110100001 11100100 10111001 10001011 11100101 10100100 10010110 1110011110011010 10000100 11101001 10000010 10100011 11100101 10001111 1010101011101000 10011101 10110100 11101000 10011101 10110110 11100011 1000000010000010”。The base sequence can be converted into a binary sequence by a preset mapping relationship that matches the data written in S101.比如,按照A对应11,T对应10,C对应01,G对应00,可以将S102得到的碱基序列转换为二进制序列为:“11100110 10011000 10100101 11101111 10111100 10001100 1110010110110111 10110010 11100100 10111000 10001101 11100101 10000110 1000110111100110 10011000 10101111 11100110 10000011 10110011 11101000 1011000110100001 11100100 10111001 10001011 11100101 10100100 10010110 1110011110011010 10000100 11101001 10000010 10100011 11100101 10001111 1010101011101000 10011101 10110100 11101000 10011101 10110110 11100011 1000000010000010”。
然后根据二进制序列生成计算机信息序列。A sequence of computer information is then generated from the binary sequence.
根据所生成的二进制序列,结合预先所设定的编码规则,可以将二进制序列转换为对应的数据文件,包括如图片、文本、程序、音频、视频等文件。According to the generated binary sequence, combined with the preset encoding rules, the binary sequence can be converted into corresponding data files, including files such as pictures, text, programs, audio, and video.
如利用计算机程序将上述步骤S203得到的二进制序列恢复成“春,已不再是想象之外的那只蝴蝶。”的文本信息。For example, a computer program is used to recover the binary sequence obtained in the above step S203 into the text information of "Chun, no longer the unimaginable butterfly."
应理解,上述实施例中各步骤的实现,可以通过人为计算实现,也可以通过计算机程序实现。并且上述各步骤的序号的大小并不意味着执行顺序的先后,各过程的执行顺序应以其功能和内在逻辑确定,而不应对本申请实施例的实施过程构成任何限定。It should be understood that the implementation of each step in the foregoing embodiments may be implemented by human calculation or by computer programs. Moreover, the sequence numbers of the above steps do not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiment of the present application.
第二方面,参考图1,本申请实施例提供了一种DNA数据存储装置,包括数据处理模块,数据处理模块,用于获取与目标数据的二进制序列对应的碱基序列;对碱基序列进行分割,得到S个序列单元,且每个序列单元包括多个分割的序列片段,其中,S个序列单元共含有K个序列片段,且序列片段的长度为n,n、S和K均为大于或者等于2的整数;利用预设的索引信息对K个序列片段和S个序列单元进行标记,得到K个标记序列片段和S个标记序列单元,其中,索引信息包括用于表示S个所述序列单元在碱基序列中的排列顺序的第一检索序列,和用于表示属于同一序列单元中的多个序列片段在序列单元中的排列顺序的第二检索序列,K个标记序列片段用于合成存储有目标数据的K个第一DNA分子。In the second aspect, with reference to FIG. 1 , the embodiment of the present application provides a DNA data storage device, including a data processing module, and a data processing module for obtaining the base sequence corresponding to the binary sequence of the target data; Segmentation to obtain S sequence units, and each sequence unit includes multiple segmented sequence segments, wherein, S sequence units contain K sequence segments in total, and the length of the sequence segment is n, and n, S and K are all greater than Or an integer equal to 2; use the preset index information to mark K sequence fragments and S sequence units to obtain K mark sequence fragments and S mark sequence units, wherein the index information includes the information used to represent the S sequence units The first search sequence for the arrangement order of the sequence unit in the base sequence, and the second search sequence for indicating the arrangement order of multiple sequence fragments belonging to the same sequence unit in the sequence unit, and the K marker sequence fragments are used for K first DNA molecules storing target data are synthesized.
在一些实施例中,数据处理模块,还用于根据第二检索序列对属于同一标记序列单元的每个标记序列片段对应的序列片段进行拼接,得到序列单元;根据第一检索序列将得到的S个序列单元进行拼接,得到碱基序列;将碱基序列转换为目标数据。In some embodiments, the data processing module is further configured to splice the sequence segments corresponding to each tag sequence segment belonging to the same tag sequence unit according to the second search sequence to obtain the sequence unit; the obtained S according to the first search sequence The sequence units are spliced to obtain the base sequence; the base sequence is converted into the target data.
在一些实施例中,存储装置还包括DNA合成模块,用于将K个标记序列片段合成存储有目标数据的K个第一DNA分子。在一些实施例中,存储装置还包括DNA分子存储模块,用于将K个第一DNA分子存储在S个第一物理空间,其中,同属于一个序列单元的标记序列片段对应的第一DNA分子存储在同一个第一物理空间,不属于同一个序列单元的标记序列片段对应的第一DNA分子存储在不同的第一物理空间。In some embodiments, the storage device further includes a DNA synthesis module for synthesizing the K marker sequence fragments into K first DNA molecules storing target data. In some embodiments, the storage device further includes a DNA molecule storage module, configured to store K first DNA molecules in S first physical spaces, wherein the first DNA molecules corresponding to the marker sequence fragments belonging to one sequence unit Stored in the same first physical space, the first DNA molecules corresponding to marker sequence fragments that do not belong to the same sequence unit are stored in different first physical spaces.
在一些实施例中,DNA分子存储模块还用于将第二DNA分子存储在第二物理空间。In some embodiments, the DNA molecule storage module is also used to store a second DNA molecule in a second physical space.
在一些实施例中,存储装置还包括DNA分子测序模块,对每个第一物理空间中存储的多个第一DNA分子进行测序,得到多个标记序列片段;In some embodiments, the storage device further includes a DNA molecule sequencing module, which performs sequencing on a plurality of first DNA molecules stored in each first physical space to obtain a plurality of marker sequence fragments;
图1含有上述模块的装置能利用DNA进行数据存储写入,与图2所示的DNA数据存储写入的方法对应。The device containing the above modules in FIG. 1 can use DNA for data storage and writing, which corresponds to the method for DNA data storage and writing shown in FIG. 2 .
本申请实施例还提供了一种DNA数据存储设备。如图7所示,本实施例提供的终端设备70包括:处理器710、存储器720以及存储在存储器720中并可在处理器710上运行的计算机程序721。处理器710执行计算机程序721时实现上述DNA数据的存储方法各个实施例中的步骤,例如图2所示的步骤S101至S103。The embodiment of the present application also provides a DNA data storage device. As shown in FIG. 7 , the terminal device 70 provided in this embodiment includes: a
示例性的,计算机程序721可以被分割成一个或多个模块/单元,一个或者多个模块/单元被存储在存储器720中,并由处理器710执行,以完成本申请。一个或多个模块/单元可以是能够完成特定功能的一系列计算机程序指令段,该指令段可以用于描述计算机程序721在终端设备中的执行过程。例如,计算机程序721可以被分割成数据处理模块。在一些实施例中,计算机程序721还可以被分割成数据处理模块、DNA分子合成模块、DNA分子存储模块和DNA分子存储模块,各模块具体功能如文所述。为了节约篇幅,此处不再赘述。Exemplarily, the
本领域技术人员可以理解,图7仅仅是终端设备70的一种示例,并不构成对终端设备70的限定,可以包括比图示更多或更少的部件,或者组合某些部件,或者不同的部件。Those skilled in the art can understand that FIG. 7 is only an example of a terminal device 70, and does not constitute a limitation to the terminal device 70. It may include more or less components than those shown in the figure, or combine certain components, or be different. parts.
处理器710可以是中央处理单元(Central Processing Unit,CPU),还可以是其他通用处理器、数字信号处理器(Digital Signal Processor,DSP)、专用集成电路(Application Specific Integrated Circuit,ASIC)、现成可编程门阵列(Field-Programmable Gate Array,FPGA)或者其他可编程逻辑器件、分立门或者晶体管逻辑器件、分立硬件组件等。通用处理器可以是微处理器或者该处理器也可以是任何常规的处理器等。The
存储器720可以是终端设备70的内部存储单元,例如终端设备70的硬盘或内存。存储器720也可以是终端设备70的外部存储设备,例如终端设备70上配备的插接式硬盘,智能存储卡(Smart Media Card,SMC),安全数字(Secure Digital,SD)卡,闪存卡(Flash Card)等等。进一步地,存储器720还可以既包括终端设备70的内部存储单元也包括外部存储设备。存储器720用于存储计算机程序721以及终端设备70所需的其他程序和数据。存储器720还可以用于暂时地存储已经输出或者将要输出的数据。The
本申请实施例还提供了一种计算机可读存储介质,计算机可读存储介质存储有计算机程序,计算机程序被处理器执行时实现前述各实施例的处理方法。An embodiment of the present application also provides a computer-readable storage medium, where a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the processing methods of the foregoing embodiments are implemented.
本申请实施例还提供了一种计算机程序产品,当计算机程序产品在终端设备上运行时,使得终端设备执行前述各实施例的DNA数据的存储方法。The embodiment of the present application also provides a computer program product, which, when the computer program product is run on the terminal device, enables the terminal device to execute the methods for storing DNA data in the foregoing embodiments.
本申请实施例提供一种DNA硬盘,参考图5,包括多个物理空间,物理空间由物理材料制成,每个物理空间用于存储DNA分子。An embodiment of the present application provides a DNA hard disk. Referring to FIG. 5 , it includes multiple physical spaces made of physical materials, and each physical space is used to store DNA molecules.
本申请实施例中,物理材料制成的物理空间包裹DNA分子,DNA分子通过物理材料隔离。其中,物理空间的形状没有严格限定,可以是圆形的、方形的,还可以是其他任意形状。In the embodiment of the present application, the physical space made of physical materials wraps DNA molecules, and the DNA molecules are isolated by physical materials. Wherein, the shape of the physical space is not strictly limited, and may be circular, square, or any other shape.
在一些实施例中,物理空间中存储的DNA分子包括上文中的第一DNA分子,此时,物理空间为第一物理空间。对应的,第一DNA分子包括第二检索序列和序列片段。每一个第一物理空间还存储有上文所述的第二检索序列。In some embodiments, the DNA molecules stored in the physical space include the above-mentioned first DNA molecule, and in this case, the physical space is the first physical space. Correspondingly, the first DNA molecule includes the second search sequence and sequence fragments. Each first physical space also stores the above-mentioned second retrieval sequence.
在一些实施例中,物理空间中存储的DNA分子包括上文所述的第二DNA分子,此时,物理空间为第二物理空间。In some embodiments, the DNA molecules stored in the physical space include the above-mentioned second DNA molecules, and in this case, the physical space is the second physical space.
在一些实施例中,物理材料选自SiO2、金属氧化物、高分子聚合材料中的至少一种。其中,高分子聚合材料包括树脂,但不限于树脂。In some embodiments, the physical material is at least one selected from SiO 2 , metal oxides, and polymer materials. Wherein, the high molecular polymer material includes resin, but is not limited to resin.
以上仅为本申请的可选实施例而已,并不用于限制本申请。对于本领域的技术人员来说,本申请可以有各种更改和变化。凡在本申请的精神和原则之内,所作的任何修改、等同替换、改进等,均应包含在本申请的权利要求范围之内。The above are only optional embodiments of the application, and are not intended to limit the application. For those skilled in the art, various modifications and changes may occur in this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included within the scope of the claims of the present application.
Claims (14)
- A method for storing DNA data, comprising:acquiring a base sequence corresponding to a binary sequence of target data;dividing the base sequence to obtain S sequence units, wherein each sequence unit comprises a plurality of divided sequence fragments, the S sequence units contain K sequence fragments, the length of each sequence fragment is n, and n, S and K are integers greater than or equal to 2;labeling the K sequence segments and the S sequence units by using preset index information to obtain K labeled sequence segments and S labeled sequence units, wherein the index information comprises a first retrieval sequence used for representing the arrangement sequence of the S sequence units in the base sequence and a second retrieval sequence used for representing the arrangement sequence of a plurality of sequence segments belonging to the same sequence unit in the sequence unit, and the K labeled sequence segments are used for synthesizing K first DNA molecules storing the target data;the storage method further comprises the following steps:after synthesizing K marked sequence segments into K first DNA molecules storing the target data, storing the K first DNA molecules in S first physical spaces, wherein the first DNA molecules corresponding to the marked sequence segments belonging to the same sequence unit are stored in the same first physical space, and the first DNA molecules corresponding to the marked sequence segments not belonging to the same sequence unit are stored in different first physical spaces.
- 2. The method for storing DNA data according to claim 1, wherein the labeling of the plurality of sequence fragments belonging to the same sequence unit with the second search sequence comprises:splicing a second search sequence on either side of the sequence fragment, orAnd simultaneously splicing retrieval base groups on two sides of the sequence fragment, wherein the retrieval base groups on the two sides form the second retrieval sequence.
- 3. The method for storing DNA data according to claim 1, wherein the first search sequence includes i DNA sequence fragments, i is an integer of 1 or more, and each of the DNA sequence fragments includes a first base sequence serving as an index marker and a second base sequence for identifying the sequence unit number.
- 4. The method according to claim 1, wherein s first physical spaces are integrated in one DNA hard disk.
- 5. The method for storing DNA data according to claim 1, further comprising:storing a second DNA molecule in a second physical space corresponding to the first physical space, the second DNA molecule storing the index information.
- 6. The method for storing DNA data according to any one of claims 1 to 5, wherein the method for decoding K first DNA molecules comprises:sequencing a plurality of the first DNA molecules stored in each of the first physical spaces to obtain a plurality of the marker sequence fragments;splicing the sequence fragments corresponding to each of the labeled sequence fragments belonging to the same labeled sequence unit according to the second retrieval sequence to obtain the sequence unit; splicing the obtained S sequence units according to the first retrieval sequence to obtain the base sequence;converting the base sequence into the target data.
- 7. A DNA data storage device is characterized by comprising a data processing module,the data processing module is used for acquiring a base sequence corresponding to the binary sequence of the target data; dividing the base sequence to obtain S sequence units, wherein each sequence unit comprises a plurality of divided sequence fragments, the S sequence units contain K sequence fragments in total, the length of each sequence fragment is n, and n, S and K are integers greater than or equal to 2; labeling the K sequence segments and the S sequence units by using preset index information to obtain K labeled sequence segments and S labeled sequence units, wherein the index information comprises a first retrieval sequence used for representing the arrangement sequence of the S sequence units in the base sequence and a second retrieval sequence used for representing the arrangement sequence of a plurality of sequence segments belonging to the same sequence unit in the sequence unit, and the K labeled sequence segments are used for synthesizing K first DNA molecules storing the target data;the device further comprises: a DNA molecule storage module,for storing K of said first DNA molecules in S first physical spaces, wherein said first DNA molecules corresponding to said marker sequence fragments belonging to one of said sequence units are stored in the same one of said first physical spaces and said first DNA molecules corresponding to said marker sequence fragments not belonging to the same one of said sequence units are stored in a different one of said first physical spaces.
- 8. The DNA data storage device of claim 7, further comprising: and the DNA molecule synthesis module is used for synthesizing the K marking sequence segments into K first DNA molecules in which the target data are stored.
- 9. The DNA data storage device of claim 7, wherein the DNA molecule storage module is further configured to store a second DNA molecule in a second physical space.
- 10. The DNA data storage device of claim 7, further comprising a DNA molecule sequencing module for sequencing a plurality of the first DNA molecules stored in each of the first physical spaces to obtain a plurality of the marker sequence fragments;the data processing module is further configured to splice the sequence segments corresponding to each of the labeled sequence segments belonging to the same labeled sequence unit according to the second search sequence to obtain the sequence unit; splicing the obtained S sequence units according to the first retrieval sequence to obtain the base sequence; converting the base sequence into the target data.
- 11. A DNA data storage device comprising a terminal device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the storage method of DNA data according to any one of claims 1 to 6 when executing the computer program.
- 12. A computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the method for storing DNA data according to any one of claims 1 to 6.
- 13. A DNA hard disk, characterized by comprising a plurality of physical spaces made of physical materials, wherein each physical space is used for storing DNA molecules, and the DNA molecules are stored according to the storage method of any one of claims 1 to 6.
- 14. The DNA hard disk of claim 13 characterized in that the physical material is selected from SiO 2 At least one of metal oxide and polymer material.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202110929436.9A CN113782102B (en) | 2021-08-13 | 2021-08-13 | DNA data storage method, device, equipment and readable storage medium |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202110929436.9A CN113782102B (en) | 2021-08-13 | 2021-08-13 | DNA data storage method, device, equipment and readable storage medium |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| CN113782102A CN113782102A (en) | 2021-12-10 |
| CN113782102B true CN113782102B (en) | 2022-12-13 |
Family
ID=78837721
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| CN202110929436.9A Expired - Fee Related CN113782102B (en) | 2021-08-13 | 2021-08-13 | DNA data storage method, device, equipment and readable storage medium |
Country Status (1)
| Country | Link |
|---|---|
| CN (1) | CN113782102B (en) |
Families Citing this family (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN114958828B (en) * | 2022-06-14 | 2024-04-19 | 深圳先进技术研究院 | Data information storage method based on DNA molecular medium |
| CN114758703B (en) * | 2022-06-14 | 2022-09-13 | 深圳先进技术研究院 | Data information storage method based on recombinant plasmid DNA molecules |
Citations (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2005247900A (en) * | 2004-03-01 | 2005-09-15 | Ltt Bio-Pharma Co Ltd | Method for judging seal or sign put by using dna-containing ink |
| CN101702240A (en) * | 2009-11-26 | 2010-05-05 | 大连大学 | Image encryption method based on DNA sub-sequence operation |
| CN106845158A (en) * | 2017-02-17 | 2017-06-13 | 苏州泓迅生物科技股份有限公司 | A kind of method that information Store is carried out using DNA |
| CN111095423A (en) * | 2017-08-25 | 2020-05-01 | 深圳华大生命科学研究院 | Encoding/decoding method, device and data processing device |
| CN111091876A (en) * | 2019-12-16 | 2020-05-01 | 中国科学院深圳先进技术研究院 | A DNA storage method, system and electronic device |
| CN111858510A (en) * | 2020-07-16 | 2020-10-30 | 中国科学院北京基因组研究所(国家生物信息中心) | DNA movable type storage system and method |
| CN112288090A (en) * | 2020-10-22 | 2021-01-29 | 中国科学院深圳先进技术研究院 | Method and device for processing DNA sequence with data information |
| CN112673428A (en) * | 2019-05-31 | 2021-04-16 | 伊鲁米那股份有限公司 | Storage device, system and method |
Family Cites Families (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US11093547B2 (en) * | 2018-06-19 | 2021-08-17 | Intel Corporation | Data storage based on encoded DNA sequences |
| CN112749247B (en) * | 2019-10-31 | 2023-08-18 | 中国科学院深圳先进技术研究院 | Method and device for storing and reading text information |
-
2021
- 2021-08-13 CN CN202110929436.9A patent/CN113782102B/en not_active Expired - Fee Related
Patent Citations (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2005247900A (en) * | 2004-03-01 | 2005-09-15 | Ltt Bio-Pharma Co Ltd | Method for judging seal or sign put by using dna-containing ink |
| CN101702240A (en) * | 2009-11-26 | 2010-05-05 | 大连大学 | Image encryption method based on DNA sub-sequence operation |
| CN106845158A (en) * | 2017-02-17 | 2017-06-13 | 苏州泓迅生物科技股份有限公司 | A kind of method that information Store is carried out using DNA |
| CN111095423A (en) * | 2017-08-25 | 2020-05-01 | 深圳华大生命科学研究院 | Encoding/decoding method, device and data processing device |
| CN112673428A (en) * | 2019-05-31 | 2021-04-16 | 伊鲁米那股份有限公司 | Storage device, system and method |
| CN111091876A (en) * | 2019-12-16 | 2020-05-01 | 中国科学院深圳先进技术研究院 | A DNA storage method, system and electronic device |
| CN111858510A (en) * | 2020-07-16 | 2020-10-30 | 中国科学院北京基因组研究所(国家生物信息中心) | DNA movable type storage system and method |
| CN112288090A (en) * | 2020-10-22 | 2021-01-29 | 中国科学院深圳先进技术研究院 | Method and device for processing DNA sequence with data information |
Non-Patent Citations (4)
| Title |
|---|
| Addressing Information Using Data Hiding for DNA-based Storage Systems;Takahiro Ota等;《2020 International Symposium on Information Theory and Its Applications (ISITA)》;20210802;第509-513页 * |
| DNA数据存储技术原理及其研究进展;滕越等;《生物化学与生物物理进展》;20210531;第48卷(第5期);第494-504页 * |
| Mendel: A Distributed Storage Framework for Similarity Searching over Sequencing Data;Cameron Tolooee等;《2016 IEEE International Parallel and Distributed Processing Symposium (IPDPS)》;20160721;第790-799页 * |
| 人工DNA合成技术:DNA数据存储的基石;黄小罗等;《合成生物学》;20210228;第2卷(第3期);第335-353页 * |
Also Published As
| Publication number | Publication date |
|---|---|
| CN113782102A (en) | 2021-12-10 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN109830263B (en) | DNA storage method based on oligonucleotide sequence coding storage | |
| CN112382340B (en) | Encoding and decoding method and encoding and decoding device for DNA data storage | |
| CN112288090B (en) | Method and device for processing DNA sequences containing data information | |
| Wang et al. | High capacity DNA data storage with variable-length Oligonucleotides using repeat accumulate code and hybrid mapping | |
| US20180373839A1 (en) | Systems and methods for encoding genomic graph information | |
| US20170249345A1 (en) | A biomolecule based data storage system | |
| CN113744804B (en) | Method and device for storing data by using DNA and storage equipment | |
| CN110088839B (en) | Efficient data structures for bioinformatics information representation | |
| CN105022935A (en) | Encoding method and decoding method for performing information storage by means of DNA | |
| CN110121577A (en) | Methods and systems for representing and processing bioinformatic data using reference sequences | |
| Gervasio et al. | How close are we to storing data in DNA? | |
| CN112527736A (en) | Data storage method and data recovery method based on DNA and terminal equipment | |
| CN110168652B (en) | Methods and systems for storing and accessing bioinformatics data | |
| CN111095423B (en) | Encoding/decoding method, device and data processing device | |
| CN113782102A (en) | DNA data storage method, device, device and readable storage medium | |
| WO2022109879A1 (en) | Encoding and decoding method and encoding and decoding device between binary information and base sequence for dna data storage | |
| Wang et al. | Mainstream encoding–decoding methods of DNA data storage | |
| CN109658981A (en) | A kind of data classification method of unicellular sequencing | |
| Wu et al. | HD-code: End-to-end high density code for DNA storage | |
| CN100566177C (en) | Conversion method and system | |
| CN114822695B (en) | Encoding method and encoding device for DNA storage | |
| Beck et al. | Finding data in DNA: computer forensic investigations of living organisms | |
| WO2023015550A1 (en) | Dna data storage method and apparatus, device, and readable storage medium | |
| Goel | A compression algorithm for DNA that uses ASCII values | |
| WO2022082573A1 (en) | Method and apparatus for processing dna sequence storing data information |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| PB01 | Publication | ||
| PB01 | Publication | ||
| SE01 | Entry into force of request for substantive examination | ||
| SE01 | Entry into force of request for substantive examination | ||
| TA01 | Transfer of patent application right |
Effective date of registration: 20220517 Address after: 518000 4th floor, Zhuohong building, Zhenmei community, Xinhu street, Guangming District, Shenzhen, Guangdong Applicant after: Zhongke carbon yuan (Shenzhen) Biotechnology Co.,Ltd. Address before: 1068 No. 518055 Guangdong city in Shenzhen Province, Nanshan District City Xili University School Avenue Applicant before: SHENZHEN INSTITUTES OF ADVANCED TECHNOLOGY |
|
| TA01 | Transfer of patent application right | ||
| GR01 | Patent grant | ||
| GR01 | Patent grant | ||
| CF01 | Termination of patent right due to non-payment of annual fee |
Granted publication date: 20221213 |
|
| CF01 | Termination of patent right due to non-payment of annual fee |