WO2015199440A1 - 핵산염기서열 보안 방법, 장치 및 이를 저장한 기록매체 - Google Patents
핵산염기서열 보안 방법, 장치 및 이를 저장한 기록매체 Download PDFInfo
- Publication number
- WO2015199440A1 WO2015199440A1 PCT/KR2015/006427 KR2015006427W WO2015199440A1 WO 2015199440 A1 WO2015199440 A1 WO 2015199440A1 KR 2015006427 W KR2015006427 W KR 2015006427W WO 2015199440 A1 WO2015199440 A1 WO 2015199440A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- target
- analysis
- nucleic acid
- complexes
- complex
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F21/00—Security arrangements for protecting computers, components thereof, programs or data against unauthorised activity
- G06F21/60—Protecting data
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B99/00—Subject matter not provided for in other groups of this subclass
Definitions
- the present invention relates to a nucleic acid nucleotide sequence security technology, and more particularly, to a nucleic acid nucleotide sequence security method, apparatus and apparatus for analyzing the nucleic acid nucleotide sequence of an analysis requester without exposing the nucleic acid nucleotide sequence of the requester to the analyst. It relates to a recording medium.
- Genome or genome means the total sequence of a chromosome in an individual.
- the genome is the sum of nearly complete genetic information of a species and contains nucleic acid sequence information.
- the human genome is about 32 billion pairs of all nucleic acid sequences, including all the genes and other non-genes needed to make a human individual, with 22 pairs, 44 autosomes, and one pair of sex chromosomes (X, Y). And mitochondria.
- the nucleic acid of the genome is a double helix consisting of adenine (A), thiamin (T), guanine (G) and cytosine (C) bases together with phosphoric acid and sugars. Genetic information is encoded by the sequence of the four bases of the nucleic acid.
- Human genes each with hundreds or thousands of amino acid sequences, provide schematics for all the proteins produced by the body. Human genes are estimated to be about 30,000 to 50,000, and make up about 3 to 5% of the genome. About 30,000 to 50,000 genes of humans are encoded in nucleotide sequences, and only a part of them is revealed. The genotype of nucleic acid base sequences present in 95-97% of the genome nucleic acid sequences that do not belong to the genome of the genome has a number of sites related to the phenotype of the individual and its medical meaning has been continuously revealed. The interpretation of the medical and biological meanings of genomic information encoded by sequencing continues to develop. Therefore, an individual who primarily analyzes the nucleotide sequence may need to repeatedly search and interpret the nucleotide sequence information in order to confirm his or her correspondence to the newly discovered genome sequencing information.
- Genome sequence information is classified as sensitive information that is burdensome to expose and should not be exposed to others as much as possible.
- the calculation methods and information resources necessary for the interpretation of the genotype information owned by an individual are held by an analyst or an analysis institution having analysis equipment and an analysis method for information decoding, so that an individual can In order to calculate the phenotypic occurrence probability, the genotype should be sent to the analyst or analyst.
- Such genotype information transmission has a problem of increasing the risk of exposure of genotype information to an individual.
- Korean Patent Publication No. 2013-0075559 relates to a method for managing genetic information, and more specifically, by dividing and storing individual genetic information into a plurality of subsequences and managing them, such a partitioning is performed when hacking or leakage of stored information occurs.
- the above patent relates to security in the storage and storage of sensitive genetic information.
- Information related to the process by which a person or an organization possessing the genetic information to be analyzed transmits nucleic acid sequence information to an analyst or an analysis institution and receives the analysis result. It is distinguished from the present invention regarding security management technology.
- the disclosed patent has a problem that when the delimiter information is exposed, genetic information of an individual may be exposed, and the security of the delimiter information itself may also be weak.
- One embodiment of the present invention is to provide a nucleic acid base sequence security method that can analyze the nucleic acid base sequence of the analysis requester without exposing the nucleic acid base sequence of the analysis requester to the analyst.
- An embodiment of the present invention expresses a complex comprising a target unit and a gastrointestinal unit derived from a nucleic acid base sequence of an analysis requester, and analyzes the nucleic acid base sequence of an analysis requester through a sparse matrix defined in view of a plurality of target units.
- the present invention provides a method for securing nucleic acid base sequence.
- An embodiment of the present invention to provide a nucleic acid base sequence security method that can set the security strength of the nucleic acid base sequence by determining the number of the plurality of complexes each of the target unit and the gastrointestinal unit or the size of each of the plurality of complexes do.
- An embodiment of the present invention is to provide a nucleic acid sequence security method that can receive the analysis results for a plurality of complexes from the analyst to obtain the analysis results for the nucleic acid base sequence without exposure of the nucleic acid base sequence of the requester. .
- nucleic acid bases capable of keeping inherent information or know-how, such as odds ratios for individual bases possessed by the analyst It is intended to provide a method of sequence security.
- An embodiment of the present invention is to provide a nucleic acid base sequence security method that can improve the level of security while minimizing the amount of calculation of the nucleic acid base sequence security device.
- the method for securing nucleic acid base sequences may comprise (a) at least one of a target element derived from the nucleic acid sequence of the assay requester or a disguising element that is the same as or different from the target unit. Generating a plurality of complexes comprising; and (b) providing the analyst with the generated plurality of complexes.
- the step (a) may include expressing the plurality of complexes as sparse matrices for the plurality of target units.
- the step (a) may include the target unit when the target unit is included in the complex.
- the method may further include determining a location.
- the step (a) may further comprise defining a target monomer cell in the sparse matrix as a base-locus set comprising at least one base and a locus for the target monomer.
- the step (a) may further comprise defining a target unit cell in the sparse matrix as at least one base set associated with the locus.
- the step (a) may further include dynamically determining the position of the target monomer cell in the sparse matrix to generate the target monomer map required in the decoding process.
- the step (a) may comprise the step of generating the target unit by extracting at least one base and the locus from the nucleic acid base sequence.
- the step (a) may include generating the target monomer by segmenting the nucleic acid base sequence into partial nucleotide sequences.
- the step (a) may include generating at least one gastrointestinal unit in the complex based on the similarity with the target unit.
- the step (a) may further include generating at least one gastrointestinal unit whose genetic distance or evolutionary distance from the target unit is less than or equal to a specific distance.
- the step (a) may include determining the number of the plurality of complexes or the size of each of the plurality of complexes according to the security strength set by the analysis requester.
- Step (b) may include dividing the generated plurality of complexes and providing them to a plurality of direct or indirect analysts.
- the method may further include (c) receiving a plurality of analysis composites representing analysis results of the plurality of complexes from the analyzer to obtain an analysis result of the nucleic acid base sequence.
- Step (c) may include determining a plurality of target analysis elements representing analysis results of the plurality of target monomers based on the target monomer map.
- the step (c) may further include calculating the posterior odds of the analysis requester by combining the determined plurality of target analysis units.
- the nucleic acid sequence security device may comprise a plurality of target elements each of which comprises at least one of a target element derived from the nucleic acid sequence of the assay requester, or a disguising element that is the same as or different from the target unit. And a complex providing unit for generating complexes of the complex and a complex providing unit for providing an analyst with the generated plurality of complexes.
- the apparatus may further include an analysis complex analyzer configured to receive a plurality of analysis complexes representing analysis results of the plurality of complexes from the analyzer to obtain an analysis result of the nucleic acid base sequence.
- an analysis complex analyzer configured to receive a plurality of analysis complexes representing analysis results of the plurality of complexes from the analyzer to obtain an analysis result of the nucleic acid base sequence.
- the analysis complex analyzer may determine a plurality of target analysis elements representing analysis results of a plurality of target monomers based on a target monomer map.
- the analysis complex analyzer may calculate the posterior odds of the analysis requestor by combining the determined plurality of target analysis units.
- the recording medium recording the computer program relating to the nucleic acid sequence security method is a target element derived from the nucleic acid sequence of the requester of the analysis or a disguising element identical or different to the target unit. And a function of generating a plurality of composites including at least one of the above and providing the analyst with the generated plurality of complexes.
- the recording medium may further include a function of receiving a plurality of analysis complexes representing an analysis result of the plurality of complexes from the analyzer to obtain an analysis result of the nucleic acid base sequence.
- the disclosed technique can have the following effects. However, since a specific embodiment does not mean to include all of the following effects or only the following effects, it should not be understood that the scope of the disclosed technology is limited by this.
- the nucleic acid base sequence security method may analyze the nucleic acid base sequence of the analysis requester without exposing the nucleic acid base sequence of the analysis requester to the analyst.
- Nucleic acid base security method represents a complex comprising a target monomer and a gastrointestinal monomer derived from the nucleic acid sequence of the requester, respectively, and analyzed through a sparse matrix defined in terms of a plurality of target monomers The nucleic acid sequence of the requester can be analyzed.
- the security strength of the nucleic acid base sequence may be set by determining the number of the plurality of complexes or the size of each of the plurality of complexes each including the target unit and the gastrointestinal unit.
- the nucleic acid base sequence security method may receive an analysis result of a plurality of complexes to obtain an analysis result of the nucleic acid base sequence without exposing the nucleic acid sequence of the requester.
- Nucleic acid base security method by transmitting the analysis results in units of monomers for the nucleic acid base sequence from the analyst's point of view such as unique information such as the likelihood ratio (odds ratio) for each base possessed by the analyst You can keep your know-how.
- Nucleic acid base sequence security method can improve the level of security while minimizing the amount of calculation of the nucleic acid base sequence security device.
- FIG. 1 is a view illustrating a nucleic acid base sequence security system according to an embodiment of the present invention.
- FIG. 2 is a block diagram illustrating the nucleic acid base sequence security device of FIG. 1.
- FIG. 3 is a diagram illustrating a procedure for generating a plurality of complexes and analyzing a plurality of analysis complexes according to an embodiment of the present invention.
- FIG. 4 is a diagram illustrating a procedure of generating a plurality of complexes and analyzing a plurality of analyte complexes according to another embodiment of the present invention.
- FIG. 5A is a diagram visualizing a sparse matrix for a target unit according to an embodiment of the present invention
- FIG. 5B is a diagram visualizing a sparse matrix for a target unit according to another embodiment of the present invention.
- FIG. 6 is a flowchart illustrating a nucleic acid nucleotide sequence security method performed by the nucleic acid nucleotide sequence security device of FIG. 2.
- first and second are intended to distinguish one component from another component, and the scope of rights should not be limited by these terms.
- first component may be named a second component, and similarly, the second component may also be named a first component.
- an identification code (e.g., a, b, c, etc.) is used for convenience of description, and the identification code does not describe the order of the steps, and each step clearly indicates a specific order in context. Unless stated otherwise, they may occur out of the order noted. That is, each step may occur in the same order as specified, may be performed substantially simultaneously, or may be performed in the reverse order.
- the present invention can be embodied as computer readable code on a computer readable recording medium
- the computer readable recording medium includes all kinds of recording devices in which data can be read by a computer system.
- Examples of computer-readable recording media include ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, and the like, and are also implemented in the form of a carrier wave (for example, transmission over the Internet). It also includes.
- the computer readable recording medium can also be distributed over network coupled computer systems so that the computer readable code is stored and executed in a distributed fashion.
- Bases are also called nucleobases, or bases for short, and are cytosine, nitrogen bases found in deoxyribonucleic acid (DNA), ribonucleic acid (RNA), nucleotides, and nucleosides. Guanine (Guanine), adenine (Adenine), thymine (Thymine), may include uracil (Uracil). Nucleotides (nucleotides) are organic molecules that make up nucleic acids such as DNA and RNA, and are composed of base-sugar-phosphate bonds. Nucleic acids are a type of polymeric organic material in which nucleotides are polymerized into long chains.
- Nucleic acid base sequences can be implemented through arrays, the address of the array (ie, index) is the genomic coordinate value (hereinafter, the locus) (eg rsID number), and the data value stored in the array is the base (eg A, G). , T, C).
- the base here may include a genotype.
- the locus can be a continuous value or a list of discrete coordinate values extracted by selecting only the parts of the genome that are needed.
- Nucleotide base sequence can specify only one of the address of the array and omit the rest because the remaining address in the same array can be clearly identified if only one of the address of the array is specified if the locus is a continuous coordinate value.
- the personal genomic nucleic acid sequence includes the entire nucleic acid sequence of the genome of an individual and may correspond to the nucleic acid sequence of the requester of the analysis.
- the target nucleic acid base sequence corresponds to a plurality of target units associated with the nucleic acid base sequence of the requester, and is generated by extracting and arranging the locus-base pairs of the site of analysis from the individual genomic nucleic acid sequence and arranging them in a predetermined order, or by creating a personal genome.
- the nucleic acid base sequence can be generated by fragmenting. That is, the target nucleic acid base sequence may be generated by linking a plurality of target units in a certain order.
- the target nucleic acid base sequence may be the whole of the original personal genomic nucleic acid base sequence, a partial sequence extracted from the individual genomic nucleic acid base sequence, or a base sequence combining the partial base sequences extracted from various sites of the individual genomic nucleic acid base sequence. .
- the target nucleic acid base sequence may be individual genotype information.
- SNP single nucleotide polymorphism
- Polymorphism refers to individual differences in the nucleotide sequences present on the genome
- monobasic polymorphism refers to individual differences in the base of one of the base sequences consisting of A, T, C and G.
- nucleotide sequences include various genome variations, which are caused by substitution, addition, or deletion of bases, including single nucleotide polymorphism (SNV), short tandem repeat polymorphism (STRP), or variable number of tandem (VNTR). in the form of polyalleic variations, including repeat) and CNV (Copy numbervariation).
- SNV single nucleotide polymorphism
- SRP short tandem repeat polymorphism
- VNTR variable number of tandem
- CNV Copy numbervariation
- the target element corresponds to a partial base sequence derived from the nucleic acid base sequence of the requester.
- the target monomer may be generated by extracting at least one base and a locus from the nucleic acid sequence of the requester, or may be generated by segmenting the nucleic acid sequence of the requester into a partial base sequence.
- the disguising element may comprise the same or different base sequence as the target unit to make identification of the target unit difficult.
- the gastrointestinal unit may be generated by using a nucleic acid base sequence existing in nature or by referring to a nucleic acid base sequence existing in nature.
- the gastrointestinal unit may consist of partial sequences that do not match the target unit, may consist of partial sequences that partially match and partially mismatch the target unit, or may consist of the same partial sequences as the target unit have.
- the gastrointestinal unit is generated by referring to a nucleic acid sequence of the requester of the analysis, or by splitting at least one randomly generated nucleic acid sequence into at least two partial sequences, or a nucleic acid base of the assay requester. Referencing the nucleotide sequence of the target unit generated by segmentation or may be composed of at least one nucleic acid base sequence generated randomly.
- the complex comprises at least one of a target unit or a gastrointestinal unit that is the same as or different from the target unit. That is, the complex may include the target unit and the gastrointestinal unit together, may include only the gastrointestinal unit, or may include the target unit.
- Analytical complexes may represent analytical results for the complex.
- the assay complex may include a set of interpretation results for nucleic acid base sequences corresponding to the unit (s) included in the complex.
- the assay complex may comprise a set of nucleic acid base sequences corresponding to the unit (s) included in the complex and a sequence-interpretation pair that includes the results of the analysis together.
- the analysis complex may also include an analysis result of an embodiment in which the nucleic acid base sequence information of the unit (s) included in the complex is not indicated or omitted, and replaced with an identifier assigned for identification of the nucleic acid sequence.
- the present invention may enable a high level of security by changing the size (number of units included in the complex) or number of such complexes and assay complexes.
- FIG. 1 is a view illustrating a nucleic acid base sequence security system according to an embodiment of the present invention.
- the nucleic acid nucleotide sequence security system 100 includes a nucleic acid nucleotide sequence security device 110 and an analysis server (hereinafter, an analyst) 120, which may be connected through a network.
- an analysis server hereinafter, an analyst
- the nucleic acid nucleotide sequence security device 110 may request an analysis on a composite generated based on the nucleotide sequence of the analysis requester, and may be implemented as, for example, a desktop, a notebook, a tablet PC, or a smartphone. .
- the nucleic acid base sequence of the analysis requester may be managed through a plurality of memory regions (a sparse matrix memory region and a target monomer map region to be described later).
- the analyzer 120 is connected to the nucleic acid nucleotide sequence security device 110 through a network and receives a composite from the nucleic acid nucleotide sequence security device 110, and analyzes the received complex to produce an analysis composite.
- the base sequence security device 110 may be provided.
- FIG. 2 is a block diagram illustrating the nucleic acid base sequence security device of FIG. 1.
- the nucleic acid base sequence security device 110 includes a processor 210, a memory 220, a network interface 230, a user input device 240, a user output device 250, and a storage device 260. It includes.
- the processor 210 includes a complex generator 212, a complex provider 214, an assay complex receiver 216, an assay complex analyzer 218, and a nucleic acid nucleotide sequence security controller 219.
- the complex generating unit 212 generates a plurality of complexes each including at least one of a target unit derived from the nucleic acid sequence of the requester or a gastrointestinal unit which is the same as or different from that of the target unit.
- the position of the target monomer in the complex may be determined.
- the complex generating unit 212 generates a first complex including one target unit, a second complex including four gastrointestinal units, and a third complex including one target unit and one gastrointestinal unit And positioning of the target units included in the first complex and the third complex in the corresponding complex.
- the complex generator 212 may express the plurality of complexes as a sparse matrix for the plurality of target units.
- the complex generator 212 dynamically determines the position of the target monomer cell in the sparse matrix to generate the target monomer map 224 necessary in the decoding process (the process of obtaining the analysis result of the target monomer from the plurality of analysis complexes). can do.
- the complex generator 212 may determine the number of the plurality of complexes or the size (the number of target units or gastrointestinal units included in the complex) of the plurality of complexes or the plurality of complexes according to the set security strength.
- the number of the plurality of complexes is the number of columns of the matrix composed of the plurality of complexes, and the size of each of the plurality of complexes is related to the number of rows of the matrix composed of the plurality of complexes.
- the complex generator 212 may receive a security strength setting request from the analysis requester to generate a plurality of complexes that satisfy the requested security strength.
- the complex generating unit 212 generates a plurality of complexes each containing at least one of a target unit or a gastrointestinal unit, and transmits the same to the analyst, whereby the rate of increase of the nucleic acid sequence calculation amount of the analyzer is significantly increased compared to the rate of increase of the security level for the nucleic acid base sequence. It can provide a low security method.
- the encryption and decryption functions corresponding to the keys are kept at their own risk. When exposed, the encryption technology is disabled. Since security keys are a kind of information, their size is the key to security.
- the simplest process for decryption is the random key generator, a technique that tests all combinations. This allows theoretically all security keys to be decryptable. For example, a 4-digit security key can be decrypted in 10,000 attacks.
- the present invention targets genetic information, which is a strong identification information of an individual, it is more similar to a problem of transmitting a security key itself rather than a problem of encrypting and transmitting data.
- the present invention distributes the security key into at least two or more units which may include a camouflage key (both methods of extracting, segmenting, or combining elements constituting the security key are possible) to increase the security key.
- the complex generating unit 212 generates 10 complexes based on a nucleic acid base sequence of length 100, each complex having one target unit (derived from a nucleic acid base sequence of length 100) and four gastrointestinal units. By including the nucleic acid base sequence security key can be increased.
- the present invention does not change the length of the nucleic acid base sequence of the analysis requester input to the algorithm, while the calculation load of the analyzer 120 is increased by a multiple of the number of units included in each complex, so that the security level increase rate of j ⁇ i.
- the throughput increase i * j of the analyst 120 provides a significantly lower, very advantageous way.
- the nucleic acid nucleotide sequence security device 110 generates 10 complexes based on the nucleic acid sequence of the requester and generates 10 complexes by adding 4 gastrointestinal monomers to each target unit.
- the security level of the analyst 120 is increased by 5 * 10 times, which is significantly lower than the security level increase.
- the complex provider 214 provides the analyzer 120 with a plurality of complexes generated by the complex generator 212.
- the complex provider 214 may provide all of the generated complexes to one analyst or divide the generated complexes and provide them to a plurality of direct or indirect analysts.
- the complex providing unit 214 may generate six complexes so that the three complexes may be provided to Analyst A and the complexes may be provided to Analyst B.
- the analysis complex receiver 216 receives a plurality of analysis composites representing analysis results of the plurality of complexes from the analyzer 120.
- the complex receiver 216 may include a first analysis complex showing an analysis result for the first complex, a second analysis complex showing an analysis result for the second complex, and a third analysis showing an analysis result for the third complex.
- the complex can be received.
- the analysis complex analyzer 218 obtains an analysis result of the nucleic acid base sequence through the received plurality of analysis complexes.
- the analysis complex analyzer 218 determines a plurality of target analysis elements corresponding to the analysis results of the plurality of target units in the plurality of analysis complexes and uses the final target analysis elements for the nucleic acid base sequence of the analysis requester. Analyze results.
- the nucleic acid nucleotide sequence security control unit 219 controls the overall operation of the nucleic acid nucleotide sequence security device 110, the complex generating unit 212, complex providing unit 214, analysis complex receiving unit 216 and analysis complex analysis unit ( 218 may control the flow of data.
- the memory 220 includes a sparse matrix memory area (S.M.M.A) and a target element map area (T.E.M.A).
- S.M.M.A sparse matrix memory area
- T.E.M.A target element map area
- the sparse matrix memory region S.M.M.A corresponds to a space for storing the sparse matrix 222 for the plurality of complexes 223 and the plurality of target units included in the complex 223.
- Each of the plurality of complexes 223 may include at least one of a target unit (T.E) or a gastrointestinal unit (D.E).
- T.E target unit
- D.E gastrointestinal unit
- complexes 1, 2, 3, and 5 (223a, 223b, 223c, and 223e) contain one target unit and four gastrointestinal units
- complex 4 (223d) contains five gastrointestinal units.
- the plurality of complexes 223 include four target units and 21 gastrointestinal units, and the plurality of target units and gastrointestinal units included in the plurality of complexes 223 form a 5-row, 5-column matrix. Can be configured (each of the plurality of complexes 223 located in a column of the constructed matrix). Accordingly, the plurality of complexes 223 may be represented by the sparse matrix 222 for the plurality of target units.
- the sparse matrix 222 for the plurality of target units is the target unit of complex 1 (223a) in row 2, column 1, the target monomer of complex 2 (223b) in row 4, column 2, and the complex 3 (in row 1, column 3).
- Target monomer of 223c) and row 5, column 5, target monomer of complex 5 (223e) is included as a cell of the matrix.
- the target monomer map region T.E.M.A corresponds to a space for storing the target monomer map 224 for the sparse matrix 222 stored in the sparse matrix memory region S.M.M.A.
- the target monomer map 224 corresponds to a map of the positions of the plurality of target monomers included in the sparse matrix 222.
- the target monomer map 224 is a position from the target monomer in the first column to the target monomer in the fifth column when the number of columns of the sparse matrix 222 stored in the sparse matrix memory area SMMA corresponds to five. It includes.
- the target monomer map 224 stores the position of the target monomer with respect to the corresponding column of the sparse matrix 222 as a negative value (for example, -1) when the corresponding column of the sparse matrix 222 does not include the target monomer. Can be.
- the network interface 230 includes an environment for connecting with the analyst 120 through a network, and may include, for example, an adapter for local area network (LAN) communication.
- LAN local area network
- the user input device 240 includes an environment for receiving user input and may include, for example, an adapter such as a mouse, trackball, touch pad, graphics tablet, scanner, touch screen, keyboard, or pointing device.
- an adapter such as a mouse, trackball, touch pad, graphics tablet, scanner, touch screen, keyboard, or pointing device.
- the user output device 250 may include an environment for outputting specific information (eg, a nucleic acid sequence analysis result of the analysis requester) to the user, and may include an adapter such as a monitor or a touch screen. .
- the user input device 240 and the user output device 250 may be connected via a remote connection.
- the storage device 260 may be implemented as a nonvolatile memory such as a solid state disk (SSD) or a hard disk drive (HDD), and is used to store data necessary for the nucleic acid base sequence security device 110.
- SSD solid state disk
- HDD hard disk drive
- FIG. 3 is a diagram illustrating a procedure of generating a plurality of complexes and analyzing a plurality of analyte complexes according to an embodiment of the present invention.
- each base has a locus (the locus can be expressed as a unique address / coordinate of a base on a chromosome, a kind of coordinate value, the 1234501st position on chromosome 12, or an rsID).
- the locus can be expressed as a unique address / coordinate of a base on a chromosome, a kind of coordinate value, the 1234501st position on chromosome 12, or an rsID).
- VCF Variant Call Format
- the standardized format for containing the nucleic acid sequence of the requester of the assay and the locus + base + other information of the nucleic acid sequence included in the target unit, gastrointestinal unit, complex or assay complex is VCF (Variant CallFormat), BCF.
- GFF Gene-Finding Format
- GVF GenericFeature Format
- SAM / BAM Sequence Alignment Map / Binary version of SAM
- QUAL SCARF
- QSEQ QSEQ
- IG Maq
- SOAP SOAP
- bcf pileup
- mpileup CASAVA
- MaCH GLFv2
- BED15 BED detail
- BEDPE bedGraph
- bigBed bigWig, Chain
- GenePred table HAL
- It may include HDF5, MAF, Net, Personal Genome SNP, PSL, WIG (Wiggleformat), .2bit, .nib, CSFASTQ, CSFASTA, FASTA, FASTQ formats, or extensions thereof.
- the target nucleic acid sequence may be sent by selecting only a selected specific base (marker) such as SNP (single nucleotide polymorphism), or may send a continuous whole sequence. Since the gene sequence is the sequence itself in which the bases are arranged in a certain order, and the present invention relates to the security of the transferred target nucleic acid base sequence, it is possible whether the transfer sequence is the original sequence or a partially selected sequence.
- the target nucleic acid sequence is a single nucleotide polymorphism (SNP) of a specific gene or nongenic locus, a dimeric mutation or short tandem repeat polymorphism (STRP) or VNTR (various number) including substitution, addition or deletion of bases. bases determined in polyalleic variants, including oftandem repeats.
- the complex generator 212 extracts at least one base and the locus from the nucleic acid base sequence of the requester to generate a target unit.
- the complex generating unit 212 extracts at least one base and the locus from the nucleic acid base sequence of the requester to generate a plurality of complexes.
- the nucleic acid bases of the target units T.E1 and ACGCA having the nucleic acid base sequence of GAAT are extracted.
- Target monomer T.E2 having a sequence
- target monomer T.E3 having a nucleic acid base sequence of TCCTGAT
- target monomer T.E4 having a nucleic acid base sequence of GACAC
- target monomer T.E5 having a nucleic acid base sequence of CCAGCG Can be.
- the complex generating unit 212 generates a gastrointestinal unit constituting a plurality of complexes.
- the complex generating unit 212 may generate at least one gastrointestinal unit in the complex based on the similarity with the generated target unit.
- the complex generator 212 may generate at least one gastrointestinal unit whose genetic distance or evolutionary distance from the generated target unit is less than a specific distance.
- the complex generating unit 212 includes a complex C1 including the target unit T.E1, a complex C2 including the target unit T.E2, a complex C3 including the target unit TE 3, and a target unit T.E4.
- Complex C4 target monomer, complex C5 comprising target monomer TE 5 and complex C6 comprising only a plurality of gastrointestinal units without the target monomer can be generated.
- the complex generator 212 may determine the position of the target unit included in the complex while generating a plurality of complexes C1 to C6.
- the complex generating unit 212 includes the target unit T.E1 at position 3 of complex C1, the target unit T.E2 at position 1 of complex C2, and the target unit T.E3 at position 4 of complex C3. In position, the target unit T.E4 can be placed at position 1 of complex C4 and finally the target unit T.E5 can be placed at position 3 of complex C5.
- the generated plurality of complexes C1 to C6 are provided to the analyst, and the analyst 120 analyzes the provided complexes C1 to C6 to indicate an analysis result of the plurality of complexes C1 to C6.
- a plurality of analysis composites (analysis composite 1 to analysis composite 6 and A.C1 to A.C6) are provided to the analysis composite receiver 216.
- the analyzer 120 generates possible analysis result values for the fragmented nucleic acid base sequences. For example, when the complex generator 212 generates 10 complexes, and each complex includes one target unit and four gastrointestinal units, the analyst 120 generates and generates 50 analysis results.
- the plurality of analysis complexes including the analyzed result is transmitted to the nucleic acid base sequence security device (110).
- the nucleic acid base sequence security device 110 receives 50 analysis results and refers to each analysis complex by referring to a target monomer base sequence or a target monomer generation rule (for example, target monomer position information such as the target monomer map 224).
- 10 results of the analysis or analysis of the target monomers (10 target analysis units) can be extracted and combined to obtain the same result as the final result of the analysis or analysis of the nucleic acid sequence of the requester.
- a function for combining 10 result values of the target monomers extracted from the analysis complex may be promised in advance, generated / transmitted by an analysis institution, or requested by a client (eg, an analysis requester).
- the function includes various calculation methods, such as multiplying ten numbers, adding or calculating an average value.
- the plurality of analysis complexes A.C1 to A.C6 correspond to a plurality of target units T.E1 to T.E5 respectively included in the plurality of complexes C1 to C6.
- Target interpreting monomers R13, R21, R34, R41 and R53. Since the nucleic acid nucleotide sequence security device 110 holds the nucleic acid nucleotide sequence information of the target unit and the gastrointestinal unit requested for analysis, the nucleic acid nucleotide sequence information of the corresponding target unit for each complex (in FIG.
- T.E1 is GAAT
- T Clearly extract analysis results (here R13, R21, R34, R41, and R53) that are logically linked to sequences that match ACGCA, T.E3 to TCCTGAT, T.E4 GACAC, and T.E5 to CCAGCG can do.
- the analyzer 120 may reduce the amount of information through encoding or the like (logically still identifiable) without necessarily including and transmitting the nucleic acid sequence information of the target unit and the gastrointestinal unit in each unit included in the plurality of analysis complexes. Can be reduced or encoded).
- the plurality of target interpreting units is a nucleotide sequence resulting from the nucleotide sequence constituting the plurality of target monomers and the analysis result of the corresponding nucleic acid sequence It can mean a set of data pairs.
- the plurality of target interpreting units are not shown or omitted nucleic acid sequence information constituting the plurality of target units and given for identification of the corresponding nucleic acid base sequence. Refers to a set of identifier-result data pairs replaced with identifiers.
- the analysis complex analyzer 218 may include the plurality of target monomers T.E1 to T.E5 in the plurality of analysis complexes A.C1 to A.C6 received from the analyst based on the target monomer map 224.
- a plurality of target analysis units R13, R21, R34, R41, and R53 corresponding to can be determined.
- the analysis complex analyzer 218 calculates a plurality of determined target analysis units (R13, R21, R34, R41, and R53) using a specific function f (Rk), and performs a final analysis on the nucleic acid base sequence of the requester of the analysis.
- the results can be derived.
- the analysis complex analyzer 218 may combine the determined plurality of target analysis units (R13, R21, R34, R41, and R53) to calculate the posterior odds of the analysis requester and calculate the calculated analysis.
- the requester's post ozone may provide a final analysis of the nucleic acid base sequence of the requester.
- the post ozone calculation process of the analysis requester performed by the analysis complex analyzer 218 will be described in detail.
- the analysis complex analysis unit 218 may calculate the posterior odds of the nucleic acid base sequence of the requester according to Bayes Theorem (the disease risk rate of an individual having the nucleic acid base sequence).
- Bayes Theorem the disease risk rate of an individual having the nucleic acid base sequence.
- Posteror Odds may be expressed as a product of Prior Odds and Likelihood Ratio, as shown in Equation 1 below.
- the dictionary ozone may be replaced with a neutral value of 1 when there is no prior knowledge. Therefore, the post-oz can be obtained through a chain calculation by continuously multiplying the likelihood ratio by the pre-oz value 1.
- the likelihood ratio according to the genotype may be calculated as in Equation 2 below.
- a nucleic acid base sequence consisting of bases G (a1), G (a2), ..., G (an) obtained from n locus a1, a2, ... an [a1, a2, .. given the likelihood ratios LR (G (a1)), LR (G (a2)), ..., LR (G (an)) for each position of The post ozone corresponding to the risk may be calculated by Equation 3 below.
- Rk of Equation 4 may correspond to an analysis result value of the analysis result base sequence for the target monomer included in the plurality of analysis complexes.
- Equation 4 is only k and Rk, and thus may be predetermined between the analysis requester and the analyst 120, may be transmitted by the analyst 120 with a plurality of analysis complexes, A corresponding equation required for the calculation may be requested together with the analysis result values R1 to Rk of each of the at least one target unit.
- Equation 4 is established regardless of the number of target units or the length of the nucleic acid base sequence of each target unit. Therefore, the nucleic acid nucleotide sequence security device 110 may determine the number of target monomers to be any number less than or equal to n. However, segmenting the positions of nucleic acid sequences of the requestor for analysis one by one is likelihood LR (G (a1)) and LR (G (a2)) for each position, which are the core intellectual property from the perspective of the analyst 120 performing the analysis. This is not appropriate because it exposes all of LR (G (an)). The nucleic acid base sequence security device 110 needs to generate a target unit having a length of at least two for intellectual property protection, such as an algorithm used for data processing by the analyst 120.
- the length of the target monomer can vary depending on the objective of the security strength setting. Since the security strength increases exponentially with respect to the target monomer length, it is easy to set a high security level. On the other hand, if the length of the target unit is too long to approach the length of the nucleic acid sequence of the requester, the probability of exposure of the target nucleic acid sequence of the sender is relatively increased, and the number of gastrointestinal units contained in the complex must be increased to defend against this. . However, since the number of gastrointestinal units and the security strength only increase linearly, the number of gasoline units must be increased greatly, which is difficult to see in a preferable way by increasing the analysis load of the analyst 120 exponentially.
- the present invention adjusts the number of complexes to protect the target nucleic acid sequence, the number of target units included in the plurality of complexes, or the number of gastrointestinal units included in each complex, thereby controlling the privacy of the requester and the analyst 120.
- IP asset security level requirements That is, the present invention may satisfy the requirement of the nucleic acid sequence security of the analysis requester and the knowledge asset security of the analyst 120 by adjusting the number of target units, the number of gastrointestinal units, and the number of complexes.
- FIG. 3 illustrates a case in which the complex generating unit 212 generates six complexes.
- the complex C1 to the complex C5 include one target unit and about four gastrointestinal units, and C6 includes only four gastrointestinal units. It is included.
- 32-bit security level can be satisfied by generating four complexes, eight complexes each comprising 16 units, four complexes each of 256 units, or two complexes each comprising 65546 units.
- the present invention may first determine the information protection level, and then determine the number of complexes and the number of units included in the complex in consideration of the calculation load of the analyst 120.
- FIG. 4 is a diagram illustrating a procedure of generating a plurality of complexes and analyzing a plurality of analyte complexes according to another embodiment of the present invention.
- the complex generator 212 segments the nucleic acid base sequence of the analysis requester into partial nucleotide sequences to generate the target unit.
- the complex generating unit 212 targets T.E1 having a nucleic acid base sequence of GGAA and a nucleic acid base sequence of TCAAC in a manner of segmenting a nucleic acid base sequence of an analysis requester into a partial base sequence to generate a plurality of complexes.
- a target unit T.E3 having a nucleic acid base sequence of units T.E2, CGGCGGA, a target unit T.E4 having a nucleic acid base sequence of CTGAT, or a target unit T.E5 having a nucleic acid base sequence of TACACCC can be generated.
- the procedure performed in the nucleic acid base sequence security device 110 is as shown in FIG. 3.
- FIG. 5A is a diagram visualizing a sparse matrix for a target unit according to an embodiment of the present invention
- FIG. 5B is a diagram visualizing a sparse matrix for a target unit according to another embodiment of the present invention.
- the complex generating unit 212 may define a target unit cell in a sparse matrix as a base-location set including at least one base and a locus for the target unit.
- a target monomer cell in a sparse matrix is a nucleic acid base sequence consisting of base A with a locus of 12, base G with a locus of 15, and base G with a locus of 332, if the target monomer corresponds to ⁇ (A , 12), (G, 15), (G, 332) ⁇ .
- the complex generator 212 may define a target unit cell in the sparse matrix as at least one base set associated with the locus.
- a target monomer cell in a sparse matrix corresponds to a nucleic acid base sequence whose target monomer is composed of base A with the locus of 12 loci, base G with the locus of 15, and base G with the locus of 332. It can be defined as a set of ⁇ A, G, G ⁇ .
- the loci 12,15 and 332 of the bases can be expressed in common in the matrix.
- FIG. 6 is a flowchart illustrating a nucleic acid nucleotide sequence security method performed by the nucleic acid nucleotide sequence security device of FIG. 2.
- the complex generator 212 generates a plurality of complexes each including at least one of a target unit derived from the nucleic acid sequence of the requester or a target unit identical to or different from the target unit (S601).
- the complex generating unit 212 may express the plurality of complexes as a sparse matrix for the plurality of target units (S602).
- the complex generation unit 212 may determine the position of the target unit when the target unit is included in the complex (S603).
- the complex generation unit 212 may generate a target monomer map necessary in the decoding process by dynamically determining a position of the target monomer cell in the sparse matrix while expressing the plurality of complexes as a sparse matrix for the plurality of target units. (S604)
- the complex providing unit 214 may provide the analyst with the generated plurality of complexes, and the analysis complex receiving unit 216 may receive a plurality of analysis complexes indicating analysis results of the plurality of complexes from the analyst (S605 and S606). .
- the analysis complex analysis unit 218 may determine a plurality of target analysis units representing analysis results of the plurality of target units based on the location information of the target units stored in the target unit map (S607). Analysis complex analysis unit 218 may combine the determined plurality of target analysis units to calculate the post-hose of the analysis requester, and obtain the final analysis result of the nucleic acid base sequence of the analysis requester through the calculated post-operation (S608 and S609).
- the present invention relates to a nucleic acid nucleotide sequence security technology, and more particularly, to a nucleic acid nucleotide sequence security method, apparatus and apparatus for analyzing the nucleic acid nucleotide sequence of an analysis requester without exposing the nucleic acid nucleotide sequence of the requester to the analyst. It relates to a recording medium.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Health & Medical Sciences (AREA)
- General Health & Medical Sciences (AREA)
- Physics & Mathematics (AREA)
- Computer Security & Cryptography (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Computer Hardware Design (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Bioethics (AREA)
- Life Sciences & Earth Sciences (AREA)
- Software Systems (AREA)
- Bioinformatics & Computational Biology (AREA)
- Biotechnology (AREA)
- Evolutionary Biology (AREA)
- Medical Informatics (AREA)
- Spectroscopy & Molecular Physics (AREA)
- Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
- Apparatus Associated With Microorganisms And Enzymes (AREA)
Abstract
핵산염기서열 보안 방법은 (a) 각각이 분석 요청자의 핵산염기서열로부터 파생된 표적 단위체(target element) 또는 상기 표적 단위체와 동일하거나 또는 다른 위장 단위체(disguising element) 중 적어도 하나를 포함하는 복수의 복합체들(composites)을 생성하는 단계 및(b) 상기 생성된 복수의 복합체들을 분석자에게 제공하는 단계를 포함한다.
Description
본 발명은 핵산염기서열 보안 기술에 관한 것으로, 보다 상세하게는 분석 요청자의 핵산염기서열을 분석자에게 노출하지 않고 분석 요청자의 핵산염기서열을 분석할 수 있는 핵산염기서열 보안 방법, 장치 및 이를 저장한 기록매체에 관한 것이다.
게놈 (Genome) 또는 유전체는 한 개체가 갖는 염색체의 총 염기서열을 의미한다. 유전체는 한 생물종의 거의 완전한 유전정보의 총합이고 핵산염기서열 정보를 저장하고 있다. 인간의 유전체는 한 인간 개체를 만들기 위해 필요한 모든 유전자들과 유전자 이외의 부분을 포함하는 약 32억쌍 정도의 모든 핵산염기서열로, 22쌍 44개의 상염색체와 1쌍의 성염색체(X,Y) 및 미토콘드리아에 나뉘어져 있다. 유전체의 핵산은 인산, 당과 함께 아데닌(A), 티아민(T), 구아닌(G) 및 시토신(C) 염기로 이루어진 이중나선형의 물질이다. 유전정보는 핵산의 위 4 가지 염기의 서열의 배열에 의해 부호화된다.
사람의 유전자는 각각 수백개 또는 수천개의 아미노산배열을 가지며 몸에서 생산하는 모든 단백질의 설계도를 제공한다. 사람의 유전자는 약 3만~5만개로 추정되며 전체 유전체의 약 3~5%를 차지한다. 약 3만~5만개에 달하는 사람의 유전자는 염기서열로 부호화되어 있고 이중 극히 일 부분만이 그 의미가 밝혀져 있다. 유전체의 핵산염기서열 중 유전자 영역에 속하지 않는 95~97% 부위에 존재하는 핵산염기서열의 유전형에도 개체의 표현형에 관련되는 수많은 부위가 있고 그에 대한 의학적 의미도 지속적으로 밝혀지고 있다. 염기서열로 부호화 되어있는 유전체 정보의 의학적, 생물학적 의미의 해석은 지속적으로 발전하고 있다. 따라서, 염기서열을 일차분석한 개인은 새롭게 밝혀진 유전체 염기서열 해석정보에 대한 본인의 해당여부를 확인하기 위해 본인의 염기서열정보에 대한 반복적인 조회 및 해석을 필요로 할 수 있다.
염기서열분석기술 발달은 분석 시에 요구되는 비용을 감소시켰고 각 개인의 인간 유전체 분석과 유전체 정보의 활용을 가능하게 하였다. 이에 따라, 많은 사람들은 자신의 유전적 특성인 유전체 염기서열정보를 보유하고 있고 질병 발생률을 포함한 관련된 표현형 위험률은 개인이 보유한 유전체 염기서열 정보를 통해 계산될 수 있다. 따라서, 개인의 유전체 염기서열 정보는 노출이 부담스러운 민감정보로 분류되며 가급적 타인에게 노출되지 않는 것이 좋다. 그러나, 개인이 소유한 자신의 유전형 정보의 해석에 필요한 계산방법 및 그에 필요한 정보자원은 정보해독용 분석장비와 분석방법 등을 보유한 분석자 또는 분석기관이 보유하고 있어서, 개인은 자신의 유전형 정보로부터 자신의 표현형 발생확률을 계산하기 위해 분석자 또는 분석기관에게 자신의 유전형 정보를 전송해야 한다. 이러한 유전형 정보 전송은 개인의 유전형 정보 노출위험을 증가시킨다는 문제가 있다.
대한민국공개특허 2013-0075559는 유전자 정보 관리 방법에 관한것으로, 보다 구체적으로, 개인의 유전자 정보를 복수개의 부분서열로 분할하고 저장하여 관리함으로써, 저장 정보에 대한 해킹 혹은 유출 등이 일어난 경우에 그 분할저장된 정보를 분할전의 원 상태로 복원하는 것을 어렵게 하는 방법을 개시한다. 상기의 특허는 민감한 유전자 정보의 저장보관시의 보안에 관한 것으로, 분석대상 유전정보를 보유한 개인 또는 기관이 핵산염기서열 정보를 분석자 또는 분석기관으로 전송하고 그 분석결과를 전송받는 과정에 관여되는 정보 보안 관리 기술에 관한 본 발명과 구별된다. 상기 공개특허는 구분자 정보가 노출되면 개인의 유전자 정보가 노출될 수 있고, 구분자 정보 자체의 보안도 취약할 수 있다는 문제점을 갖는다.
본 발명의 일 실시예는 분석 요청자의 핵산염기서열을 분석자에게 노출하지 않고 분석 요청자의 핵산염기서열을 분석할 수 있는 핵산염기서열 보안 방법을 제공하고자 한다.
본 발명의 일 실시예는 분석 요청자의 핵산염기서열로부터 파생된 표적 단위체와 위장 단위체를 각각 포함하는 복합체를 표현하고 복수의 표적 단위체들의 관점에서 정의된 희소 행렬을 통해 분석 요청자의 핵산염기서열을 분석할 수 있는 핵산염기서열 보안 방법을 제공하고자 한다.
본 발명의 일 실시예는 표적 단위체와 위장 단위체를 각각 포함하는 복수의 복합체들의 개수 또는 복수의 복합체들 각각의 크기를 결정하여 핵산염기서열의 보안강도를 설정할 수 있는 핵산염기서열 보안 방법을 제공하고자 한다.
본 발명의 일 실시예는 분석자로부터 복수의 복합체들에 대한 해석결과를 수신하여 분석 요청자의 핵산염기서열의 노출 없이 핵산염기서열에 대한 분석 결과를 획득할 수 있는 핵산염기서열 보안 방법을 제공하고자 한다.
본 발명의 일 실시예는 분석자의 관점에서 핵산염기서열에 대한 단위체 단위로 해석결과를 송신함으로써 분석자가 보유한 개별 염기에 대한 우도비(odds ratio)와 같은 고유의 정보나 노하우를 지킬 수 있는 핵산염기서열 보안 방법을 제공하고자 한다.
본 발명의 일 실시예는 핵산염기서열 보안 장치의 계산량을 최소화하면서도 보안의 수준을 향상시킬 수 있는 핵산염기서열 보안 방법을 제공하고자 한다.
실시예들 중에서, 핵산염기서열 보안 방법은 (a) 각각이 분석 요청자의 핵산염기서열로부터 파생된 표적 단위체(target element) 또는 상기 표적 단위체와 동일하거나 또는 다른 위장 단위체(disguising element) 중 적어도 하나를 포함하는 복수의 복합체들(composites)을 생성하는 단계 및(b) 상기 생성된 복수의 복합체들을 분석자에게 제공하는 단계를 포함한다.
상기 (a) 단계는 상기 복수의 복합체들을 복수의 표적 단위체들에 대한 희소 행렬로 표현하는 단계를 포함할 수 있다.상기 (a) 단계는 해당 표적 단위체가 해당 복합체에 포함되면 상기 해당 표적 단위체의 위치를 결정하는 단계를 더 포함할 수 있다.
일 실시예에서, 상기 (a) 단계는 상기 희소 행렬에 있는 표적 단위체 셀을 상기 표적 단위체에 대한 적어도 하나의 염기와 유전좌위를 포함하는 염기-좌위 집합으로 정의하는 단계를 더 포함할 수 있다. 다른 일 실시예에서, 상기 (a) 단계는 상기 희소 행렬에 있는 표적 단위체 셀을 유전좌위와 연관된 적어도 하나의 염기 집합으로 정의하는 단계를 더 포함할 수 있다. 상기 (a) 단계는 상기 희소 행렬에 있는 표적 단위체 셀의 위치를 동적으로 결정하여 디코딩 과정에서 필요한 표적 단위체 맵을 생성하는 단계를 더 포함할 수 있다.
일 실시예에서, 상기 (a) 단계는 상기 핵산염기서열로부터 적어도 하나의 염기와 유전좌위를 추출하여 상기 표적 단위체를 생성하는 단계를 포함할 수 있다. 다른 일 실시예에서, 상기 (a) 단계는 상기 핵산염기서열을 부분 염기서열로 분절하여 상기 표적 단위체를 생성하는 단계를 포함할 수 있다.
상기 (a) 단계는 상기 표적 단위체와의 유사도를 기초로 해당 복합체에 있는 적어도 하나의 위장 단위체를 생성하는 단계를 포함할 수 있다. 상기 (a) 단계는상기 표적 단위체와의 유전자 거리 또는 진화적 거리가 특정 거리 이하인 적어도 하나의 위장 단위체를 생성하는 단계를 더 포함할 수 있다. 상기 (a) 단계는상기 분석 요청자에 의하여 설정된 보안 강도에 따라 상기 복수의 복합체들의 개수 또는 상기 복수의 복합체들 각각의 크기를 결정하는 단계를 포함할 수 있다.
상기 (b) 단계는 상기 생성된 복수의 복합체들을 분할하여 복수의 직접 또는 간접 분석자들에게 제공하는 단계를 포함할 수 있다.
상기 방법은 (c) 상기 분석자로부터 상기 복수의 복합체들에 대한 해석결과를 나타내는 복수의 분석 복합체(analysis composites)들을 수신하여 상기 핵산염기서열에 대한 분석 결과를 획득하는 단계를 더 포함할 수 있다.
상기 (c) 단계는 표적 단위체 맵을 기초로 복수의 표적 단위체들에 대한 해석결과를 나타내는 복수의 표적 분석 단위체들(target analysis elements)을 결정하는 단계를 포함할 수 있다. 상기 (c) 단계는 상기 결정된 복수의 표적 분석 단위체들을 조합하여 상기 분석 요청자의 사후오즈(Posterior Odds)를 계산하는 단계를 더 포함할 수 있다.
실시예들 중에서, 핵산염기서열 보안 장치는 각각이 분석 요청자의 핵산염기서열로부터 파생된 표적 단위체(target element) 또는 상기 표적 단위체와 동일하거나 또는 다른 위장 단위체(disguising element) 중 적어도 하나를 포함하는 복수의 복합체들(composites)을 생성하는 복합체 생성부 및 상기 생성된 복수의 복합체들을 분석자에게 제공하는 복합체 제공부를 포함한다.
상기 장치는 상기 분석자로부터 상기 복수의 복합체들에 대한 해석결과를 나타내는 복수의 분석 복합체들을 수신하여 상기 핵산염기서열에 대한 분석 결과를 획득하는 분석 복합체 해석부를 더 포함할 수 있다.
상기 분석 복합체 해석부는 표적 단위체 맵을 기초로 복수의 표적 단위체들에 대한 해석결과를 나타내는 복수의 표적 분석 단위체들(target analysis elements)을 결정할 수 있다. 상기 분석 복합체 해석부는 상기 결정된 복수의 표적 분석 단위체들을 조합하여 상기 분석 요청자의 사후오즈(Posterior Odds)를 계산할 수 있다.
실시예들 중에서, 핵산염기서열 보안 방법에 관한 컴퓨터 프로그램을 기록한 기록매체는 각각이 분석 요청자의 핵산염기서열로부터 파생된 표적 단위체(target element) 또는 상기 표적 단위체와 동일하거나 또는 다른 위장 단위체(disguising element) 중 적어도 하나를 포함하는 복수의 복합체들(composites)을 생성하는 기능 및상기 생성된 복수의 복합체들을 분석자에게 제공하는 기능을 포함한다.
상기 기록매체는 상기 분석자로부터 상기 복수의 복합체들에 대한 해석결과를 나타내는 복수의 분석 복합체들을 수신하여 상기 핵산염기서열에 대한 분석 결과를 획득하는 기능을 더 포함할 수 있다.
개시된 기술은 다음의 효과를 가질 수 있다. 다만, 특정 실시예가 다음의 효과를 전부 포함하여야 한다거나 다음의 효과만을 포함하여야 한다는 의미는 아니므로, 개시된 기술의 권리범위는 이에 의하여 제한되는 것으로 이해되어서는 아니 될 것이다.
본 발명의 일 실시예에 따른 핵산염기서열 보안 방법은 분석 요청자의 핵산염기서열을 분석자에게 노출하지 않고 분석 요청자의 핵산염기서열을 분석할수 있다.
본 발명의 일 실시예에 따른 핵산염기서열 보안 방법은 분석 요청자의 핵산염기서열로부터 파생된 표적 단위체와 위장 단위체를 각각 포함하는 복합체를 표현하고 복수의 표적 단위체들의 관점에서 정의된 희소 행렬을 통해 분석 요청자의 핵산염기서열을 분석할 수 있다.
본 발명의 일 실시예에 따른 핵산염기서열 보안 방법은 표적 단위체와 위장 단위체를 각각 포함하는 복수의 복합체들의 개수 또는 복수의 복합체들 각각의 크기를 결정하여 핵산염기서열의 보안강도를 설정할 수 있다.
본 발명의 일 실시예에 따른 핵산염기서열 보안 방법은 복수의 복합체들에 대한 해석결과를 수신하여 분석 요청자의 핵산염기서열의 노출 없이 핵산염기서열에 대한 분석 결과를 획득할 수 있다.
본 발명의 일 실시예에 따른 핵산염기서열 보안 방법은 분석자의 관점에서 핵산염기서열에 대한 단위체 단위로 해석결과를 송신함으로써 분석자가 보유한 개별 염기에 대한 우도비(odds ratio)와 같은 고유의 정보나 노하우를 지킬 수 있다.
본 발명의 일 실시예에 따른 핵산염기서열 보안 방법은 핵산염기서열 보안 장치의 계산량을 최소화하면서도 보안의 수준을 향상시킬 수 있다.
도 1은 본 발명의 일 실시예에 따른 핵산염기서열 보안 시스템을 설명하는 도면이다.
도 2는 도 1에 있는 핵산염기서열 보안장치를 설명하는 블록도이다.
도 3은 본 발명의 일 실시예에 따른 복수의 복합체들 생성 및 복수의 분석 복합체들 해석 프로시저를 설명하는 도면이다.
도4은 본 발명의 다른 일 실시예에 따른 복수의 복합체들 생성 및 복수의 분석 복합체들 해석프로시저를 설명하는 도면이다.
도 5a는 본 발명의 일 실시예에 따른 표적 단위체에 대한 희소 행렬을 시각화한 도면이고, 도 5b는 본 발명의 다른 일 실시예에 따른 표적 단위체에 대한 희소 행렬을 시각화한 도면이다.
도 6은 도 2에 있는 핵산염기서열 보안 장치에 의하여 수행되는 핵산염기서열 보안 방법을 설명하는 흐름도이다.
본 발명에 관한 설명은 구조적 내지 기능적 설명을 위한 실시예에 불과하므로, 본 발명의 권리범위는 본문에 설명된 실시예에 의하여 제한되는 것으로 해석되어서는 아니 된다. 즉, 실시예는 다양한 변경이 가능하고 여러 가지 형태를 가질 수 있으므로 본 발명의 권리범위는 기술적 사상을 실현할 수 있는 균등물들을 포함하는 것으로 이해되어야 한다. 또한, 본 발명에서 제시된 목적 또는 효과는 특정 실시예가 이를 전부 포함하여야 한다거나 그러한 효과만을 포함하여야 한다는 의미는 아니므로, 본 발명의 권리범위는 이에 의하여 제한되는 것으로 이해되어서는 아니 될 것이다.
한편, 본 출원에서 서술되는 용어의 의미는 다음과 같이 이해되어야 할 것이다.
"제1", "제2" 등의 용어는 하나의 구성요소를 다른 구성요소로부터 구별하기 위한 것으로, 이들 용어들에 의해 권리범위가 한정되어서는 아니 된다. 예를 들어, 제1 구성요소는 제2 구성요소로 명명될 수 있고, 유사하게 제2 구성요소도 제1 구성요소로 명명될 수 있다.
어떤 구성요소가 다른 구성요소에 "연결되어"있다고 언급된 때에는, 그 다른 구성요소에 직접적으로 연결될 수도 있지만, 중간에 다른 구성요소가 존재할 수도 있다고 이해되어야 할 것이다. 반면에, 어떤 구성요소가 다른 구성요소에 "직접 연결되어"있다고 언급된 때에는 중간에 다른 구성요소가 존재하지 않는 것으로 이해되어야 할 것이다. 한편, 구성요소들 간의 관계를 설명하는 다른 표현들, 즉 "~사이에"와 "바로 ~사이에" 또는 "~에 이웃하는"과 "~에 직접 이웃하는" 등도 마찬가지로 해석되어야 한다.
단수의 표현은 문맥상 명백하게 다르게 뜻하지 않는 한 복수의 표현을 포함하는 것으로 이해되어야 하고, "포함하다"또는 "가지다" 등의 용어는 실시된 특징, 숫자, 단계, 동작, 구성요소, 부분품 또는 이들을 조합한 것이 존재함을 지정하려는 것이며, 하나 또는 그 이상의 다른 특징이나 숫자, 단계, 동작, 구성요소, 부분품 또는 이들을 조합한 것들의 존재 또는 부가 가능성을 미리 배제하지 않는 것으로 이해되어야 한다.
각 단계들에 있어 식별부호(예를 들어, a, b, c 등)는 설명의 편의를 위하여 사용되는 것으로 식별부호는 각 단계들의 순서를 설명하는 것이 아니며, 각 단계들은 문맥상 명백하게 특정 순서를 기재하지 않는 이상 명기된 순서와 다르게 일어날 수 있다. 즉, 각 단계들은 명기된 순서와 동일하게 일어날 수도 있고 실질적으로 동시에 수행될 수도 있으며 반대의 순서대로 수행될 수도 있다.
본 발명은 컴퓨터가 읽을 수 있는 기록매체에 컴퓨터가 읽을 수 있는 코드로서 구현될 수 있고, 컴퓨터가 읽을 수 있는 기록 매체는 컴퓨터 시스템에 의하여 읽혀질 수 있는 데이터가 저장되는 모든 종류의 기록 장치를 포함한다. 컴퓨터가 읽을 수 있는 기록 매체의 예로는 ROM, RAM, CD-ROM, 자기 테이프, 플로피 디스크, 광 데이터 저장 장치 등이 있으며, 또한, 캐리어 웨이브(예를 들어 인터넷을 통한 전송)의 형태로 구현되는 것도 포함한다. 또한, 컴퓨터가 읽을 수 있는 기록 매체는 네트워크로 연결된 컴퓨터 시스템에 분산되어, 분산 방식으로 컴퓨터가 읽을 수 있는 코드가 저장되고 실행될 수 있다.
염기는 핵염기(nucleobase) 혹은 줄여서 염기(base)로 불리기도 하며, DNA(deoxyribonucleic acid), RNA(ribonucleic acid), 뉴클레오타이드 및 뉴클레오사이드에서 발견되는 질소염기(nitrogenous bases)인 시토신(Cytosine), 구아닌(Guanine), 아데닌(Adenine), 티민(Thymine), 우라실(Uracil)을 포함할 수 있다. 뉴클레오타이드(nucleotide)는 DNA, RNA와 같은핵산을 구성하는 유기 분자로, 염기-당-인산의 결합으로 이루어진다. 핵산은 뉴클레오타이드가 긴 사슬 모양으로 중합된 고분자 유기물의 한 종류이다.
핵산염기서열은 배열을 통해 구현될 수 있고, 배열의 주소(즉, 인덱스)는 유전체 좌표값(이하, 유전좌위)(예, rsID 번호), 배열에 저장된 자료 값은 염기(예, A, G, T, C)로 구성된다. 한편, 여기에서, 염기는 유전자형을 포함할 수 있다. 유전좌위는 연속된 값일 수도 있고, 유전체의 서로 다른부위에서 필요한 부분들만을 골라서 추출한 불연속 좌표값의 목록일 수도 있다. 핵산염기서열은 그 유전좌위가 연속된 좌표값인 경우 배열의 주소 중 한 개만 명시해도 같은 배열내의 나머지 주소는 명확히 알 수 있으므로 배열의 주소 중 한 개만 명시하고 나머지는 생략할 수 있다.
개인 유전체 핵산염기서열은 한 개인의 유전체(genome)의 전체 핵산 염기서열을 포함하고, 분석 요청자의 핵산염기서열에 해당할 수 있다.
표적 핵산염기서열은 분석 요청자의 핵산염기서열과 연관된 복수의 표적단위체들에 해당하고, 개인 유전체 핵산염기서열에서 분석 대상 부위의 유전좌위-염기 쌍을 추출하고 일정한 순서로 배열하여 생성하거나 또는 개인 유전체 핵산염기서열을 분절하여 생성할 수 있다. 즉, 표적 핵산염기서열은 복수의 표적 단위체들을 일정한 순서로 연결하여 생성될 수 있다. 표적 핵산염기서열은 원래의 개인 유전체 핵산염기서열의 전체이거나, 개인 유전체 핵산염기서열에서 추출해낸 부분 염기서열이거나 또는 개인 유전체 핵산염기서열의 여러 부위에서 추출한 부분 염기서열을 조합한 염기서열일 수 있다.
본 발명의 일 실시예에서, 표적 핵산염기서열은 개인별 유전형 정보일 수 있다. 특히, 인간 유전체에 대한 염기서열 분석이 이루어진 뒤, 단순히 유전체 전체 염기서열 해독에 대한 분석뿐만 아니라 인종별, 개인별 다양성을 기반으로 한 단일염기다형성 (single nucleotide polymorphism, SNP) 유전체 염기서열 변이에 대한 분석이 활발하게 진행되는 현실이 고려되었다.다형성은 유전체상에 존재하는 염기서열의 개인간 차이를 말하는 것으로, 단일염기다형성은 A, T, C 및 G로 이루어진 염기서열 중 하나의 염기에 개인간 차이가 있는 것으로 유전자 다형성 중 그 수가 가장 많다. 인간의 유전자는 약 99.9% 일치하지만 약 0.1%의 단일 염기다형성 차이로 인해 체질, 외모, 질병 등 개인과 인종의 유전적 특성을 나타내게 되며 이로 인해 예컨대 사람마다 동일한 약을 사용해도 약의 효능 및 반응이 다르게 된다. 염기서열의 개인간 차이는 다양한 유전체 변이를 포함하고 이는 염기의 치환, 부가 또는 결실에 의한 것으로, 단일 염기다형성을 포함하는 SNV(Single Nucleotide Variation), STRP (short tandem repeatpolymorphism) 또는 VNTR (various number of tandem repeat) 및 CNV (Copy numbervariation)를 포함하는 다수체 (polyalleic) 변이의 형태로 나타날 수 있다.
표적단위체(target element)는 분석 요청자의 핵산염기서열에서 파생된 부분 염기서열에 해당한다. 표적 단위체는 분석 요청자의 핵산염기서열로부터 적어도 하나의 염기와 유전좌위를 추출하여 생성되거나 또는 분석 요청자의 핵산염기서열을 부분 염기서열로 분절하여 생성될 수 있다.
위장단위체(disguising element)는 표적 단위체의 식별을 어렵게 하기 위해 표적 단위체와 동일하거나 또는 다른 염기서열을 포함할 수 있다. 위장 단위체는 실제 자연계에 존재하는 핵산염기서열을 활용하거나 또는 실제 자연계에 존재하는 핵산염기서열을 참고하여 생성될 수 있다. 일 실시예에서, 위장 단위체는 표적 단위체와 일치하지 않는 부분 염기서열로 구성되거나, 표적 단위체와 부분적으로 일치하고 부분적으로 불일치하는 부분 염기서열로 구성되거나 또는 표적 단위체와 동일한 부분 염기서열로 구성될 수 있다. 다른 일 실시예에서, 위장 단위체는 분석 요청자의 핵산염기서열을 참고하거나, 무작위적으로 생성된 적어도 하나의 핵산염기서열을 적어도 두 개의 부분 염기서열들로 분절하여 생성되거나 또는 분석 요청자의 핵산염기서열을 분절하여 생성된 표적 단위체의 염기서열을 참고하거나 또는 무작위적으로 생성된 적어도 하나의 핵산염기서열로 구성될 수 있다.
복합체는 표적 단위체 또는 표적 단위체와 동일하거나 또는 다른 위장 단위체 중 적어도 하나를 포함한다. 즉, 복합체는 표적 단위체와 위장 단위체를 함께 포함하거나, 위장 단위체만을 포함하거나 또는 표적 단위체를 포함할 수 있다.
분석 복합체는 복합체에 대한 해석결과를 나타낼 수 있다. 예를 들어, 분석 복합체는 복합체에포함된단위체(들)에 각각 해당하는 핵산염기서열에 대한 해석결과의 집합을 포함할 수 있다. 다른 예를 들어, 분석 복합체는 복합체에포함된단위체(들)에 각각 해당하는 핵산염기서열 및 해당 해석결과를 함께 포함한 염기서열-해석결과 쌍의 집합을 포함할 수 있다. 분석 복합체는 복합체에 포함된 단위체(들)의 핵산염기서열 정보를 표기하지 않거나 생략하고 해당 핵산 염기서열의 식별을 위해 부여된 식별자로 대체한 양태의 해석결과도 포함할 수 있다. 본 발명은 이러한 복합체 및 분석 복합체의 크기(복합체에 포함되는 단위체의 개수) 또는 개수 변화를 통한 높은 수준의 보안을 가능하게 할 수 있다.
여기서 사용되는 모든 용어들은 다르게 정의되지 않는 한, 본 발명이 속하는 분야에서 통상의 지식을 가진 자에 의해 일반적으로 이해되는 것과 동일한 의미를 가진다. 일반적으로 사용되는 사전에 정의되어 있는 용어들은 관련 기술의 문맥상 가지는 의미와 일치하는 것으로 해석되어야 하며, 본 출원에서 명백하게 정의하지 않는 한 이상적이거나 과도하게 형식적인 의미를 지니는 것으로 해석될 수 없다.
도 1은 본 발명의 일 실시예에 따른 핵산염기서열 보안 시스템을 설명하는 도면이다.
도 1을 참조하면, 핵산염기서열 보안 시스템(100)은 핵산염기서열 보안장치(110) 및 분석 서버(이하, 분석자)(120)를 포함하고, 이들은 네트워크를 통해 연결될 수 있다.
핵산염기서열 보안 장치(110)는 분석 요청자의핵산염기서열을 기초로 생성된 복합체(composite)에 관한 분석을 요청할 수 있고, 예를 들어, 데스크탑, 노트북, 태블릿 PC 또는 스마트폰으로 구현될 수 있다. 분석 요청자의 핵산염기서열은 복수의 메모리 영역들(후술될 희소 행렬 메모리 영역 및 표적 단위체 맵 영역)을 통해 관리될 수 있다.
분석자(120)는 핵산염기서열 보안 장치(110)와 네트워크를 통해 연결되고 핵산염기서열 보안 장치(110)로부터 복합체(composite)를수신하고, 수신한복합체를 분석하여 분석 복합체(analysis composite)를 핵산염기서열 보안 장치(110)에 제공할 수 있다.
도 2는 도 1에 있는 핵산염기서열 보안 장치를 설명하는 블록도이다.
도 2를 참조하면, 핵산염기서열 보안 장치(110)는 프로세서(210), 메모리(220), 네트워크 인터페이스(230), 사용자 입력 장치(240) 및 사용자 출력 장치(250) 및 저장장치(260)를 포함한다.
프로세서(210)는 복합체 생성부(212), 복합체 제공부(214), 분석 복합체 수신부(216), 분석 복합체 해석부(218) 및 핵산염기서열 보안 제어부(219)를 포함한다.
복합체 생성부(212)는 각각이 분석 요청자의 핵산염기서열로부터 파생된 표적 단위체 또는 표적 단위체와 동일하거나 또는 다른 위장 단위체 중 적어도 하나를 포함하는 복수의 복합체들을 생성한다.복합체 생성부(212)는 생성된 복합체에 표적 단위체가 포함되면 해당 복합체에서의 해당 표적 단위체의 위치를 결정할 수 있다. 예를 들어, 복합체 생성부(212)는 한 개의 표적 단위체를 포함하는 제1 복합체, 네 개의 위장 단위체를 포함하는 제2 복합체 및 한 개의 표적 단위체와 한 개의 위장 단위체를 포함하는 제3 복합체를 생성하고 제1 복합체 및 제3 복합체에 포함된 표적 단위체의 해당 복합체에서의 위치를 결정할 수 있다.
복합체 생성부(212)는 복수의 복합체들을 복수의 표적 단위체들에 대한 희소 행렬로 표현할 수 있다. 복합체 생성부(212)는 희소 행렬에 있는 표적 단위체 셀의 위치를 동적으로 결정하여 디코딩 과정(복수의 분석 복합체로부터 표적 단위체에 대한 해석결과를 획득하는 과정)에서 필요한 표적 단위체 맵(224)을 생성할 수 있다.
복합체 생성부(212)는 설정된 보안 강도에 따라 복수의 복합체들의 개수 또는 복수의 복합체들 각각의 크기(복합체에 포함된 표적 단위체 또는 위장 단위체의 개수)를 결정할 수 있다. 복수의 복합체들의 개수는 복수의 복합체들로 구성된 행렬의 열의 개수, 복수의 복합체들 각각의 크기는 복수의 복합체들로 구성된 행렬의 행의 개수와 관련된다. 복합체 생성부(212)는 분석 요청자로부터 보안 강도 설정 요청을 수신하여 요청된 보안 강도를 만족시키는 복수의 복합체들을 생성할 수 있다.
복합체 생성부(212)는 각각이 표적 단위체 또는 위장 단위체 중 적어도 하나를 포함하는 복수의 복합체들을 생성하여 분석자에게 전송함으로써, 핵산염기서열에 대한 보안 등급 증가율에 비해서 분석자의 핵산염기서열 계산량 증가율은 현저히 낮은 보안 방법을 제공할 수 있다.
일반적인 암호화 기술은 정보 K를 암호화 함수(key) E로 변환한 E(K) 형태로 전송하고, 수신자는 복호화함수(key) D로 복호화 K = D(E(K)) 하는 과정이며, 이때 보안키에 해당하는 암호화, 복호화 함수는 각자의 책임으로 보관하는 것이며, 노출되면 암호화 기술은 무력화된다. 보안키는 일종의 정보이므로 그 크기가 보안의 핵심이다. 암호 해독을 위한 가장 단순한 공정은 랜덤 키 발생기(randome key generator, 모든 조합을 다 테스트 해보는 기법)이다. 이를 통해, 이론적으로 모든 보안키는 해독 가능하다. 예를 들어 4자리수 보안키는 1만번의 공격으로 해독할 수 있다. 따라서 현대보안 알고리즘은 32비트 보안키와 같은 매우 큰 보안키를 사용함으로써 현존하는 컴퓨터 기술로는 적당한 시간내(예를 들어, 1억년)에 모든 경우의 수를 도저히 테스트해볼 수 없도록 한다. 결국, 해결과제의 크기(Problem Size)를 현실적으로 계산 불가능한 수준으로 크게 만드는 것이다.
본 발명의 실시예는 개인의 강력한 식별정보인 유전자 정보를 대상으로 하므로, 데이터를 암호화해서 전송하는 문제에 관한 것이라기 보다는 보안키 자체를 전송하는 문제와 더 유사하다. 본 발명은 보안키를 위장키가 포함될 수 있는 적어도 2개 이상의 단위체로 분산(보안키를 구성하는 요소를 추출하거나, 분절하거나 또는 추출하여 조합하는 방식이 모두 가능)시킴으로써 보안키를 크게하는 효과를 갖는다. 예를 들어, 복합체 생성부(212)는 길이가 100인 핵산염기서열을 기초로 10개의 복합체를 생성하고 각 복합체는 1개의 표적 단위체(길이가 100인 핵산염기서열에서 유래)와 4개의 위장단위체를 포함함으로써 핵산염기서열 보안키를 크게 할 수 있다.
예를 들어, 생성된 복합체의 개수가 i, 각 복합체가 포함하는 단위체의 수가 j 인 경우 뒤섞여 버린 것을 모두 조합 추출하는 경우의 수는 j^i개 만큼 존재하게 되므로, 조합 가능한 j^i개의 키 중 원본 키 한 개를 찾아내는 문제로 변환되어 해결과제의 크기가 기하급수적으로 증가하게 된다. 예를 들어, i=10, j=10 일 때 문제크기는 10^10즉 100억 가지로 증가하는 것이다. 비교를 위해 현존하는 보안기술인 16 비트보안 (65536), 32비트(43억), 64비트(1.8x1019)를 고려하면, 64비트 보안과 유사수준을 얻기 위해서 i=19, j=10으로 충분하다. 즉 각 복합체가 포함하는 단위체의 개수가 2인 경우, 복합체수는 보안 비트수와 동일해지며, 이때 보안수준은 1/j^i이 된다.
이와 같이 해결과제의 크기가 커지는 경우 분석자 또는 분석기관인 수신측의 계산부하의 증가문제도 함께 해결하는 것이 중요하다. 분석자 또는 분석기관이 원래의 표적염기서열 한 개를 직접 받은경우에 수행하는 해석 또는 분석 연산량을 의미하는 계산복잡도는 해당 알고리즘에 따라 달라지며, 입력한 분석 요청자의 핵산염기서열의 길이에 무관한 상수(K), 해당 길이에 비례하는 선형알고리즘, 해당 길이의 제곱에 비례하는 이차함수(quadratic) 또는 해당 길이의 기하급수승 등에 해당할 수 있다.
즉, 본 발명은 알고리즘에 입력되는 분석 요청자의 핵산염기서열의 길이는 변하지 않는 반면, 분석자(120)의 계산부하는 각 복합체에 포함된 단위체 부가 배수만큼 증가하게 되므로 j^i인 보안등급 증가율에 비해서 분석자(120)의 계산량 증가율(i*j)은 현저히 낮은 매우 유리한 방식을 제공한다. 예를 들어, 핵산염기서열 보안 장치(110)는 분석 요청자의 핵산염기서열을 기초로 10개의 표적 단위체를 생성하고 표적 단위체 마다 4개의 위장 단위체를 부가하여 10개의 복합체를 생성한 경우, 5^10배 증가된 보안 등급을 가지는데 이때, 분석자(120)의 계산량은 5*10배 증가하고 이는 보안 등급 증가량에 비해 현저히 낮은 수치이다.
복합체 제공부(214)는 복합체 생성부(212)가 생성한 복수의 복합체들을 분석자(120)에게 제공한다. 복합체 제공부(214)는 생성된 복수의 복합체들 전부를 한 분석자에게 제공하거나 생성된 복수의 복합체들을 분할하여 복수의 직접 또는 간접 분석자들에게 제공할 수 있다. 예를 들어, 복합체 제공부(214)는 6개의 복합체를 생성하여 3개의 복합체는 분석자 A에게, 복합체는 분석자 B에게 제공할 수 있다.
분석 복합체 수신부(216)는 분석자(120)로부터 복수의 복합체들에 대한 해석결과를 나타내는 복수의 분석 복합체들(analysis composites)을 수신한다. 예를 들어, 복합체 수신부(216)는 제1 복합체에 대한 해석결과를 나타내는 제1 분석 복합체, 제2 복합체에 대한 해석결과를 나타내는 제2 분석 복합체 및 제3 복합체에 대한 해석결과를 나타내는 제3 분석 복합체를 수신할 수 있다.
분석 복합체 해석부(218)는 수신된 복수의 분석 복합체들을 통해 핵산염기서열에 대한 분석 결과를 획득한다. 분석 복합체 해석부(218)는 복수의 분석 복합체들에서 복수의 표적 단위체들의 해석결과에 해당하는 복수의 표적 해석 단위체들(target analysis elements)를 결정하고 이를 이용하여 분석 요청자의 핵산염기서열에 대한 최종 분석결과를 도출할 수 있다.
핵산염기서열 보안 제어부(219)는 핵산염기서열 보안 장치(110)의 전체적인 동작을 제어하고, 복합체 생성부(212), 복합체 제공부(214), 분석 복합체 수신부(216) 및 분석 복합체 해석부(218) 간의 데이터 흐름을 제어할 수 있다.
메모리(220)는 희소 행렬 메모리 영역(Sparse Matrix Memory Area, S.M.M.A) 및 표적 단위체 맵 영역(Target Element Map, T.E.M.A)을 포함한다.
희소 행렬 메모리 영역(S.M.M.A)는 복수의 복합체들(composites, 223) 및 복합체(223)에 포함된 복수의 표적 단위체들에 대한 희소 행렬(222)을 저장하는 공간에 해당한다.
복수의 복합체들(223) 각각은 표적 단위체(T.E) 또는 위장 단위체(D.E) 중 적어도 하나를 포함할 수 있다. 예를 들어, 복합체 1, 2, 3 및 5(223a, 223b, 223c 및 223e)는 하나의 표적 단위체와 네 개의 위장 단위체를, 복합체 4(223d)는 다섯 개의 위장 단위체를 포함한다.
도 2에서, 복수의 복합체들(223)은 4개의 표적 단위체와 21개의 위장 단위체를 포함하고, 복수의 복합체들(223)에 포함된 복수의 표적 단위체들과 위장 단위체들은 5행 5열 행렬을 구성할 수 있다(복수의 복합체들(223) 각각은 구성된 행렬의 열에 위치). 따라서, 복수의 복합체들(223)은 복수의 표적 단위체들에 대한 희소 행렬(222)로 표현될 수 있다.
도 2에서, 복수의 표적 단위체들에 대한 희소 행렬(222)은 2행 1열에 복합체 1(223a)의 표적 단위체, 4행 2열에 복합체 2(223b)의 표적 단위체, 1행 3열에 복합체 3(223c)의 표적 단위체 및 5행 5열에 복합체 5(223e)의 표적 단위체를 행렬의 셀로 포함한다.
표적 단위체 맵 영역(T.E.M.A)은 희소 행렬 메모리 영역(S.M.M.A)에 저장된 희소 행렬(222)에 대한 표적 단위체 맵(224)을 저장하는 공간에 해당한다.
표적 단위체 맵(224)는 희소 행렬(222)에 포함된 복수의 표적 단위체들의 위치에 대한 맵에 해당한다. 예를 들어, 표적 단위체 맵(224)은 희소 행렬 메모리 영역(S.M.M.A)에 저장된 희소 행렬(222)의 열의 개수가 5개에 해당하면 제1 열에 있는 표적 단위체부터 제5 열에 있는 표적 단위체까지의 위치를 포함한다. 표적 단위체 맵(224)은 희소 행렬(222)의 해당 열이 표적 단위체를 포함하고 있지 않는 경우 해당 열에 대한 표적 단위체의 위치를 마이너스 값(예를 들어 -1)으로 저장하여 표적 단위체의 유무를 구분할 수 있다.
네트워크 인터페이스(230)는 네트워크를 통해 분석자(120)와 연결하기 위한 환경을 포함하고, 예를 들어, LAN(Local Area Network) 통신을 위한 어댑터를 포함할 수 있다.
사용자 입력 장치(240)는 사용자 입력을 수신하기 위한 환경을 포함하고, 예를 들어, 마우스, 트랙볼, 터치 패드, 그래픽 태블릿, 스캐너, 터치 스크린, 키보드 또는 포인팅 장치와 같은 어댑터를 포함할 수 있다.
사용자 출력 장치(250)는 사용자에게 특정 정보(예를 들어, 분석 요청자의 핵산염기서열 분석결과)를 출력하기 위한 환경을 포함하고, 예를 들어, 모니터 또는 터치스크린과 같은 어댑터를 포함할 수 있다. 일 실시예에서, 사용자 입력 장치(240)와 사용자 출력 장치(250)는 원격 접속을 통해 접속될 수 있다.
저장장치(260)는 SSD(Solid State Disk) 또는 HDD(Hard Disk Drive)와 같은 비휘발성 메모리로 구현될 수 있고, 핵산염기서열 보안 장치(110)에 필요한 데이터를 저장하는데 사용된다.
도 3은본 발명의 일 실시예에 따른 복수의 복합체들 생성 및 복수의 분석 복합체들 해석프로시저를 설명하는 도면이다.
도 3에서, 각 염기는 유전좌위(유전좌위는 염기의 고유한 주소/좌표로 염색체상의 위치, 일종의 좌표값, 12번 염색체의 1234501번째 위치 또는 rsID 등의 값으로 표기될 수 있다)를 가지고 있으며, 이처럼 유전좌위 + 염기 + 기타 정보를 담는 파일 포맷은 표준화되어 사용되는 많은 파일이 있다. 예를 들면 VCF (Variant Call Format)은 전체 서열이 아니라, 추출된 개인별 변이 부위만 좌표와 검출된 유전형 + 기타의 정보로 표현하는 파일 포맷이다. 일 실시예에서, 분석 요청자의 핵산염기서열 및 상기 표적 단위체, 위장 단위체, 복합체 또는 분석 복합체에 포함된 핵산염기서열의 유전좌위 + 염기 + 기타정보를 담는 표준화된 포맷은 VCF(Variant CallFormat), BCF(Binary version of VCF), GFF(Gene-Finding Format, GenericFeature Format, 현재버전은 4.1), GTF(Gene Transfer Format), GVF(GenomeVariation Format), SAM/BAM(Sequence Alignment Map/ Binary version of SAM),QUAL, SCARF, QSEQ, IG, Maq, SOAP, bcf, pileup, mpileup, CASAVA, MaCH, GLFv2,GPFv2, axt, BED, BED15, BED detail, BEDPE, bedGraph, bigBed, bigWig, Chain,GenePred table, HAL, HDF5, MAF, Net, Personal Genome SNP, PSL, WIG (Wiggleformat), .2bit, .nib, CSFASTQ, CSFASTA, FASTA, FASTQ 형식 또는 그 확장형식들을 포함할 수 있다. 이때 표적 핵산염기서열은 SNP(single nucleotide polymorphism) 등 선택된 특정 염기(마커)만 선별해서 보낼 수도 있고 연속된 전체서열을 보낼 수도 있다. 유전자서열은 염기가 일정순서로 나열된 서열 자체이고, 또한 본 발명은 전송된 표적 핵산염기서열의 보안에 관한 것이기 때문에, 그 전송서열이 원 서열이든, 일부만 선택된 서열이든 모두 가능한 것이다. 일 실시예에서, 표적 핵산염기서열은 특정 유전자 또는 비유전자 좌위의 SNP(single nucleotide polymorphism), 염기의 치환, 부가 또는 결실을 포함하는 이수체 돌연변이 또는 STRP (short tandem repeat polymorphism) 또는 VNTR (various number oftandem repeat)를 포함하는 다수체 (polyalleic) 변이에서 결정된 염기일수 있다.
복합체 생성부(212)는 분석 요청자의 핵산염기서열로부터 적어도 하나의 염기와 유전좌위를 추출하여 표적 단위체를 생성한다. 복합체 생성부(212)는 복수의 복합체들을 생성하기 위하여 분석 요청자의 핵산염기서열에서 적어도 하나의 염기와 유전좌위를 추출하는 방식으로 GAAT의 핵산염기서열을 가지는 표적 단위체 T.E1, ACGCA의 핵산염기서열을 가지는 표적 단위체 T.E2, TCCTGAT의 핵산염기서열을 가지는 표적 단위체 T.E3, GACAC의 핵산염기서열을 가지는 표적 단위체 T.E4, CCAGCG의 핵산염기서열을 가지는 표적 단위체T.E5를 생성할 수 있다.
다음으로, 복합체 생성부(212)는 복수의 복합체들을 구성하는 위장 단위체를 생성한다. 일 실시예에서, 복합체 생성부(212)는 생성된 표적 단위체와의 유사도를 기초로 해당 복합체에 있는 적어도 하나의 위장 단위체를 생성할 수 있다. 다른 일 실시예에서, 복합체 생성부(212)는 생성된 표적 단위체와의 유전자 거리 또는 진화적 거리가 특정 거리 이하인 적어도 하나의 위장 단위체를 생성할 수 있다.
최종적으로, 복합체 생성부(212)는 표적 단위체 T.E1을 포함하는 복합체 C1, 표적 단위체 T.E2를 포함하는 복합체 C2, 표적 단위체 T.E 3를 포함하는 복합체 C3, 표적 단위체 T.E4를 포함하는 복합체 C4, 표적 단위체, 표적 단위체 T.E 5를 포함하는 복합체 C5 및 표적 단위체를 포함하지 않고 복수의 위장 단위체들로만 구성된 복합체 C6를 생성할 수 있다. 복합체 생성부(212)는 복수의 복합체들(C1~C6)을 생성하면서 복합체에 포함된 표적 단위체의 위치를 결정 할 수 있다. 도 3에서, 복합체 생성부(212)는 표적 단위체 T.E1을 복합체 C1의 3번 위치에, 표적 단위체 T.E2을 복합체 C2의 1번 위치에, 표적 단위체 T.E3를 복합체 C3의 4번 위치에, 표적 단위체 T.E4를 복합체 C4의 1번 위치에 마지막으로 표적 단위체 T.E5를 복합체 C5의 3번 위치에 배치할 수 있다.
생성된 복수의 복합체들(C1~C6)는 분석자에게 제공되고 분석자(120)는 제공받은 복수의 복합체들(C1~C6)를 분석하여 복수의 복합체들(C1~C6)에 대한 해석결과를 나타내는 복수의 분석 복합체들(analysis composite 1~analysis composite 6, A.C1~A.C6)을 분석 복합체 수신부(216)에 제공한다. 즉, 분석자(120)는 분절된 핵산염기서열들에 대한 제공 가능한 분석결과값을 생성한다. 예를 들어, 복합체 생성부(212)가 10개의 복합체를 생성하고, 각 복합체는 1개의 표적 단위체와 4개의 위장단위체를 포함하고 있는 경우, 분석자(120)는 50개의 분석결과값을 생성하고 생성된 분석결과값을 포함하는 복수의 분석 복합체들을 핵산염기서열 보안 장치(110)로 전송하게 된다. 핵산염기서열 보안 장치(110)는 50개의 분석결과값을 수신하여 표적 단위체 염기서열 또는 표적 단위체 생성규칙(예를 들어, 표적 단위체 맵(224)과 같은 표적 단위체 위치 정보)을 참조하여 각 분석 복합체에서 표적 단위체의 해석 또는 분석의 결과값 10개(10개의 표적 분석 단위체)를 추출하고, 이들을 조합하여 분석 요청자의 핵산염기서열에 대한 해석 또는 분석의 최종결과물과 동일한 결과값을 획득할 수 있다. 이때, 분석 복합체에서 추출한 표적 단위체의 결과값 10개를 조합하기 위한 함수는, 함수를 미리 약속해두거나, 분석기관이 생성/전송하거나 클라이언트(예를들어, 분석 요청자)가 요청할 수 있다. 여기에서, 함수는 10개의 수를 곱하거나, 합하거나 또는 평균치를 계산하는 등의 다양한 계산 방법을 포함한다
도 3에서, 복수의 분석 복합체들(A.C1~A.C6)은 복수의 복합체들(C1~C6)에 포함된 복수의 표적 단위체들(T.E1~T.E5)에 각각 대응하는 복수의 표적 해석 단위체(R13, R21, R34, R41 및 R53)을 포함한다. 핵산염기서열 보안 장치(110)는 분석을 요청한 표적 단위체와 위장 단위체의 핵산염기서열 정보를 보유하고 있으므로, 각 복합체별 해당 표적단위체의 핵산염기서열 정보(도3에서, T.E1은 GAAT, T.E2는 ACGCA, T.E3는 TCCTGAT, T.E4 GACAC, T.E5는 CCAGCG에 해당)와 일치하는 서열과 논리적으로 연결된 분석결과 (여기서는 R13, R21,R34, R41, R53)를 용이하게 추출할 수 있다. 이때 분석자(120)는 복수의 분석 복합체들에 포함된 각 단위체에 표적 단위체 및 위장 단위체의 핵산염기서열정보를 반드시 다 포함시켜 전송할 필요는 없이 부호화 등을 통해 정보량을 줄여(논리적으로는 여전히 식별가능하게 줄이거나 부호화) 전송하는 것이 가능하다.
일 실시예에서, 복수의 표적 해석 단위체(R13, R21, R34, R41 및 R53)는 복수의 표적 단위체들을 구성하는 핵산염기서열과 해당 핵산염기서열에 대한 분석결과값을 연결하는 염기서열-결과값 자료쌍의 집합을 의미할 수 있다. 다른 일 실시예에서, 복수의 표적 해석 단위체(R13, R21, R34, R41 및 R53)는 복수의 표적 단위체들을 구성하는 핵산염기서열 정보가 표기되지 않거나 생략되고 해당 핵산염기서열의 식별을 위해 부여된 식별자로 대체된 식별자-결과값 자료 쌍의 집합을 의미할 수 있다.
분석 복합체 해석부(218)는 표적 단위체 맵(224)을 기초로 분석자로부터 수신한 복수의 분석 복합체들(A.C1~A.C6)에서 복수의 표적 단위체들(T.E1~T.E5)에 대응하는 복수의 표적 해석 단위체(R13, R21, R34, R41 및 R53)을 결정할 수 있다.
분석 복합체 해석부(218)는 결정된 복수의 표적 분석 단위체들((R13, R21, R34, R41 및 R53))을 특정 함수 f(Rk)를 이용하여 연산하고 분석 요청자의 핵산염기서열에 대한 최종 분석 결과를 도출할 수 있다. 예를 들어 분석 복합체 해석부(218)는 결정된 복수의 표적 분석 단위체들((R13, R21, R34, R41 및 R53))을 조합하여 분석 요청자의 사후오즈(Posterior Odds)를 계산할 수 있고 계산된 분석 요청자의 사후오즈는 분석 요청자의 핵산염기서열에 대한 최종 분석 결과를 제공할 수 있다. 이하, 분석 복합체 해석부(218)에서 수행되는 분석 요청자의 사후오즈 계산과정을 상세히 설명한다.
분석 복합체 해석부(218)는 베이즈 이론(Bayes Theorem)에 따라 분석 요청자의 핵산염기서열의 사후오즈(Posterior Odds, 해당 핵산염기서열을 가진 개인의 질병위험률)를 계산할 수 있다. 사후오즈(Posteror Odds)는 베이즈 이론(Bayes Theorem)에 따르면 하기의 수학식 1과 같이 사전오즈(Prior Odds)와 우도비(Likelihood Ratio)의 곱으로 표현될 수 있다.
[수학식 1]
Posteror Odds = Prior Odds x Likelihood Ratio
이때, 사전오즈는 사전지식이 없는 경우 중립 값인 1로 치환될 수 있다. 따라서, 사후오즈는사전오즈 값 1에 우도비를 계속 곱해 나가는 연쇄계산(Chain calculation)을 통해 얻어질 수 있다. 여기에서, 유전형에 따른 우도비는 하기의 수학식 2와 같이 계산될 수 있다.
[수학식 2]
우도비 = (증례군의유전형확률)/(대조군유전형확률)
예를들어, n 개의 유전좌위 a1, a2,... an에서획득한 염기 G(a1), G(a2),..., G(an)으로 구성된 핵산염기서열 [a1, a2,... an]의 각 위치에 대한 우도비 LR(G(a1)), LR(G(a2)),..., LR(G(an))가 주어진 경우, 해당 핵산염기서열을 가진 개인의 질병위험률에 해당하는 사후오즈는 하기의 수학식 3과 같이 계산될 수 있다.
[수학식 3]
분석 복합체 해석부(218)는 수학식 3을 이용하여 분석 요청자의 핵산염기서열의 사후오즈를 계산할 수 있다. 예를 들어, 핵산염기서열 보안 장치(110)가 분석 요청자의 핵산염기서열 [a1, a2,... an]로부터 k개의표적단위체를 결정한 경우, 곱셈의 교환법칙(a x b = b x a)과, 곱셈의 결합법칙((a xb) x c = a x (b x c))에의해, 분석 복합체 해석부(218)는 하기의 수학식 4와 같이k개의 표적 단위체의 사후오즈(Rk)를 곱함(프로덕트 연산수행)으로써 분석 요청자의 핵산염기서열에 대한 최종 사후오즈를 계산할 수 있다.
[수학식 4]
본 발명에서, 수학식 4의 Rk는 복수의 분석 복합체들에 포함된 표적 단위체에 대한 분석결과 염기서열의 해석결과 값에 해당할 수 있다.
수학식 4는 변수가k 및 Rk 뿐이므로, 분석 요청자와 분석자(120) 사이에서 사전에 미리 정해질 수 있고, 분석자(120)에 의해 복수의 분석 복합체들과 함께 전송될 수도 있으며, 분석 요청자에 의해 적어도 하나의 표적 단위체 각각의 해석결과값(R1~Rk)과 함께 연산에 필요한 해당 수식이 요청될 수도 있다.
또한, 수학식 4는 표적 단위체의 개수나, 각 표적 단위체의 핵산염기서열의 길이에 무관하게 성립한다. 그러므로 핵산염기서열 보안 장치(110)는 표적 단위체의 개수를 n보다 작거나 같은 모든 수로 결정하는 것이 가능하다. 그러나 분석 요청자의 핵산염기서열의 위치를 하나씩 분절하는 것은 분석을 수행하는 분석자(120)의 관점에서 핵심적인 지식재산인 각 위치에 대한 우도 LR(G(a1)), LR(G(a2)),..., LR(G(an))를 모두 노출하는 것이 되므로 적절하지 않다. 핵산염기서열 보안 장치(110)는 분석자(120)의 데이터 처리에 사용되는 알고리즘등과 같은 지식재산보호를 위해서 최소 2 이상의 길이를 갖는 표적 단위체를 생성할 필요가 있다. 표적 단위체의 길이는 보안강도 설정의 목표에 따라 변동될 수 있다. 보안강도는 표적 단위체 길이에 대해 지수적으로 증가하므로 손쉽게 고도의 보안등급을 설정할 수 있다. 반면 표적 단위체의 길이가 너무 길어져 분석 요청자의 핵산염기서열의 길이에 근접하게 되면 상대적으로 송신자의 표적 핵산염기서열 노출확률이 증가하게 되어, 이를 방어하기 위해서는 복합체에 포함된 위장 단위체의 수를 늘려야 한다. 그러나 위장 단위체 수와 보안강도는 선형적으로 밖에는 증가하지 않으므로, 그 수를 매우 크게 증가시켜야 하고, 이는 분석자(120)의 분석부하를 기하급수적으로 증가시켜 바람직한 방법으로 보기 어렵다. 그러므로 본 발명은 표적 핵산염기서열을 보호하기 위한 복합체의 개수, 복수의 복합체들에 포함된 표적 단위체의 개수 또는 각 복합체에 포함된 위장 단위체의 개수를 조절하여 분석 요청자와 분석자(120) 각각의 프라이버시와 지식자산 보안등급 요구를 충족할 수 있다. 즉, 본 발명은 표적 단위체의 개수, 위장 단위체의 개수 및 복수의 복합체들의 개수를 조절하여 분석 요청자의 핵산염기서열 보안 및 분석자(120)의 지식자산 보안 요구를 동시에 충족시킬 수 있다.
예를 들어, 도 3은 복합체 생성부(212)가 6개의 복합체를 생성한 경우로, 복합체 C1부터 복합체 C5까지는 하나의 표적 단위체와 약 4개의 위장 단위체를 포함하고, C6는 4개의 위장 단위체만을 포함하고 있다. 이때 표적 단위체를 포함하지 않는 복합체 C6를 포함하는 6개의 복합체의 표적 핵산염기서열의 정보보호 등급은 4^6= 2^12= 1/4096으로 12비트 보안수준에 해당한다. 보안수준 목표를 매우 강력한 32비트로 정한 경우, 2^32= 4^16= 16^8= 256^4=65546^2이므로 각 2개의 단위체를 포함하는 32개의 복합체, 각 4개의 단위체를 포함하는 16개의 복합체, 각 16개의 단위체를 포함하는 8개의 복합체, 각 256개의 단위체를 포함하는 4개의 복합체 또는 각 65546개의 단위체를 포함하는 2개의 복합체를 생성하여 32비트의 보안수준이 만족될 수 있다.
그러므로 본 발명은 먼저 정보보호등급을 정한 후, 분석자(120)의 계산부하를 고려한 복합체의 개수 및 복합체에 포함되는 단위체의 개수를 결정할 수 있다.
도 4는 본 발명의 다른 일 실시예에 따른 복수의 복합체들 생성 및 복수의 분석 복합체들 해석프로시저를 설명하는 도면이다.
도 4에서, 복합체 생성부(212)는 분석 요청자의 핵산염기서열을 부분 염기서열로 분절하여 상기 표적 단위체를 생성한다. 복합체 생성부(212)는 복수의 복합체들을 생성하기 위하여 분석 요청자의 핵산염기서열을 부분 염기서열로 분절하는 방식으로 GGAA의 핵산염기서열을 가지는 표적 단위체 T.E1, TCAAC의 핵산염기서열을 가지는 표적 단위체 T.E2, CGGCGGA의 핵산염기서열을 가지는 표적 단위체 T.E3, CTGAT의 핵산염기서열을 가지는 표적 단위체 T.E4, TACACCC의 핵산염기서열을 가지는 표적 단위체 T.E5를 생성할 수 있다. 이후에 핵산염기서열 보안 장치(110)에서 수행되는 프로시저는 도 3과 같다.
도 5a는 본 발명의 일 실시예에 따른 표적 단위체에 대한 희소 행렬을 시각화한 도면이고, 도 5b는 본 발명의 다른 일 실시예에 따른 표적 단위체에 대한 희소 행렬을 시각화한 도면이다.
도 5a에서, 복합체 생성부(212)는 희소 행렬에 있는 표적 단위체 셀을 표적 단위체에 대한 적어도 하나의 염기와 유전좌위를 포함하는 염기-좌위 집합으로 정의할 수 있다. 예를 들어, 희소 행렬에 있는 표적 단위체 셀은 해당 표적 단위체가 유전좌위가 12인 염기 A, 유전좌위가 15인 염기 G 및 유전좌위가 332인 염기 G로 구성된 핵산염기서열에 해당하면 {(A,12),(G,15),(G,332)}의 집합으로 정의될 수 있다.
도 5b에서, 복합체 생성부(212)는 희소 행렬에 있는 표적 단위체 셀을 유전좌위와 연관된 적어도 하나의 염기 집합으로 정의할 수 있다. 예를 들어, 희소 행렬에 있는 표적 단위체 셀은 해당 표적 단위체가 해당 표적 단위체가 유전좌위가 12인 염기 A, 유전좌위가 15인 염기 G 및 유전좌위가 332인 염기 G로 구성된 핵산염기서열에 해당하면 {A,G,G}의 집합으로 정의될 수 있다. 이때, 해당 염기들의 유전좌위 12,15 및 332는 행렬에서 공통으로 표현될 수 있다.
도 6은 도 2에 있는 핵산염기서열 보안 장치에 의하여 수행되는 핵산염기서열 보안 방법을 설명하는 흐름도이다.
복합체 생성부(212)는 각각이 분석 요청자의 핵산염기서열로부터 파생된 표적 단위체 또는 표적 단위체와 동일하거나 또는 다른 위장 단위체 중 적어도 하나를 포함하는 복수의 복합체들을 생성한다(S601). 이때, 복합체 생성부(212)는 복수의 복합체들을 복수의 표적 단위체들에 대한 희소 행렬로 표현할 수 있다(S602).
복합체 생성부(212)는 표적 단위체가 해당 복합체에 포함되면 해당 표적 단위체의 위치를 결정할 수 있다(S603). 복합체 생성부(212)는 복수의 복합체들을 복수의 표적 단위체들에 대한 희소 행렬로 표현하면서, 희소 행렬에 있는 표적 단위체 셀의 위치를 동적으로 결정하여 디코딩 과정에서 필요한 표적 단위체 맵을 생성할 수 있다(S604)
복합체 제공부(214)는 생성된 복수의 복합체들을 분석자에게 제공하고 해석 복합체 수신부(216)은 분석자로부터 복수의 복합체들에 대한 해석결과를 나타내는 복수의 분석 복합체들을 수신할 수 있다(S605 및 S606).
해석 복합체 해석부(218)는 표적 단위체 맵에 저장된 표적 단위체의 위치 정보를 기초로 복수의 표적 단위체들에 대한 해석결과를 나타내는 복수의 표적 분석 단위체들을 결정할 수 있다(S607). 해석 복합체 해석부(218)는 결정된 복수의 표적 분석 단위체들을 조합하여 분석 요청자의 사후오즈를 계산하고, 계산된 사후오즈를통해 분석 요청자의 핵산염기서열의 최종 분석결과를 획득할 수 있다(S608 및 S609).
상기에서는 본 출원의 바람직한 실시예를 참조하여 설명하였지만, 해당 기술 분야의 숙련된 당업자는 하기의 특허 청구의 범위에 기재된 본 발명의 사상 및 영역으로부터 벗어나지 않는 범위 내에서 본 출원을 다양하게 수정 및 변경시킬 수 있음을 이해할 수 있을 것이다.
본 발명은 핵산염기서열 보안 기술에 관한 것으로, 보다 상세하게는 분석 요청자의 핵산염기서열을 분석자에게 노출하지 않고 분석 요청자의 핵산염기서열을 분석할 수 있는 핵산염기서열 보안 방법, 장치 및 이를 저장한 기록매체에 관한 것이다.
Claims (21)
- (a) 각각이 분석 요청자의 핵산염기서열로부터 파생된 표적 단위체(target element) 또는 상기 표적 단위체와 동일하거나 또는 다른 위장 단위체(disguising element) 중 적어도 하나를 포함하는 복수의 복합체들(composites)을 생성하는 단계; 및(b) 상기 생성된 복수의 복합체들을 분석자에게 제공하는 단계를 포함하는 핵산염기서열 보안 방법.
- 제1항에 있어서, 상기 (a) 단계는상기 복수의 복합체들을 복수의 표적 단위체들에 대한 희소 행렬로 표현하는 단계를 포함하는 것을 특징으로 하는 핵산염기서열 보안 방법.
- 제2항에 있어서, 상기 (a) 단계는해당 표적 단위체가 해당 복합체에 포함되면 상기 해당 표적 단위체의 위치를 결정하는 단계를 더 포함하는 것을 특징으로 하는 핵산염기서열 보안 방법.
- 제2항에 있어서, 상기 (a) 단계는상기 희소 행렬에 있는 표적 단위체 셀을 상기 표적 단위체에 대한 적어도 하나의 염기와 유전좌위를 포함하는 염기-좌위 집합으로 정의하는 단계를 더 포함하는 것을 특징으로 하는 핵산염기서열 보안 방법.
- 제2항에 있어서, 상기 (a) 단계는상기 희소 행렬에 있는 표적 단위체 셀을 유전좌위와 연관된 적어도 하나의 염기 집합으로 정의하는 단계를 더 포함하는 것을 특징으로 하는 핵산염기서열 보안 방법.
- 제2항에 있어서, 상기 (a) 단계는상기 희소 행렬에 있는 표적 단위체 셀의 위치를 동적으로 결정하여 디코딩 과정에서 필요한 표적 단위체 맵을 생성하는 단계를 더 포함하는 것을 특징으로 하는 핵산염기서열 보안 방법.
- 제1항에 있어서, 상기 (a) 단계는상기 핵산염기서열로부터 적어도 하나의 염기와 유전좌위를 추출하여 상기 표적 단위체를 생성하는 단계를 포함하는 것을 특징으로 하는 핵산염기서열 보안 방법.
- 제1항에 있어서, 상기 (a) 단계는상기 핵산염기서열을 부분 염기서열로 분절하여 상기 표적 단위체를 생성하는 단계를 포함하는 것을 특징으로 하는 핵산염기서열 보안 방법.
- 제1항에 있어서, 상기 (a) 단계는상기 표적 단위체와의 유사도를 기초로 해당 복합체에 있는 적어도 하나의 위장 단위체를 생성하는 단계를 포함하는 것을 특징으로 하는 핵산염기서열 보안 방법.
- 제9항에 있어서, 상기 (a) 단계는상기 표적 단위체와의 유전자 거리 또는 진화적 거리가 특정 거리 이하인 적어도 하나의 위장 단위체를 생성하는 단계를 더 포함하는 것을 특징으로 하는 핵산염기서열 보안 방법.
- 제1항에 있어서, 상기 (a) 단계는상기 분석 요청자에 의하여 설정된 보안 강도에 따라 상기 복수의 복합체들의 개수 또는 상기 복수의 복합체들 각각의 크기를 결정하는 단계를 포함하는 것을 특징으로 하는 핵산염기서열 보안 방법.
- 제1항에 있어서,(c) 상기 분석자로부터 상기 복수의 복합체들에 대한 해석결과를 나타내는 복수의 분석 복합체들(analysis composites)을 수신하여 상기 핵산염기서열에 대한 분석 결과를 획득하는 단계를 더 포함하는 것을 특징으로 하는 핵산염기서열 보안 방법.
- 제12항에 있어서, 상기 (c) 단계는표적 단위체 맵을 기초로 복수의 표적 단위체들에 대한 해석결과를 나타내는 복수의 표적 분석 단위체들(target analysis elements)을 결정하는 단계를 포함하는 것을 특징으로 하는 핵산염기서열 보안 방법.
- 제13항에 있어서, 상기 (c) 단계는상기 결정된 복수의 표적 분석 단위체들을 조합하여 상기 분석 요청자의 사후오즈(Posterior Odds)를 계산하는 단계를 더 포함하는 것을 특징으로 하는 핵산염기서열 보안 방법.
- 제1항에 있어서, 상기 (b) 단계는상기 생성된 복수의 복합체들을 분할하여 복수의 직접 또는 간접 분석자들에게 제공하는 단계를 포함하는 것을 특징으로 하는 핵산염기서열 보안 방법.
- 각각이 분석 요청자의 핵산염기서열로부터 파생된 표적 단위체(target element) 또는 상기 표적 단위체와 동일하거나 또는 다른 위장 단위체(disguising element) 중 적어도 하나를 포함하는 복수의 복합체들(composites)을 생성하는 복합체 생성부; 및상기 생성된 복수의 복합체들을 분석자에게 제공하는 복합체 제공부를 포함하는 핵산염기서열 보안 장치.
- 제16항에 있어서,상기 분석자로부터 상기 복수의 복합체들에 대한 해석결과를 나타내는 복수의 분석 복합체들을 수신하여 상기 핵산염기서열에 대한 분석 결과를 획득하는 분석 복합체 해석부를 더 포함하는 것을 특징으로 하는 핵산염기서열 보안 장치.
- 제17항에 있어서, 상기 분석 복합체 해석부는표적 단위체 맵을 기초로 복수의 표적 단위체들에 대한 해석결과를 나타내는 복수의 표적 분석 단위체들(target analysis elements)을 결정하는 것을 특징으로 하는 핵산염기서열 보안 장치.
- 제18항에 있어서, 상기 분석 복합체 해석부는상기 결정된 복수의 표적 분석 단위체들을 조합하여 상기 분석 요청자의 사후오즈(Posterior Odds)를 계산하는 것을 특징으로 하는 핵산염기서열 보안 장치.
- 각각이 분석 요청자의 핵산염기서열로부터 파생된 표적 단위체(target element) 또는 상기 표적 단위체와 동일하거나 또는 다른 위장 단위체(disguising element) 중 적어도 하나를 포함하는 복수의 복합체들(composites)을 생성하는 기능; 및상기 생성된 복수의 복합체들을 분석자에게 제공하는 기능을 포함하는 핵산염기서열 보안 방법에 관한 컴퓨터 프로그램을 기록한 기록매체.
- 제20항에 있어서,상기 분석자로부터 상기 복수의 복합체들에 대한 해석결과를 나타내는 복수의 분석 복합체들을 수신하여 상기 핵산염기서열에 대한 분석 결과를 획득하는 기능을 더 포함하는 것을 특징으로 하는 핵산염기서열 보안 방법에 관한 컴퓨터 프로그램을 기록한 기록매체.
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US15/321,135 US10896743B2 (en) | 2014-06-24 | 2015-06-24 | Secure communication of nucleic acid sequence information through a network |
| US17/111,444 US20210104298A1 (en) | 2014-06-24 | 2020-12-03 | Secure communication of nucleic acid sequence information through a network |
Applications Claiming Priority (4)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| KR10-2014-0077019 | 2014-06-24 | ||
| KR20140077019 | 2014-06-24 | ||
| KR10-2015-0088901 | 2015-06-23 | ||
| KR1020150088901A KR101788673B1 (ko) | 2014-06-24 | 2015-06-23 | 핵산염기서열 보안 방법, 장치 및 이를 저장한 기록매체 |
Related Child Applications (2)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US15/321,135 A-371-Of-International US10896743B2 (en) | 2014-06-24 | 2015-06-24 | Secure communication of nucleic acid sequence information through a network |
| US17/111,444 Continuation US20210104298A1 (en) | 2014-06-24 | 2020-12-03 | Secure communication of nucleic acid sequence information through a network |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2015199440A1 true WO2015199440A1 (ko) | 2015-12-30 |
Family
ID=54938443
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/KR2015/006427 Ceased WO2015199440A1 (ko) | 2014-06-24 | 2015-06-24 | 핵산염기서열 보안 방법, 장치 및 이를 저장한 기록매체 |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2015199440A1 (ko) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2018094115A1 (en) * | 2016-11-16 | 2018-05-24 | Catalog Technologies, Inc. | Systems for nucleic acid-based data storage |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2003242154A (ja) * | 2002-02-18 | 2003-08-29 | Celestar Lexico-Sciences Inc | 遺伝子発現情報管理装置、遺伝子発現情報管理方法、プログラム、および、記録媒体 |
| KR20130075559A (ko) * | 2011-12-27 | 2013-07-05 | 주식회사 마크로젠 | 유전자 정보 관리 장치 및 방법 |
| KR20130122816A (ko) * | 2012-05-01 | 2013-11-11 | 강원대학교산학협력단 | 유전자 염기서열 압축장치 및 압축방법 |
| WO2013178801A2 (en) * | 2012-06-01 | 2013-12-05 | European Molecular Biology Laboratory | High-capacity storage of digital information in dna |
-
2015
- 2015-06-24 WO PCT/KR2015/006427 patent/WO2015199440A1/ko not_active Ceased
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2003242154A (ja) * | 2002-02-18 | 2003-08-29 | Celestar Lexico-Sciences Inc | 遺伝子発現情報管理装置、遺伝子発現情報管理方法、プログラム、および、記録媒体 |
| KR20130075559A (ko) * | 2011-12-27 | 2013-07-05 | 주식회사 마크로젠 | 유전자 정보 관리 장치 및 방법 |
| KR20130122816A (ko) * | 2012-05-01 | 2013-11-11 | 강원대학교산학협력단 | 유전자 염기서열 압축장치 및 압축방법 |
| WO2013178801A2 (en) * | 2012-06-01 | 2013-12-05 | European Molecular Biology Laboratory | High-capacity storage of digital information in dna |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2018094115A1 (en) * | 2016-11-16 | 2018-05-24 | Catalog Technologies, Inc. | Systems for nucleic acid-based data storage |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US20250226056A1 (en) | Variant classifier based on deep neural networks | |
| US11702708B2 (en) | Systems and methods for analyzing viral nucleic acids | |
| KR101788673B1 (ko) | 핵산염기서열 보안 방법, 장치 및 이를 저장한 기록매체 | |
| US20210057045A1 (en) | Determining the Clinical Significance of Variant Sequences | |
| Tang et al. | Infection control in the new age of genomic epidemiology | |
| Song et al. | Deep-level phylogeny of Cicadomorpha inferred from mitochondrial genomes sequenced by NGS | |
| Iqbal et al. | De novo assembly and genotyping of variants using colored de Bruijn graphs | |
| Sun et al. | ESPRIT: estimating species richness using large collections of 16S rRNA pyrosequences | |
| Page et al. | BamBam: genome sequence analysis tools for biologists | |
| US9935765B2 (en) | Device, system and method for securing and comparing genomic data | |
| US11004544B2 (en) | Method of providing biological data, method of encrypting biological data, and method of processing biological data | |
| US20160019339A1 (en) | Bioinformatics tools, systems and methods for sequence assembly | |
| KR102828110B1 (ko) | 개선된 컴퓨팅 장치 | |
| Baptista et al. | Is reliance on an inaccurate genome sequence sabotaging your experiments? | |
| Peralta et al. | Sniploid: A utility to exploit high‐throughput SNP data derived from RNA‐seq in allopolyploid species | |
| CA3020669A1 (en) | Systems and methods for biological data management | |
| Jani et al. | IslandCafe: compositional anomaly and feature enrichment assessment for delineation of genomic islands | |
| Chaguza et al. | RCandy: an R package for visualizing homologous recombinations in bacterial genomes | |
| Goussarov et al. | PaSiT: a novel approach based on short-oligonucleotide frequencies for efficient bacterial identification and typing | |
| Bharti et al. | MTCID: a database of genetic polymorphisms in clinical isolates of Mycobacterium tuberculosis | |
| JP2008529538A (ja) | 相補性デュプリコンの増幅を含む遺伝子分析方法 | |
| Tassios et al. | Bacterial next generation sequencing (NGS) made easy | |
| Zhang et al. | Reading the underlying information from massive metagenomic sequencing data | |
| Sánchez-Busó et al. | pyngoST: fast, simultaneous and accurate multiple sequence typing of Neisseria gonorrhoeae genome collections | |
| Shah et al. | SNP-VISTA: an interactive SNP visualization tool |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 15811382 Country of ref document: EP Kind code of ref document: A1 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 15321135 Country of ref document: US |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 15811382 Country of ref document: EP Kind code of ref document: A1 |



