WO2023030233A1 - 一种拷贝数变异的检测方法及其应用 - Google Patents
一种拷贝数变异的检测方法及其应用 Download PDFInfo
- Publication number
- WO2023030233A1 WO2023030233A1 PCT/CN2022/115447 CN2022115447W WO2023030233A1 WO 2023030233 A1 WO2023030233 A1 WO 2023030233A1 CN 2022115447 W CN2022115447 W CN 2022115447W WO 2023030233 A1 WO2023030233 A1 WO 2023030233A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- sample
- copy number
- tested
- target
- sequencing data
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B20/00—ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
- G16B20/30—Detection of binding sites or motifs
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q1/00—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
- C12Q1/68—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
- C12Q1/6869—Methods for sequencing
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B20/00—ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
- G16B20/10—Ploidy or copy number detection
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B20/00—ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
- G16B20/20—Allele or variant detection, e.g. single nucleotide polymorphism [SNP] detection
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B30/00—ICT specially adapted for sequence analysis involving nucleotides or amino acids
- G16B30/10—Sequence alignment; Homology search
Definitions
- This application relates to the field of biological information, in particular to a method for detecting copy number variation and its application.
- Copy number variation is one of the common types of variation in the human genome. Copy number variation includes two types of variation, amplification and deletion of gene copy number.
- the detection of gene copy number variation can be used to monitor the genome status of the subject, and can also be used to discover the association between specific diseases and certain genomic variations.
- gene copy number variation may lead to a variety of common genetic diseases, such as BRCA1/2 gene deletion may lead to the risk of hereditary breast cancer; gene copy number variation may affect the occurrence and development of tumors, such as HER2 gene amplification is not only associated with It is related to the occurrence and development of tumors, and it is also an important clinical treatment monitoring and prognostic indicator, and is an important target of tumor targeted therapy.
- the detection of copy number variation can play a vital role in the monitoring of the genome status of subjects, genome-wide association studies, prevention of genetic diseases, and precise treatment of tumors.
- subjects carrying certain copy number variations may have a higher lifetime risk of developing a disease (eg, tumor) than the general population. Therefore, the copy number variation detection method can be used to screen out subjects with higher risks, who can receive individualized monitoring of the disease, so as to achieve the purpose of early diagnosis and early treatment.
- the purpose of the present application is to provide a method for abnormal detection of gene copy number to address the above-mentioned deficiencies in the prior art. This method can at least reduce batch effects, errors and/or improve the stability of copy number detection results, which is of great significance for detecting driving events related to copy number abnormalities and interpreting tumor genome evolution information.
- the present application provides a method for detecting copy number variation and its application.
- the present application provides a copy number state analysis method, the method includes, dividing the target interval of the sample to be tested into several window areas, obtaining the sequencing data of the control window area in the sample group to be tested, based on the The sequencing data of the control window area is used to determine the copy number status of the target gene of the sample to be tested; optionally, the control window area includes a window area with a low coverage fluctuation level.
- the present application provides a copy number status analysis device, comprising the following modules: a receiving module, used to obtain the sequencing data of the sample group to be tested; a determination module, used to determine the target gene in the sample to be tested; a judging module, It is used for determining the copy number state of the target gene in the sample to be tested according to the sequencing data of the sample group to be tested.
- the present application provides a storage medium, which records a program capable of running the method described in the present application.
- Figures 1A-1B show the detection results based on the construction of the reference baseline method and the method of the present application. Each box plot represents the distribution of BRCA1 gene exon copy values in 30 samples. Groups A and B represent different batches of probe capture.
- Figure 1A shows the copy number distribution of each exon of the BRCA1 gene calculated based on the method of constructing a reference baseline.
- Figure 1B shows the distribution of copy numbers of each exon of the BRCA1 gene calculated by the method of the present application.
- Figures 2A-2B show the results of sample detection based on the difference between the reference baseline method and the method of this application for NGS library construction. Abscissa: chromosome coordinates; ordinate: estimated copy number (CN) values.
- Figure 2A shows the results of detecting copy number variations based on the construction of a reference baseline method.
- Figure 2B shows the results of detecting copy number variation by the method of the present application. Boxes indicate called copy number variations.
- Figures 3A-3B show the detection results of samples based on different thresholds set in the screening stability window. Abscissa: chromosome coordinates; ordinate: estimated copy number (CN) values.
- Figure 3A shows the results of detecting copy number variation in samples with a threshold set at 0.05.
- Figure 3B shows the results of detecting copy number variation in samples with a threshold set at 0.15. Boxes indicate called copy number variations.
- Figures 4A-4J show the detection results of 10 cases of simulated positive copy number variation samples after constructing batch baselines. Abscissa: chromosome coordinates; ordinate: estimated copy number (CN) values. The box in the figure indicates the detected copy number variation.
- Figures 5A-5F show an example of the copy number distribution diagram of part of the data of the detection results of the simulated samples in this application.
- Figures 6A-6C show an example of the copy number distribution diagram of part of the data of the test results of the standard sample in this application.
- Figures 7A-7C show examples of copy number distribution diagrams of part of the data of the test results of real samples in this application.
- Figures 8A-8F show examples of copy number distribution diagrams using different baseline detection results for standard sample 1 in this application.
- next-generation gene sequencing high-throughput sequencing
- next-generation sequencing generally refer to the second-generation high-throughput sequencing technology and the higher-throughput sequencing methods developed thereafter.
- Next-generation sequencing Platforms include but are not limited to existing sequencing platforms such as Illumina. With the continuous development of sequencing technology, those skilled in the art can understand that the sequencing methods and devices of other methods can also be used for this method. For example, second-generation gene sequencing It can have the advantages of high sensitivity, high throughput, high sequencing depth, or low cost.
- Massively Parallel Signature Sequencing Massively Parallel Signature Sequencing, MPSS
- Polony Sequencing 454pyro sequencing
- Illumina (Solexa) sequencing Illumina (Solexa) sequencing
- Ion semi conductor sequencing DNA nano-ball sequencing
- Complete Genomics DNA nanoarrays Complete Genomics DNA nanoarrays and combined probe anchored ligation sequencing, etc.
- the second-generation gene sequencing can make it possible to analyze the transcriptome and genome of a species in detail, so it is also called deep sequencing (deep sequencing)
- the method of the present application can also be applied to first-generation gene sequencing, second-generation gene sequencing, third-generation gene sequencing or single-molecule sequencing (SMS).
- SMS single-molecule sequencing
- database usually refers to an organized entity of related data, Regardless of the representation of the data or organized entities.
- said organized entities of related data may take the form of tables, maps, grids, groups, datagrams, files, documents, lists, or any other form.
- the database may include any data collected and stored in a computer-accessible form.
- the term "computing module” generally refers to a functional module for computing.
- the calculation module can calculate the output value or obtain a conclusion or result according to the input value, for example, the calculation module can be mainly used for calculating the output value.
- a computing module can be tangible, such as a processor of an electronic computer, a computer or electronic device with a processor, or a computer network, or it can be a program, command line or software package stored on an electronic medium.
- processing module generally refers to a functional module for data processing.
- the processing module may be based on processing the input value into statistically significant data, for example, it may be a classification of data for the input value.
- a processing module may be tangible, such as an electronic or magnetic medium for storing data, and a processor of an electronic computer, a computer or electronic device with a processor, or a computer network, or it may be a program stored on an electronic medium, command line or package.
- the term "judgment module” generally refers to a functional module for obtaining relevant judgment results.
- the judging module may calculate an output value or obtain a conclusion or a result according to an input value, for example, the judging module may be mainly used to obtain a conclusion or a result.
- the judging module can be tangible, such as a processor of an electronic computer, a computer with a processor or an electronic device or a computer network, or it can be a program, a command line or a software package stored on an electronic medium.
- sample obtaining module generally refers to a functional module for obtaining said sample of a subject.
- the sample obtaining module may include reagents and/or instruments required to obtain the sample (eg, tissue sample, blood sample, saliva, pleural effusion, peritoneal effusion, cerebrospinal fluid, etc.).
- lancets, blood collection tubes, and/or blood sample transport boxes may be included.
- the device of the present application may not contain or contain one or more of the sample obtaining modules, and may optionally have the function of outputting the measured value of the sample described in the present application.
- the term "receiving module” generally refers to a functional module for obtaining said measured values in said sample.
- the receiving module may input the samples described in this application (such as tissue samples, blood samples, saliva, pleural effusion, peritoneal effusion, cerebrospinal fluid, etc.).
- the receiving module may input the measured values of the samples described in the present application (such as tissue samples, blood samples, saliva, pleural effusion, peritoneal effusion, cerebrospinal fluid, etc.).
- the receiving module can detect the state of the sample.
- the data receiving module may optionally perform the gene sequencing described in this application (eg, next-generation gene sequencing) on the sample.
- the data receiving module may optionally include reagents and/or instruments required for the gene sequencing.
- the data receiving module can optionally detect sequencing depth, sequencing read length count or copy number.
- the term "copy number variation” generally refers to the amplification or deletion of the copy number of a target interval, a target gene or a target interval in a target gene.
- the copy number variation analysis method provided in this application can be used for therapeutic or diagnostic purposes.
- the copy number variation analysis method provided in the present application can be used for non-therapeutic or diagnostic purposes, such as determining whether there is a copy number variation through sequencing results.
- sliding window method generally refers to a method for dividing a window area.
- a full-length area may be divided into multiple windows according to the same or different window area lengths.
- the full-length region can be divided into multiple windows according to the same or different step lengths.
- the full-length region can be divided into multiple windows according to the same window region length and the same step size.
- a quality-qualified sample may refer to a sample with a qualified average sequencing depth, minimum sequencing depth, and/or uniformity of coverage.
- a qualified average sequencing depth can refer to samples with an average sequencing depth of about 100x or greater.
- a minimum sequencing depth qualified sample can refer to a sample with a minimum sequencing depth of about 30x or greater.
- a qualified sample with coverage uniformity may refer to a sample in which the number of bases greater than or equal to 20% of the average sequencing depth of the sample accounts for about 90% or more of the total number of bases in the sample.
- the term "unqualified target interval” generally refers to an interval with low sequencing quality.
- disqualified intervals may not be suitable for analysis of copy number variation.
- an unqualified interval may not be suitable for use as a reference or for the construction of a baseline.
- screening out unqualified intervals can improve the accuracy of test results; in other cases, not screening out unqualified intervals can also obtain test results with certain accuracy.
- an unqualified interval may refer to an interval with a low sequencing depth; for example, an unqualified interval may refer to an interval in which the interval varies greatly among different samples.
- the term "region with low capture efficiency" generally refers to a region that is not easily captured by probes that are used. For example, when there is a specific sequence combination in a certain range of sequences, it may be difficult to be captured by nucleic acid probes.
- the interval with low capture efficiency may refer to an interval with low sequencing depth.
- a low capture efficiency interval can refer to an interval with a sequencing read count of about 5 or less.
- the term "unstable interval” generally refers to an interval in which the sequencing results of different samples vary greatly. For example, it may be an interval in which multiple sequencing results in the same sample have large differences. For example, it may be an interval in which the sequencing results of different samples in the same batch vary greatly. For example, it may be an interval in which the sequencing results of different batches of different samples vary greatly. For example, it may be an interval in which the sequencing results of different reference samples differ greatly.
- the method for determining the unstable interval may be to calculate the ratio of the standard deviation to the mean value of the sequencing depth of a certain interval in different samples, and determine whether the ratio is greater than a certain threshold, such as the threshold can be 0.8, or a person skilled in the art Adjust according to the actual sequencing situation.
- sample to be tested generally refers to a sample that needs to be detected and determined whether there is a copy number variation in one or more gene regions on the sample.
- the sample to be tested or its data can be pre-stored in the memory before testing.
- human reference genome generally refers to the human genome that can function as a reference in gene sequencing.
- the information of the human reference genome can refer to UCSC (University of California, Santa Cruz).
- the human reference genome can have different versions, for example, it can be hg19, GRCH37 or ensembl 75.
- GC content generally refers to the ratio of guanine G and cytosine C to all nucleotides of the sequence in a gene sequence (base sequence).
- targeted sequencing panel or “panel” generally refers to a group/set of detection objects.
- one or more target intervals are captured and detected by designing one or more probes, and such one or more probes can form a targeted sequencing panel.
- targeted sequencing panels can be designed arbitrarily for target genes, target intervals or regions of interest, for example, for several exon regions.
- a probe may refer to an oligonucleotide that is complementary to an oligonucleotide or target nucleic acid in a region of interest under study.
- the target interval is the interval for which the probe was designed.
- sequencing depth generally refers to the number of times a specific region (eg, a specific gene, a specific interval, or a specific base) is detected.
- the sequencing depth may refer to a base sequence detected by sequencing. For example, by comparing the sequencing depth to the human reference genome, and optionally removing duplicates, the number of sequencing reads on a specific gene, a specific interval, or a specific base position can be determined and counted as the sequencing depth.
- sequencing depth can be correlated to sequencing depth. For example, sequencing depth can be affected by copy number status.
- sequencing data generally refers to data of short sequences obtained after sequencing.
- the sequencing data includes the base sequence of a sequenced short sequence (sequencing read), the number of sequencing reads, and the like.
- sequencing bias generally refers to the sequencing data bias generated in different intervals.
- the special arrangement or base ratio of the sequence in the interval can affect the sequencing read length count of the interval. For example, when a bin contains higher or lower GC content, the sequencing read counts for that bin may be biased relative to bins with GC content close to 50%.
- the term “distribution similarity” may refer to the distribution similarity of two sets of data.
- the distribution similarity in this application may refer to the similarity of the sequencing read length counts between the reference sample group and the sample to be tested in one or more intervals.
- the term “statistical distance” may refer to the distance between data values of two sets of data.
- the statistical distance may refer to the statistic of the difference between the sequencing read length counts of the reference sample group and the sample to be tested in one or more intervals.
- statistical distance can be calculated by Euclidean distance, Chebyshev distance, Mahalanobis distance, etc.
- the term "statistical value” may refer to an analytical value calculated from the data value of a sample.
- the statistical value in the present application may refer to mean value, variance, standard deviation, median value, mode value and the like. Those skilled in the art select one or more statistical values for data analysis according to actual conditions.
- probability distribution generally refers to the distribution law of the values of random variables.
- probability distributions can take different forms depending on the type of random variable they belong to.
- the normal distribution can be used as a probability distribution of a random variable.
- the term “smoothing” generally refers to a method of data processing that reduces the deviation between one or more of the differences described herein. For example, it may refer to a method of fitting scatter data to a smoothed line. For example, it can be analyzed and smoothed by the method of locally weighted regression. For example, after smoothing, the bias caused by a certain variable (such as GC content) to the sample sequencing data can be eliminated or weakened by eliminating the inherent influence of the variable (such as GC content) on the sample sequencing data.
- the smoothing process may include obtaining an average value of a certain number of difference values described in this application.
- the smoothing process may include selecting data values corresponding to different lengths according to a certain interval length, and calculating a difference between different data values.
- the smoothing process may include dividing the accumulated value of the difference within a certain length range by the interval length to obtain a ratio.
- said ratio may be considered as the average difference of said differences over the length range.
- regression generally refers to a method of statistical analysis of the relationship between variables.
- the present application can obtain a linear or non-linear relationship between sample sequencing data and a certain variable (such as GC content) through regression analysis.
- the relationship between the sequencing data of a sample and a certain variable (such as GC content) can be obtained through locally weighted regression, and the sequencing data of the sample can be adjusted/corrected through this relationship.
- correction in this application may refer to processing the sequencing data of the sample according to the relationship between the sequencing data of the sample and a certain variable so as to eliminate or reduce the bias caused by the variable to the sequencing data of the sample.
- the term "locally weighted regression” generally refers to a regression analysis method that locally introduces weights in the regression analysis of input variables and target variables.
- the local weighted regression can perform local weighted regression analysis and processing on X according to Y through the (loess(X ⁇ Y)) algorithm.
- noise reduction generally refers to the removal or reduction of noise data in data.
- noisy data generally appear as high-frequency signals
- cluster analysis generally refers to dividing similar objects into different groups by means of classification, so that member objects in the same group have similar attributes.
- K-means clustering generally refers to a method of cluster analysis. For example, through K-means clustering, a group of data can be classified into several (K) cluster analysis methods according to K cluster centers, and the sum of the distance between each data and its nearest cluster center is the smallest.
- transform analysis generally refers to a method of analyzing data.
- transformation analysis can analyze and use data for further processing by transforming the original distribution of the data into a distribution in the transformed domain that is easy to solve or process.
- transform analysis can involve discrete wavelet transforms.
- discrete wavelet transform generally refers to the discretization of the scale and translation of the underlying wavelet.
- discrete wavelet transform can be used as a method of denoising.
- normalization generally refers to a way of transforming data.
- normalization may refer to the process of transforming different sets of data to some fixed range.
- standardization may refer to the process of transforming data of different groups to the same median value.
- standardization in this application may refer to a processing method of transforming sequencing data of different samples into data at a level close to the median value.
- the term "significance test” generally refers to the way of judging whether the difference between the sample and the hypothesized distribution is significant.
- the significance test can be used to determine whether the copy number variation of the sample to be tested is a significant difference.
- normal probability distribution generally refers to the probability distribution of a random variable.
- the probability of occurrence of a random variable can be determined through the normality probability distribution and the normality probability distribution density function.
- the probability of the existence of the copy number variation in the target interval of the sample to be tested can be confirmed through a normal probability distribution.
- the term "Grubbs test” generally refers to a method of judging and/or screening out outliers. For example, by judging whether a certain value conforms to the overall distribution range, it can be determined whether the value belongs to an outlier.
- T-test generally refers to a form of statistical hypothesis testing with a Student's t distribution.
- the T-test can be used to confirm that the copy number variation of a certain target gene in the sample to be tested is significant.
- the term "about” generally refers to a range of 0.5%-10% above or below the specified value, such as 0.5%, 1%, 1.5%, 2%, 2.5%, above or below the specified value. 3%, 3.5%, 4%, 4.5%, 5%, 5.5%, 6%, 6.5%, 7%, 7.5%, 8%, 8.5%, 9%, 9.5%, or 10%.
- the present application provides a method for analyzing copy number status
- the present application provides an analysis method of copy number status, which may include obtaining the sequencing data of the sample group to be tested; determining the target gene in the sample group to be tested; determining the sequence data of the sample group to be tested The copy number status of the target gene in the sample.
- the application provides a method for analyzing copy number status, which may include the following steps:
- the application provides a method for analyzing copy number status, which may include the following steps:
- the control window area may include the first 4 or more windows covering the fluctuation level, and the ratio of the absolute deviation median to the median of the sequencing data of the window area of the qualified sample can be determined.
- the coverage fluctuation level, or the ratio of the absolute deviation median to the median of the sequencing data of all the qualified samples in the control window area may be about 0.15 or less.
- the application provides a method for analyzing copy number status, which may include the following steps:
- Step (S1-1) Obtain the sequencing data of the window area of all samples in the sample group to be tested;
- Step (S1-2) Obtain qualified samples in the sample group to be tested, and the qualified samples Can include average sequencing depth, minimum sequencing depth and/or coverage uniformity qualified samples;
- the control window area may include the first 4 or more windows covering the fluctuation level, and the ratio of the absolute deviation median to the median of the sequencing data of the window area of the qualified sample can be determined.
- Covering the fluctuation level, or the ratio of the absolute deviation median to the median of the sequencing data of all the qualified samples in the control window area may be about 0.15 or less; step (S2-1) : based on the sequencing data of the control window area, determine the normalization coefficient; step (S2-2): based on the normalization coefficient, determine the copy number of each window area of the sample to be tested; step (S2-3 ): determining the significance of the copy number variation of the sample to be tested based on the sequencing data of each window area of the sample to be tested and the sequencing data of other samples in the sample group to be tested in the corresponding window area.
- the sequencing data may include sequencing depth.
- the copy number status may comprise copy number amplifications and/or deletions.
- the copy number status may comprise exon copy number status.
- the set of samples to be tested may comprise about 10 or more samples.
- the sample set to be tested may comprise about 10 or more, about 12 or more, about 15 or more, about 20 or more, about 25 or more, about 50 or more Many, or about 100 or more samples.
- the application may not require a larger number of samples in the same batch.
- the sample set to be tested can comprise about 10 or less, about 12 or less, about 15 or less, about 20 or less, about 25 or less, or about 50 or less Fewer samples.
- the copy number status analysis method of the present application can have a higher tolerance for the copy number variation level of the sample to be tested. For example, a sample containing about 30% copy number variation can be evaluated by the analysis method of the present application.
- samples containing 10% or less, 15% or less, 20% or less, 25% or less, or 30% or less copy number variation can be evaluated by the analytical methods of the present application.
- the source of the sample in this application can be any sample containing nucleic acid, such as tissue, blood, saliva, pleural effusion, peritoneal effusion, cerebrospinal fluid, and the like.
- the step (S1) of the method of the present application may further include the step (S1-1): obtaining the sequencing data of the window regions of all samples in the sample group to be tested.
- the gene sequencing of the present application may include an optional high-throughput sequencing method or module or device.
- the sequencing can be selected from the group consisting of Solexa sequencing technology, 454 sequencing technology, SOLiD sequencing technology, Complete Genomics sequencing method and semiconductor (Ion Torrent) sequencing technology and their corresponding devices.
- the step (S1-1) of the method of the present application may include dividing the region where the target gene is located into the window region by a sliding window method.
- the step size of the windowing method can be about 24 bases.
- the window region can be about 120 bases in length.
- the step (S1-1) of the method of the present application may include obtaining the average sequencing depth of each window region after removing repeated sequencing fragments.
- the step (S1) of the method of the present application may also include the step (S1-2): obtaining qualified samples in the sample group to be tested, and the qualified samples may include the average sequencing depth, minimum sequencing depth and / or cover homogeneity acceptable samples.
- said average sequencing depth qualified samples include samples whose average sequencing depth can be about 100x or greater.
- the minimum sequencing depth qualified samples include samples for which the minimum sequencing depth can be about 30x or greater.
- the various thresholds for acceptable quality can be adjusted according to the sequencing situation.
- the coverage uniformity may be related to the sequencing depth of each base of the sample.
- the coverage uniformity can be calculated by the percentage of the number of bases greater than or equal to 20% of the average sequencing depth of the sample to the total number of bases in the sample.
- the samples that pass coverage uniformity can include samples that have a coverage uniformity of about 90% or greater.
- the samples that pass coverage uniformity can include samples that have a coverage uniformity of about 90% or greater.
- the samples passing coverage uniformity can include samples having a coverage uniformity of about 90% or greater, about 92% or greater, about 95% or greater, about 97% or greater, or about 99% or greater.
- the number of qualified samples in the sample group to be tested may be 10 or more.
- the step (S1) of the method of the present application may further include a step (S1-3): standardizing the sequencing data of the window regions of all samples in the sample group to be tested.
- the normalization may include normalizing the sequencing data of each window region of the sample based on the average sequencing depth of all window regions of the sample, and/or normalizing the GC content of each window region of the sample to The sequencing data is normalized for each window region of the sample.
- the normalization may comprise dividing the sequencing data on each window region of the sample by the sum of the sequencing data on all window regions of the sample, and multiplying by a factor.
- the factor can be set according to the size of all intervals.
- the factor may optionally be 1E+07.
- the factor may optionally be 1E+100, 1E+20, 1E+10, 1E+09, 1E+08, 1E+07, 1E+06, 1E+05, 1E+04, 1E+03, or 1E+02.
- the normalization may include normalizing the sequencing data of each window area of the sample by a regression method based on the GC content.
- the regression may comprise locally weighted regression.
- control window region may comprise a window region with a low level of coverage fluctuation.
- the coverage fluctuation level may be determined based on the statistical values of the sequencing data in the window area of the quality-qualified samples. For example, the coverage fluctuation level may be determined based on the dispersion of the sequencing data in the window area of the qualified samples. For example, the coverage fluctuation level may be determined based on the absolute deviation median and/or median of the sequencing data of the window region of the qualified samples. For example, the coverage fluctuation level may be determined based on the ratio of the absolute deviation median to the median of the sequencing data of the window area of the qualified sample.
- the window areas of the qualified samples are sorted according to the coverage fluctuation level from low to high, and the control window area may include two or more windows before the coverage fluctuation level.
- the window areas of the quality-qualified samples are sorted according to the coverage fluctuation level from low to high, and the control window area may include 4 or more windows before the coverage fluctuation level.
- the ratio of the median-to-median absolute deviation of the sequencing data for all the quality-qualified samples in the control window region may be about 0.15 or less.
- the ratio of the absolute deviation median to the median of the sequencing data of all the qualified samples in the control window area can be about 0.15 or less, about 0.14 or less, about 0.13 or Less, about 0.12 or less, about 0.11 or less, about 0.10 or less, about 0.09 or less, about 0.08 or less, about 0.07 or less, about 0.06 or less, or about 0.05 or less Small.
- the ratio of the absolute deviation median to the median of the sequencing data of all the qualified samples in the control window area may be from about 0.05 to about 0.15, from about 0.07 to about 0.15, from about 0.10 to About 0.15, about 0.12 to about 0.15, about 0.05 to about 0.12, about 0.07 to about 0.12, about 0.10 to about 0.12, about 0.05 to about 0.10, about 0.07 to about 0.10, or about 0.05 to about 0.07.
- the step (S2) of the present application may further include the step (S2-1): determining a normalization coefficient based on the sequencing data in the control window area.
- the normalization coefficient can be determined by calculating the average value of the sequencing data of all the quality-qualified samples in the control window area.
- the coverage level values of abnormal samples in the control window area may be screened out.
- the abnormal coverage level value may be the coverage level value of each of the control window areas judged to be an abnormal sample by an outlier value analysis method.
- the outlier analysis method may include Grubbs test.
- each window can contain the coverage level value of qualified samples in the batch in this window, and then the Grubbs test can be used to check whether these coverage level values contain outliers, and if so, the outliers can be removed. Then, for the remaining coverage level values, you can continue to repeat the Grubbs test to determine whether there are any abnormalities until no abnormal values appear.
- the number of samples remaining after screening out the abnormal samples may be 40% or more, 70% or more, 80% or more, 90% or more, 95% or more of the number of samples before screening out , or 99% or more.
- the step (S2) of the present application may further include the step (S2-2): determining the copy number of each window region of the sample to be tested based on the normalization coefficient.
- the step (S2-2) of the present application may include determining the normalization coefficient of the sample to be tested by normalizing the sequencing data of each window region of the sample to be tested based on the normalization coefficient. The copy number for each window region.
- the normalization method may include dividing the sequencing data of the sample to be tested in the window area by the normalization coefficient of the window area, and multiplying by the ploidy.
- the ploidy can be 1.
- the ploidy can be adjusted according to specific circumstances.
- the ploidy can be 2.
- step (S2) described in this application may also include step (S2-3): based on the sequencing data of each window area of the sample to be tested and the sequencing data of other samples in the sample group to be tested in the corresponding window area, determine the The significance of the copy number variation of test samples.
- the step (S2-3) of the present application may include determining a copy number variation candidate region based on the copy number of each window region of the sample to be tested.
- the copy number variation candidate region can be determined by region segmentation.
- the region segmentation may include determining the front and rear endpoints of the copy number variation candidate region through a cyclic binary segmentation algorithm.
- the step (S2-3) described in this application may include determining the copy number variation based on the sequencing data of the window region in the candidate region of the copy number variation of the sample to be tested and the sequencing data of other samples in the sample group to be tested in the corresponding window region. Significance of variance.
- the significance of the copy number variation can be determined by means of a significance test.
- the significance test may comprise a T-test.
- the present application also provides a copy number status analysis device, which may include the following modules: a receiving module, used to obtain the sequencing data of the sample group to be tested; a determination module, used to determine the target gene in the sample to be tested; A judging module, configured to determine the copy number status of the target gene in the sample to be tested according to the sequencing data of the sample group to be tested.
- the modules therein can be configured to be executed based on the program stored in the storage medium to implement the copy number status analysis method described in the present application.
- the present application provides a copy number status analysis method, which may include the following steps: (S1) obtaining the sequencing data of the sample to be tested and/or the sequencing data of multiple reference samples; (S2) dividing the reference samples into Two or more reference sample groups; (S3) determine the reference sample group closest to the sample to be tested; (S4) determine the sample to be tested based on the sequencing data of the reference sample group closest to the sample to be tested The copy number status of the target gene of the sample.
- the present application provides a copy number status analysis device, which may include the following modules: (M1) a receiving module for obtaining sequencing data of a sample to be tested and/or sequencing data of multiple reference samples; (M2) a processing module , for dividing the reference sample into two or more reference sample groups; (M3) calculation module, for determining the reference sample group closest to the sample to be tested; (M4) judging module, for based on the The sequencing data of the reference sample group closest to the sample to be tested is used to determine the copy number status of the target gene of the sample to be tested.
- M1 a receiving module for obtaining sequencing data of a sample to be tested and/or sequencing data of multiple reference samples
- M2 a processing module , for dividing the reference sample into two or more reference sample groups
- (M3) calculation module for determining the reference sample group closest to the sample to be tested
- (M4) judging module for based on the The sequencing data of the reference sample group closest to the sample to be tested is used to determine the copy number status of the target gene of the sample
- the application provides a copy number status analysis method, which may include the following steps:
- step (S1) Acquiring the sequencing data of the sample to be tested and/or the sequencing data of a plurality of reference samples; step (S1-1): obtaining the sequencing data of the sample to be tested and/or the reference sample by gene sequencing; Step (S1-2): correcting the sequencing data of the sample to be tested and/or the reference sample;
- step (S2) dividing the reference sample into two or more reference sample groups; step (S2-1): grouping the reference samples; step (S2-2): confirming the sequencing data of the reference sample group statistic value;
- Step (S4) Based on the sequencing data of the reference sample group closest to the sample to be tested, determine the copy number status of the target gene of the sample to be tested Step (S4-1): determine that the target gene of the sample to be tested is in The copy number CN i on the target interval i; step (S4-3): determine the copy number CN g of the sample to be tested on the target gene; step (S4-4): determine the copy number CN g on the target interval Measure the probability of the existence of the copy number variation of the sample; Step (S4-5): determine the ratio sigRatio of the presence of significant copy number amplification or deletion of the sample to be tested on the target gene; Step (S4-6 ): determine the statistical test parameters for the existence of the copy number variation of the sample to be tested on the target gene; determine the copy number status of the target gene of the sample to be tested by the following content: when CN g ⁇ CN thA , sigRatio ⁇ sigRatio th , and p
- the present application provides a copy number status analysis device, which may include a module implementing the copy number status analysis method of the present application.
- the application provides a copy number status analysis method, which may include the following steps:
- step (S1) Acquiring the sequencing data of the sample to be tested and/or the sequencing data of a plurality of reference samples; step (S1-1): obtaining the sequencing data of the sample to be tested and/or the reference sample by gene sequencing; Step (S1-2): correcting the sequencing data of the sample to be tested and/or the reference sample;
- step (S2) dividing the reference sample into two or more reference sample groups; step (S2-1): grouping the reference samples; step (S2-2): confirming the sequencing data of the reference sample group statistic value;
- Step (S4) Based on the sequencing data of the reference sample group closest to the sample to be tested, determine the copy number status of the target gene of the sample to be tested Step (S4-1): determine that the target gene of the sample to be tested is in The copy number CN i on the target interval i; step (S4-2): denoise the copy number on the target interval of the sample to be tested; step (S4-3): determine the The copy number CN g of the sample on the target gene; Step (S4-4): Determine the probability of the existence of the copy number variation of the sample to be tested on the target interval; Step (S4-5): Determine the probability of the copy number variation in the target gene The ratio sigRatio of the presence of significant copy number amplification or deletion of the sample to be tested above; step (S4-6): determine the statistical test parameters for the presence of copy number variation of the sample to be tested on the target gene ; Determine the copy number status of the target gene of the sample to be tested by the following content: when CN g ⁇ CN
- the present application provides a copy number status analysis device, which may include a module implementing the copy number status analysis method of the present application.
- the sequencing data of the present application may comprise sequencing read counts.
- the sequencing data of the present application may include the number of sequencing reads (reads) on the target gene or target interval.
- the step (S1) or module (M1) of the present application may include the step (S1-1) or module (M1-1): obtaining all the samples of the test sample and/or the reference sample by gene sequencing the sequencing data.
- the gene sequencing may comprise next-generation gene sequencing (NGS).
- NGS next-generation gene sequencing
- the gene sequencing of the present application may include an optional high-throughput sequencing method or module or device.
- the sequencing can be selected from the group consisting of Solexa sequencing technology, 454 sequencing technology, SOLiD sequencing technology, Complete Genomics sequencing method and semiconductor (Ion Torrent) sequencing technology and their corresponding devices.
- test sample and/or the reference sample may comprise a nucleic acid-containing sample.
- sample source of the present application can be any sample containing nucleic acid, such as tissue, blood, saliva, pleural effusion, peritoneal effusion, cerebrospinal fluid, etc.
- the step (S1-1) or module (M1-1) may include obtaining the sequencing data of each base in the target interval of the test sample and/or reference sample.
- the target interval may include an interval corresponding to the targeted sequencing panel sequence.
- the target interval may be about 20 to about 500 bases in length.
- the length of the target interval may be about 20 to about 500 bases, about 50 to about 500 bases, about 100 to about 500 bases, about 200 to about 500 bases, about 20 to about 200 bases, about 50 to about 200 bases, about 100 to about 200 bases, about 20 to about 100 bases, about 50 to about 100 bases, Or about 20 to about 50 bases.
- the number of target intervals may be at least about 100.
- the number of target intervals can be at least about 100, at least about 200, at least about 500, at least about 1000, or at least about 10000.
- the step (S1) or module (M1) may include a step (S1-2) or module (M1-2): correcting the sequencing data of the sample to be tested and/or the reference sample.
- the method of the present application may not include step (S1-2) or only include part of step (S1-2).
- the device of the present application may not include the module (M1-2) or only include part of the module (M1-2).
- the order of the following steps in the method steps (S1-2) of the present application can be arbitrary: standardize the sequencing data of the test sample and/or reference sample, make the test sample and/or reference sample The sequencing data is smoothed and the target interval with abnormal GC content is screened out.
- module order of the modules (M1-2) of the device of the present application can be arbitrary: normalize the sequencing data of the test sample and/or reference sample, make the test sample and/or reference The sequencing data of the sample is smoothed and the target interval with abnormal GC content is screened out.
- the step (S1-2) or module (M1-2) may include: standardizing or homogenizing the sequencing data of the test sample and/or reference sample.
- the standardization or normalization may comprise dividing the sequencing data on the target interval by the sum of the sequencing data on all target intervals of the sample corresponding to the target interval, and multiplying by a factor.
- the factor can be set according to the size of all intervals.
- the factor may optionally be 1E+07.
- the factor may optionally be 1E+100, 1E+20, 1E+10, 1E+09, 1E+08, 1E+07, 1E+06, 1E+05, 1E+04, 1E+03, or 1E+02.
- the step (S1-2) or module (M1-2) may include: smoothing the sequencing data of the test sample and/or reference sample.
- the smoothing may include smoothing the sequencing data of the test sample and/or reference sample through a regression method or a device recording the program based on the sequencing bias.
- the regression may comprise locally weighted regression.
- the sequencing bias may include the number of probes covered on the target interval.
- the sequencing bias may include the GC content of the target interval.
- the step (S1-2) or module (M1-2) may optionally include: screening out the target interval with abnormal GC content.
- the target interval with an abnormal GC content may include the target interval with a GC content of about 25% or less and/or the target interval with a GC content of about 75% or higher.
- said step (S2) or module (M2) may comprise a step (S2-1) or module (M2-1): grouping said reference samples.
- the reference sample may be from the sample to be tested, or from a sample other than the sample to be tested.
- a part of the samples to be tested can be divided as reference samples.
- the reference sample can be updated. For example, after analyzing the sequencing data of a new sample each time, the data of the new sample can be added to the existing database, and the database can be re-established.
- said grouping may comprise grouping said reference samples based on said sequencing data of said target interval.
- the grouping may include grouping the reference samples by a method of cluster analysis or a device recording the procedure.
- the method of cluster analysis may include K-means clustering, hierarchical clustering, density clustering, grid clustering, probability model clustering, or neural network model clustering, etc.
- the cluster analysis method or the device recording the program may include any clustering, classification and grouping method or the device recording the program.
- the number of reference samples may be about 30 or more.
- the number of reference samples may be about 50 or more.
- the number of reference samples can be about 30 or more, about 40 or more, about 50 or more, about 60 or more, about 70 or more, about 80 or more , about 90 or more, about 100 or more, about 200 or more, about 300 or more, about 400 or more, about 500 or more, or about 1000 or more .
- the grouping can include dividing into groups of about 2 or more.
- the grouping includes grouping into about 2 or more, about 3 or more, about 4 or more, about 5 or more, about 6 or more, about 7 or more, about 8 or more, about 9 or more, about 10 or more, about 20 or more, about 30 or more, about 40 or more, about 50 or more, about 60 or more, about 70 or more, about 80 or more, about 90 or more, or about 100 or more.
- the number of reference samples in each group can be about 30 or more.
- the number of reference samples in each group can be about 30 or more, about 40 or more, about 50 or more, about 60 or more, about 70 or more, about 80 or more, about 90 or more, about 100 or more, about 200 or more, about 300 or more, about 400 or more, about 500 or more, or about 1000 Or more.
- said step (S2) or module (M2) may comprise a step (S2-2) or module (M2-2): confirming the statistical value of said sequencing data of said reference sample set.
- the statistical values of the sequencing data of the reference sample group can be used as each candidate baseline.
- said validation statistics may comprise calculating the mean and/or standard deviation of said reference samples in each group over said target interval.
- the step (S2) or module (M2) may include a step (S2-3) or module (M2-3): screening out unqualified target intervals in the reference sample.
- the unqualified target interval may include an interval with low capture efficiency and/or an unstable interval.
- the disqualified target interval can comprise a target interval with a sequencing read count of about 5 or less.
- the unqualified target interval can comprise a sequencing read count of about 30 or less, about 20 or less, about 10 or less, about 5 or less, about 4 or less, about 3 or less , a target interval of about 2 or less, about 1 or less, or about 0 or less.
- the unqualified target interval may comprise a target interval having a coefficient of variation of about 0.8 or higher, the coefficient of variation being the standard deviation and the mean of the sequencing data for the reference samples in each group over the target interval ratio.
- the unqualified target interval can comprise a coefficient of variation of about 0.8 or higher, about 0.9 or higher, or about 1.0 or higher.
- the respective thresholds for capturing the low-efficiency interval and/or the unstable interval can be adjusted according to the sequencing situation.
- the step (S3) or module (M3) may include confirming the similarity between the test sample and the reference sample group.
- the confirming the similarity may include confirming the distribution similarity between the reference sample group and the test sample based on the sequencing data of the reference sample group and the test sample on the target interval.
- the similarity may include the degree of similarity between the reference sample group and the sequencing data of the test sample on the target interval.
- the confirming the similarity may include confirming the distribution similarity between the reference sample group and the test sample by calculating the statistical distance, the method of similarity algorithm or the device recording the program.
- the statistical distance may include a statistical value of the difference between the sequencing data of the reference sample group and the sample to be tested on the target interval.
- the statistical distance may include the statistical value of the absolute value of the difference between the sequencing data of the reference sample group and the sample to be tested on the target interval.
- the statistical distance may include a p-th power statistical value of the absolute value of the difference between the sequence data of the reference sample group and the sample to be tested on the target interval, where p is 1 or more big.
- the statistical value may comprise a summed value.
- the high similarity may include that the statistical distance between the reference sample group and the test sample is short on the target interval.
- the statistical distance may include Minkowski distance.
- the similarity algorithm may include cosine similarity, Pearson correlation coefficient, Spearman correlation coefficient, log likelihood similarity, cross entropy, and the like.
- the copy number status of the target gene in the sample to be tested may include the presence and/or amount of variation in the copy number of the target gene in the sample to be tested.
- the copy number variation may comprise amplification and/or deletion of copy number.
- the step (S4) or module (M4) may include a step (S4-1) or module (M4-1): determine the copy number CN i of the target gene of the sample to be tested on the target interval i .
- the determination of the CN i may include dividing the mean value of the sequencing data on the target interval of the target gene of the sample to be tested by the reference sample group closest to the sample to be tested on the corresponding target interval The mean value of the sequencing data is multiplied by the ploidy to obtain the CN i .
- the ploidy can be 2.
- the ploidy can be 1.
- the ploidy can be adjusted according to specific circumstances.
- the step (S4) or module (M4) may include a step (S4-2) or module (M4-2): denoising the copy number on the target interval of the sample to be tested.
- the denoising may include reducing the copy number on the target interval of the sample to be tested through transformation analysis, principal component analysis algorithm, singular value decomposition and/or Gaussian filtering method or a device recording the program. noise.
- the denoising may include denoising the copy number in the target interval of the sample to be tested by using a discrete wavelet transform method or a device recording the program.
- the denoising may include reducing the copy number on the target interval of the sample to be tested through methods such as transformation analysis, principal component analysis algorithm, singular value decomposition and/or Gaussian filtering, or a device recording the program. noise.
- the step (S4) or module (M4) may include a step (S4-3) or module (M4-3): determining the copy number CN g of the target gene in the sample to be tested.
- the target gene may comprise a gene in which the copy number variation to be determined occurs.
- the target gene may comprise a gene selected from the following group: ABL1, ABL2, ABRAXAS1, ACVR1, ACVR1B, AKT1, AKT2, AKT3, ALK, ALOX12B, AMER1, APC, AR, ARAF, ARFRP1, ARID1A, ARID1B, ARID2, ARID5B, ASXL1, ASXL2, ASXL3, ATG5, ATM, ATR, ATRX, AURKA, AURKB, AXIN1, AXIN2, AXL, B2M, BAP1, BARD1, BBC3, BCL10, BCL2, BCL2L1, BCL2L11, BCL2L2, BCL6, BCOR, BCORL1, BIRC3, BLM, BMPR1A, BRAF, BRCA1, BRCA2, BRD4, BRD7, BRINP3, BRIP1, BTG1, BTG2, BTK, CALR, CARD11, CASP8, CBFB, CBL, CCND1, CCND2, CCND3, CCNE1, CD
- the target gene may comprise genes selected from the following group: ALK (the transcript number may be NM_004304.4), ERBB2 (the transcript number may be NM_004448.3), EGFR (the transcript number may be NM_005228.3), FGFR1 (transcript number can be NM_023110.2), FGFR2 (transcript number can be NM_000141.4), CDK4 (transcript number can be NM_000075.3) and MET (transcript number can be NM_000245.3).
- ALK the transcript number may be NM_004304.4
- ERBB2 the transcript number may be NM_004448.3
- EGFR the transcript number may be NM_005228.3
- FGFR1 transcription number can be NM_023110.2
- FGFR2 transcription number can be NM_000141.4
- CDK4 transcription number can be NM_000075.3
- MET transcript number can be NM_000245.3
- the sample is selected from the group consisting of tissue samples, blood samples, saliva, pleural effusion, peritoneal effusion and cerebrospinal fluid.
- the step (S4-3) or module (M4-3) may include the length of the exon of the target gene based on the sample to be tested and the length of the target interval i of the sample to be tested.
- the copy number CN i determines the CN g .
- said step (S4-3) or module (M4-3) may comprise determining said CN g based on the following formula,
- i can represent the target interval
- j can represent the target exon
- n can represent the number of target intervals on the target exon j
- m can represent the number of target exons
- CN i can represent the copy of the target interval i Number
- Len j can represent the length of the target exon j.
- the step (S4) or module (M4) may comprise a step (S4-4) or module (M4-4): determining the probability of the existence of the copy number variation of the sample to be tested on the target interval.
- the probability of the existence of the copy number variation may include the probability of copy number amplification (p a ) and/or the probability of deletion (p d ) of the sample to be tested on the target interval.
- the step (S4-4) or module (M4-4) may include, based on the sequencing data of the sample to be tested on the target interval i, and the corresponding target interval The mean and standard deviation of the sequencing data of the close reference sample group, the probability of the existence of the copy number variation confirmed by the method of probability distribution or the device recording the program.
- the probability distribution may comprise a normality probability distribution.
- the probability distribution may comprise any common probability distribution.
- the probability distribution may comprise any discrete probability distribution.
- the probability distribution may comprise any continuous probability distribution.
- said step (S4) or module (M4) may comprise a step (S4-5) or module (M4-5): determining the significant copy number amplification or deletion of said test sample on said target gene The presence ratio of sigRatio.
- the step (S4-5) or module (M4-5) may comprise, dividing the number of target intervals with significant copy number variation on the target gene by the number of all target intervals on the target gene to obtain the Describe sigRatio.
- the target interval in which the significant copy number variation occurs may include the target interval in which the proportion of the copy number variation is about 30% or higher.
- the target interval with significant copy number variation can contain the proportion of the copy number variation is about 30% or higher, about 40% or higher, about 50% or higher, about 60% or higher , about 70% or higher, about 80% or higher, about 90% or higher, about 95% or higher, or about 95% or higher of the target interval.
- said step (S4) or module (M4) may comprise a step (S4-6) or module (M4-6): a statistical test for determining the presence of a copy number variation of said test sample on said target gene parameter.
- the statistical test parameters may comprise a p-value determined by a test of significance.
- the significance test may comprise a T-test.
- the significance test may be in any significant test manner, or a modified significance test manner according to the actual situation.
- the step (S4-6) or module (M4-6) may include, based on the number of the target interval of the sample to be tested on the target gene, the target gene on the target gene
- the sequencing data of each target interval of the sample, the standard deviation of the sequencing data of each target interval of the sample to be tested on the target gene, and the reference sample closest to the sample to be tested on the corresponding target gene The mean value and standard deviation of the sequencing data in the target interval of the group were confirmed by the method of T test or the device that described the program, and the p value p ttest was confirmed.
- the step (S4) or module (M4) can determine the copy number status of the target gene of the sample to be tested by the following content:
- CN thA can be from about 2.25 to about 4.
- CN thA can be about 2.25, about 2.50, about 2.75, about 3.00, about 3.25, about 3.50, about 3.75, or about 4.00.
- CN thD can be from about 1.0 to about 1.75.
- CN thD can be about 0.25, about 0.50, about 0.75, about 1.00, about 1.25, about 1.50, about 1.75.
- sigRatio th can be about 0.3 to about 1.
- sigRatio th can be about 0.3, about 0.4, about 0.5, about 0.6, about 0.7, about 0.8, about 0.9, or about 1.0.
- p th can be from about 0.05 to about 0.00001.
- p th can be about 0.05, about 0.01, about 0.001, about 0.0001, about 0.00001, about 0.000001, or about 0.0000001.
- the present application provides a method for establishing a database, which may include obtaining sequencing data of multiple reference samples, and dividing the reference samples into two or more reference sample groups.
- the database establishment method may include (S1) obtaining sequencing data of a sample to be tested and/or sequencing data of multiple reference samples; (S2) dividing the reference samples into two or more reference sample groups.
- the present application provides a device for establishing a database, which may include the following modules: a receiving module, configured to acquire sequencing data of a sample to be tested and/or sequencing data of a plurality of reference samples; a processing module, configured to convert the reference samples into into two or more reference sample groups.
- the database establishment device may include (M1) a receiving module for obtaining the sequence data of the sample to be tested and/or the sequence data of a plurality of reference samples; (M2) a processing module for dividing the reference sample into Two or more reference sample groups.
- the application provides a method for establishing a database, which may include the following steps:
- step (S1) Acquiring the sequencing data of the sample to be tested and/or the sequencing data of a plurality of reference samples; step (S1-1): obtaining the sequencing data of the sample to be tested and/or the reference sample by gene sequencing; Step (S1-2): correcting the sequencing data of the sample to be tested and/or the reference sample;
- step (S2) dividing the reference sample into two or more reference sample groups; step (S2-1): grouping the reference samples; step (S2-2): confirming the sequencing data of the reference sample group statistical value.
- the present application provides a database establishment device, which may include a module for realizing the database establishment method of the present application.
- the application provides a method for establishing a database, which may include the following steps:
- step (S1) Acquiring the sequencing data of the sample to be tested and/or the sequencing data of a plurality of reference samples; step (S1-1): obtaining the sequencing data of the sample to be tested and/or the reference sample by gene sequencing; Step (S1-2): correcting the sequencing data of the sample to be tested and/or the reference sample;
- step (S2) dividing the reference sample into two or more reference sample groups; step (S2-1): grouping the reference samples; step (S2-2): confirming the sequencing data of the reference sample group statistical value.
- the present application provides a database establishment device, which may include a module for realizing the database establishment method of the present application.
- the present application provides a device for establishing a database, which may include the following modules: (M1) a receiving module for obtaining sequencing data of a sample to be tested and/or sequencing data of a plurality of reference samples; (M2) a processing module, Used to divide the reference samples into two or more reference sample groups.
- M1 a receiving module for obtaining sequencing data of a sample to be tested and/or sequencing data of a plurality of reference samples
- M2 a processing module, Used to divide the reference samples into two or more reference sample groups.
- the present application provides a copy number status analysis method according to the information of the existing database, which may include determining the reference sample group closest to the sample to be tested from two or more reference sample groups, and The sequence data of the closest reference sample group is used to determine the copy number status of the target gene of the sample to be tested.
- the copy number status analysis method may include (S3) determining the reference sample group closest to the sample to be tested; (S4) determining the reference sample group closest to the sample to be tested based on the sequencing data. Describe the copy number status of the target gene in the sample to be tested.
- the present application provides a copy number status analysis device, which may include the following modules: a calculation module, used to determine the reference sample group closest to the sample to be tested from two or more reference sample groups; a judgment module, using Based on the sequencing data of the reference sample group closest to the sample to be tested, the copy number status of the target gene of the sample to be tested is determined.
- the copy number status analysis device may include (M3) a calculation module for determining the reference sample group closest to the sample to be tested; (M4) a judgment module for determining the sample group closest to the sample to be tested
- the sequence data of the reference sample group is used to determine the copy number status of the target gene of the sample to be tested.
- the application provides a copy number status analysis method, which may include the following steps:
- Step (S4) Based on the sequencing data of the reference sample group closest to the sample to be tested, determine the copy number status of the target gene of the sample to be tested Step (S4-1): determine that the target gene of the sample to be tested is in The copy number CN i on the target interval i; step (S4-3): determine the copy number CN g of the sample to be tested on the target gene; step (S4-4): determine the copy number CN g on the target interval Measure the probability of the existence of the copy number variation of the sample; Step (S4-5): determine the ratio sigRatio of the presence of significant copy number amplification or deletion of the sample to be tested on the target gene; Step (S4-6 ): determine the statistical test parameters for the existence of the copy number variation of the sample to be tested on the target gene; determine the copy number status of the target gene of the sample to be tested by the following content: when CN g ⁇ CN thA , sigRatio ⁇ sigRatio th , and p
- the present application provides a copy number status analysis device, which may include a module implementing the copy number status analysis method of the present application.
- the application provides a copy number status analysis method, which may include the following steps:
- Step (S4) Based on the sequencing data of the reference sample group closest to the sample to be tested, determine the copy number status of the target gene of the sample to be tested Step (S4-1): determine that the target gene of the sample to be tested is in The copy number CN i on the target interval i; step (S4-2): denoise the copy number on the target interval of the sample to be tested; step (S4-3): determine the The copy number CN g of the sample on the target gene; Step (S4-4): Determine the probability of the existence of the copy number variation of the sample to be tested on the target interval; Step (S4-5): Determine the probability of the copy number variation in the target gene The ratio sigRatio of the presence of significant copy number amplification or deletion of the sample to be tested above; step (S4-6): determine the statistical test parameters for the presence of copy number variation of the sample to be tested on the target gene ; Determine the copy number status of the target gene of the sample to be tested by the following content: when CN g ⁇ CN
- the present application provides a copy number status analysis device, which may include a module implementing the copy number status analysis method of the present application.
- the present application provides a database, which is established according to the copy number status analysis method or the database establishment method described in the present application.
- the present application also provides a storage medium, which records a program capable of running the method described in the present application.
- the present application also provides a device, which may include the storage medium described in the present application.
- the non-transitory computer readable storage medium may include a floppy disk, a flexible disk, a hard disk, a solid state storage (SSS) (such as a solid state drive (SSD)), a solid state card (SSC), a solid state module (SSM)), an enterprise high-grade flash drives, tape, or any other non-transitory magnetic media, etc.
- SSD solid state drive
- SSC solid state card
- SSM solid state module
- Non-transitory computer readable storage media may also include punched cards, paper tape, cursor sheets (or any other physical media having a pattern of holes or other optically identifiable markings), compact disc read only memory (CD-ROM) , Rewritable Disc (CD-RW), Digital Versatile Disc (DVD), Blu-ray Disc (BD) and/or any other non-transitory optical media.
- CD-ROM compact disc read only memory
- CD-RW Rewritable Disc
- DVD Digital Versatile Disc
- BD Blu-ray Disc
- the device of the present application may further include a processor coupled to the storage medium, and the processor may be configured to execute based on a program stored in the storage medium to implement the method described in the present application.
- the present application also provides a method of the present application, which can be applied in disease diagnosis, prevention and/or treatment.
- the present application also provides a method of the present application, which can be applied in the monitoring of the copy number status of the target gene.
- the present application also provides a method of the present application, which can be applied in genome-wide association studies.
- the method can be used to determine whether the subject has a copy number variation.
- any one or more of the methods of the present application may be for non-diagnostic purposes.
- any one or more of the methods of the present application may be for diagnostic purposes.
- the method can be used in clinical practice by detecting the copy number variation (for example, it can be inferred whether certain specific tumor treatment methods are suitable for the subject).
- the level of copy number variation detected by the method can be used in clinical practice in combination with biomarkers known in the art.
- the method of the present application is used to detect copy number variation.
- the copy number variation detection algorithm of this application can select a sufficient number of samples, for example, 15 cases from the same sample type and the same experimental methodology, and try to ensure that the reagent batches, experimental equipment, etc. used in the experiment are consistent sample data .
- Each participating sample data needs to come from the BAM file after comparison of NGS sequencing data.
- the repetitive DNA sequence fragments introduced by PCR during NGS library construction can be removed to obtain uniquely compared DNA fragments.
- the sliding window method is used to slide 24bp each time, and the region is divided into window regions with a probe fixed length of 120bp, and the average coverage of the uniquely aligned DNA fragments in each window is counted level.
- the average sequencing depth is required to be ⁇ 100X
- the minimum sequencing depth is ⁇ 30X
- the detection method of this application can meet the quality requirements
- the number of samples is at least 10 cases for testing.
- the coverage level of each window area can be corrected, including preliminary coverage level correction (based on the sample average coverage level), GC correction and batch correction.
- the preliminary coverage level correction is to correct the coverage levels of all samples in the batch to the same specified coverage level. Specifically, for each window area of the sample in the batch, the average coverage level obtained by sequencing is divided by the sum of the average coverage levels of all window areas in the sample, and then multiplied by a fixed factor (the factor is 1E+07).
- GC correction is performed by calculating the GC content of each window, and then using the loess regression method to perform GC bias correction on the coverage level of each window region in the sample.
- MAD median and median absolute deviation
- peripheral blood samples were selected to detect the exon copy number variation (LGR) of BRCA1 and BRCA2.
- LGR exon copy number variation
- the experiment used RNA probes to specifically capture BRCA1 and BRCA2 gene regions, and then performed high-throughput sequencing to compare the sequencing data with the standard sequence of the human genome hg19 for comparison to obtain the compared BAM file. Subsequently, the method based on the construction of the reference baseline and the method of the present application were used to detect the copy number variation. At the same time, all sample copy number variations were confirmed by the BRCA MASTR Plus Dx kit (based on multiplex PCR capture methodology), including a total of 17 LGR positive samples and 679 negative samples.
- the data of 14 cell line samples after sequencing and comparison were selected to construct batch baselines.
- the thresholds for describing the coverage fluctuation level of the description window were set to 0.05 and 0.15, respectively, and two batch baselines were constructed. Then, the samples with known LGR copy number variation (BRCA1:exon4-6del) among the 14 samples were corrected in batches with 2 batch baselines, and then the copy number variation was detected.
- a batch baseline was constructed from 10 simulated positive samples, and then batch correction and copy number variation identification were performed on the 10 simulated samples using the constructed batch baseline.
- the results of the 10 simulated samples are shown in Figures 4A-4J, and the 10 simulated copy number variations can all be detected accurately, indicating that the present application can achieve accurate detection of copy number variations in any region.
- the method of this application uses the clustering method to divide a large amount of real sample data into different sample sets according to the trend clustering of the sequencing depth, and constructs the baselines (average depth and depth fluctuation range) respectively.
- the baselines average depth and depth fluctuation range
- the discrete wavelet transform method can optionally be used to smooth and denoise the copy number to improve the signal-to-noise ratio of the sequencing data.
- the method of the present application detects copy number variation based on the specificity of sample coverage characteristics. Specifically, based on the cluster analysis of large-scale samples, this application constructs multiple control group baselines, which can avoid the baseline mismatch problem caused by inconsistent coverage depth characteristics due to differences in experiments and samples in sequencing; and can integrate multiple coverage depths Correction strategies to reduce sample-specific data differences; ultimately, quantitative analysis and statistical difference evaluation can be used to ensure the accuracy and stability of results.
- the copy number variation detection method of the present application can be applied not only to the targeted capture sequencing data of a specific gene panel, but also to the capture sequencing data of the whole exome.
- the database establishment method of the present application may include the following steps:
- Data preparation module including:
- Coverage depth correction module including 3 independent corrections with optional order:
- i represents the site on the target interval
- n represents the total number of sites on all target intervals
- RD i represents the sequencing depth of site i on the target interval
- R is a constant that can be set according to the size of all intervals to ensure that The corrected depth of the test sample is at the same level as the corrected depth of the reference sample.
- b) Smoothing the sequencing data 1 correcting the characteristics of probe laying, specifically, according to the difference in probe laying multipliers in different intervals in the probe design, such as the number of probes covered on the interval, dividing the interval, each The length of the target interval can be about 24 base pairs, and the average coverage depth RD of each target interval is calculated. According to the number of probes covered on each target interval ProbeN, the sequencing depth on the target interval is locally weighted regression (loess(RD ⁇ ProbeN)) correction, obtain the sequencing depth RD normP after probe correction;
- GC correction specifically, extending the target interval used for coverage depth calculation to a total length greater than 200bp according to the flanks, calculating the average GC ratio, and according to the GC content of each interval GC, Perform local weighted regression (loess(RD ⁇ GC)) correction on the sequencing depth RD to obtain the GC-corrected sequencing depth RD normGC ;
- Baseline building blocks specifically including the following steps:
- a) Sample clustering Existing methods generally use all reference samples as a category to construct a baseline. The method of the present application groups the reference samples, specifically based on the consistency of the coverage depth of each reference sample in the target interval, such as the degree of approximation of the sequencing of the reference samples in the target interval, cluster analysis is performed, and the reference samples Divided into different categories of reference sample groups, the clustering method can be such as K-means clustering, hierarchical clustering and other methods;
- the setting of the number of reference sample groups needs to take into account the characteristics of tumor samples and the quality of sequencing, and determine the number of reference sample groups according to the number of captured features. For example, the number of reference sample groups can be more than 2, such as 2-10.
- interval screening can be performed: calculate the coefficient of variation cv of the sequencing depth on each interval, and remove the unstable intervals with large fluctuations in the sample, where:
- this interval is considered to be an unstable region and is filtered out; at the same time, when the corrected sequencing depth is lower than 5, it is considered to be a region with low capture efficiency and is filtered out, and the final reserved interval is regarded as a stable region. interval.
- the database of the present application can be obtained, including two or more reference sample groups whose changes in the target interval are consistent.
- the advantage of the database of this application lies in: through the cluster analysis of large-scale samples, the reference samples are divided into different types of reference sample groups, and the sample-specific background baselines are respectively constructed, which greatly reduces the False positives caused by batch effects in high-throughput sequencing data in copy number variation detection increase the stability of results.
- the method for eliminating the batch effect of the present application does not need to ensure a sufficient number of samples of the same gene panel in the same batch, which greatly reduces the difficulties in practical application.
- the copy number status analysis method of the present application may include the following steps:
- a) According to the similarity between the sample to be tested and the reference sample group, determine the reference sample group closest to the sample to be tested, that is, dynamically filter the baseline: by calculating the statistical distance, such as Minkowski distance, etc. , comparing the sequencing depth of the sample to be tested on each target interval with the sequencing depth of each reference sample group on the target interval, and confirming the statistical distance between the reference sample group and the sample to be tested:
- the L p value represents the statistical distance
- i represents the target interval
- n represents the number of target intervals
- RD sample represents the sequencing depth of each target interval of the sample to be tested, Indicates the sequencing depth of each target interval of the reference sample group closest to the sample to be tested, where the ploidy can be 2.
- the copy number smoothing and noise reduction of each interval can be performed: the CNi of each interval can be smoothed and denoised using a noise reduction algorithm to improve the signal-to-noise ratio of the data.
- the noise reduction method can use discrete wavelet transform (Discrete Wavelet Transformation, DWT), principal component analysis algorithm, singular value decomposition and/or Gaussian filtering method for smooth noise reduction.
- DWT divides the signal into high-frequency signal and low-frequency signal, passes through low-pass filter and high-pass filter respectively, performs discrete wavelet transform on discrete signal, discretizes continuous wavelet and its wavelet transform, and achieves the purpose of data noise reduction. In this way, the CNi after noise reduction can be obtained;
- Evaluation of the copy number of each target gene calculate the weighted average copy number CN g of each target gene in the sample, and use the length of the target exon to correct CN i , for example:
- i represents the target interval
- j represents the target exon
- n represents the number of target intervals on target exon j
- m represents the number of target exons
- CN i represents the copy number of target interval i
- Len j represents Length of target exon j.
- the target interval with significant copy number variation includes the target interval with a proportion of copy number variation of about 30% or higher.
- Each threshold can be obtained by using a large-scale sample training.
- CN thA represents the threshold of copy number amplification, and the value can be selected from 2.25 to 4
- CN thD represents the threshold of copy number deletion, and the value can be selected from 1.0 to 1.75
- sigRatio th represents significant amplification/deletion
- the threshold of the ratio can be selected from 0.3 to 1
- p th represents the threshold of the significant T test, and the value can be selected from 0.05 to 0.00001.
- the copy number status analysis method of this application dynamically screens the reference sample group closest to the sample to be tested according to the similarity between the sample to be tested and the reference sample group as the background baseline, which can eliminate the batch effect and improve the specificity and sensitivity of detection sex.
- Database establishment 655 reference samples were used to construct the baseline, and the database construction method of this application was adopted, for example, the k-means clustering algorithm was used to divide the reference samples into 5 reference sample groups, and 5 different candidate baselines were constructed as the database.
- Construct simulation data use varBen tumor mutation data simulation software (github.com/nccl-jmli/VarBen), based on benign tissue samples, by inserting gene reads, insert target gene reads into the sequencing data, The gradient simulates the amplification of different copy numbers of the target gene, and the simulated sample list is shown in Table 4.
- target gene Number of simulated samples Simulate copy number gradients ALK 20 2.5, 2.75, 3.0, 3.5, 4.0 ERBB2 20 2.5, 2.75, 3.0, 3.5, 4.0 FGFR1 20 2.5, 2.75, 3.0, 3.5, 4.0 FGFR2 20 2.5, 2.75, 3.0, 3.5, 4.0
- Figures 5A-5F show examples of copy number distribution diagrams of part of the test result data of the present application.
- Each dot represents an interval of a gene
- gray dots represent genes with normal copy number
- black dots represent genes with copy number amplification or deletion
- the corresponding gene names are marked.
- the horizontal axis indicates the chromosome position of the gene
- the vertical axis indicates the copy number calculated based on the method of this application (the middle horizontal line indicates the copy number of the normal gene)
- the gray background indicates the background baseline (the reference sample group closest to the sample to be tested) The fluctuation range of each target interval in .
- Figures 5A-5C simulate different degrees of copy number amplification of the ERBB2 gene
- Figures 5D-5F show different levels of copy number amplification of the FGFR1 gene
- the simulated copy number gradients are 2.5, 2.75 and 3.0.
- the results show that the copy number status analysis method of the present application is used in simulated samples, all simulated genes and copy number amplifications of different gradients can be stably detected, and the copy number prediction is accurate.
- Positive standard samples This application test includes 30 cases of positive standard samples, which are derived from NCI-BL2009 cell line, using plasmid transfection to transfect the corresponding proportion of the target gene into the cell line to obtain CNV positive data, and using microdroplets Gene copy number was quantified by digital PCR (ddPCR).
- the plasmid numbers are: Life RPCI11.C-433C10BAC-EGFR, Life RPCI11.C-936I7BAC-CDK4, Life RPCI11.C-163C9BAC-MET, Life RPCI11.C-909L6BAC-ERBB2, Life RPCI11.C-957P17BAC-FGFR1.
- the list of positive standard samples is shown in Table 6.
- Database establishment 655 reference samples were used to construct the baseline, and the database construction method of this application was adopted, for example, the k-means clustering algorithm was used to divide the reference samples into 5 reference sample groups, and 5 different candidate baselines were constructed as the database.
- the copy number status of the copy number amplification positive standard sample was detected according to the copy number status analysis method of the present application, and the detection results are shown in Table 7.
- Figures 6A-6C show examples of copy number distribution diagrams of part of the test result data of the present application.
- 6A-6C show the detection results of standard samples of CNV-positive cell lines transfected with plasmids, and the ddPCR-marked copy numbers are 3, 5 and 8, respectively.
- the results show that the method of the present invention is used in the cell line standard for plasmid transfection, all genes and different copy number states can be stably detected, and the copy number prediction is accurate.
- Real data The real samples tested in this application include 20 cases of ERBB2 amplification positive samples verified by a third-party immunohistochemical method (IHC). The list of real samples is shown in Table 8.
- Database construction use 443 reference samples to construct baselines, adopt the database construction method of this application, for example, use k-means clustering algorithm, divide reference samples into reference sample groups, and construct different candidate baselines as databases.
- Figures 7A-7C show examples of copy number distribution diagrams of part of the test result data of the present application.
- 7A-7C show the detection results of real ERBB2 positive samples. The results showed that the method of the present application was applied to real samples, and all 20 samples whose IHC results were positive for HER2 could be stably detected.
- Positive standard samples The test of this application includes 3 cases of positive standard samples, which are from the same source as in Example 9, and are used to detect the test results of different baselines. The list of positive standard samples is shown in Table 10.
- Database establishment 655 reference samples were used to construct the baseline, and the database construction method of this application was adopted, for example, the k-means clustering algorithm was used to divide the reference samples into 5 reference sample groups, and 5 different candidate baselines were constructed as the database. At the same time, without using the clustering method, all reference samples were constructed as a baseline.
- Figures 8A-8F show examples of copy number distributions of standard sample 1 using different baseline detection results.
- the results show that the optimal baseline matched by the method of the present application is the closest to the sample to be tested (the distance between the sample and the baseline is the smallest), the fluctuation (SD) of the overall copy number of the sample is the lowest, the copy number distribution map is the most stable, and the noise is the smallest , indicating that the detection results of this method are more stable.
- the method of this application can stably detect all genes and different copy number states, while other baselines cannot be stably detected when the copy number is 3.
Landscapes
- Life Sciences & Earth Sciences (AREA)
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Health & Medical Sciences (AREA)
- Chemical & Material Sciences (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- General Health & Medical Sciences (AREA)
- Biophysics (AREA)
- Biotechnology (AREA)
- Analytical Chemistry (AREA)
- Bioinformatics & Computational Biology (AREA)
- Evolutionary Biology (AREA)
- Molecular Biology (AREA)
- Theoretical Computer Science (AREA)
- Spectroscopy & Molecular Physics (AREA)
- Genetics & Genomics (AREA)
- Medical Informatics (AREA)
- Organic Chemistry (AREA)
- Wood Science & Technology (AREA)
- Zoology (AREA)
- Biochemistry (AREA)
- Immunology (AREA)
- Microbiology (AREA)
- General Engineering & Computer Science (AREA)
- Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
- Apparatus Associated With Microorganisms And Enzymes (AREA)
Abstract
Description
| 附图名称 | 基因 | 拷贝数变异 |
| BRCA1外显子16-19拷贝数扩增(图4A) | BRCA1 | 外显子16-19_Dup |
| BRCA1外显子3-7拷贝数扩增(图4B) | BRCA1 | 外显子3-7_Dup |
| BRCA2外显子14-18拷贝数扩增(图4C) | BRCA2 | 外显子14-18_Dup |
| BRCA2外显子3拷贝数扩增(图4D) | BRCA2 | 外显子3_Dup |
| BRCA2外显子25-27拷贝数扩增(图4E) | BRCA2 | 外显子25-27_Dup |
| BRCA1外显子15-22拷贝数缺失(图4F) | BRCA1 | 外显子15-22_Del |
| BRCA1外显子2拷贝数缺失(图4G) | BRCA1 | 外显子2_Del |
| BRCA1外显子7-12拷贝数缺失(图4H) | BRCA1 | 外显子7-12_Del |
| BRCA2外显子15-16拷贝数缺失(图4I) | BRCA2 | 外显子15-16_Del |
| BRCA2外显子8-11拷贝数缺失(图4J) | BRCA2 | 外显子8-11_Del |
| 目标基因 | 模拟样本数 | 模拟拷贝数梯度 |
| ALK | 20 | 2.5,2.75,3.0,3.5,4.0 |
| ERBB2 | 20 | 2.5,2.75,3.0,3.5,4.0 |
| FGFR1 | 20 | 2.5,2.75,3.0,3.5,4.0 |
| FGFR2 | 20 | 2.5,2.75,3.0,3.5,4.0 |
| 基因 | 标准样本数 | ddPCR标定拷贝数 |
| CDK4 | 10 | 3,5,8 |
| ERBB2 | 10 | 3,5,8 |
| EGFR | 10 | 3,5,8 |
| FGFR1 | 10 | 3,5,8 |
| MET | 10 | 3,5,8 |
| 基因 | 样本数 | IHC结果 |
| ERBB2 | 20 | 拷贝数:3+ |
| 样本 | 基因 | ddPCR标定拷贝数 |
| 标准样本1 | CDK4,ERBB2,EGFR,FGFR1,MET | 3 |
| 标准样本2 | CDK4,ERBB2,EGFR,FGFR1,MET | 5 |
| 标准样本3 | CDK4,ERBB2,EGFR,FGFR1,MET | 8 |
Claims (19)
- 一种拷贝数状态分析方法,所述方法包含,将待测样本的目标区间划分为若干个窗口区域,获取待测样本组中的对照窗口区域的测序数据,基于所述对照窗口区域的测序数据,确定所述待测样本的目标基因的拷贝数状态;任选地,所述对照窗口区域包含覆盖波动水平低的窗口区域。
- 如权利要求1所述的方法,所述方法还包含以下步骤:(S1)获取所述待测样本的测序数据和/或多个参考样本的测序数据;(S2)将所述参考样本分为两个或以上参考样本组;(S3)确定与所述待测样本最接近的参考样本组;(S4)基于所述与待测样本最接近的参考样本组的测序数据,确定所述待测样本的目标基因的拷贝数状态。
- 如权利要求1-2中任一项所述的方法,所述方法包含,将质量合格的样品的窗口区域按照覆盖波动水平从低到高排序,所述对照窗口区域包含覆盖波动水平前2个或更多、或前4个或更多所述窗口,或者所述对照窗口区域的所有所述质量合格的样品的所述测序数据的绝对离差中位数与中位数的比值为约0.15或更小。
- 如权利要求3所述的方法,所述方法包含,基于所述质量合格的样品的窗口区域的测序数据统计值,确定所述覆盖波动水平;任选地,基于所述质量合格的样品的窗口区域的测序数据的绝对离差中位数与中位数的比值,确定所述覆盖波动水平。
- 如权利要求3-4中任一项所述的方法,所述方法包含,基于所述对照窗口区域的测序数据,确定归一化系数;任选地,通过计算所述对照窗口区域的所有所述质量合格的样品的测序数据平均值,确定所述归一化系数。
- 如权利要求5所述的方法,所述方法包含,基于所述归一化系数,确定待测样本的每一个窗口区域的拷贝数;任选地,所述归一化包含将所述窗口区域的待测样本的测序数据除以所述窗口区域的归一化系数,乘以倍性。
- 如权利要求1-6中任一项所述的方法,所述方法包含,基于待测样本的每一个窗口区域的测序数据以及相应窗口区域的待测样本组中其它样本的测序数据,确定待测样本的拷贝数变异显著性;任选地,通过T检验 显著性检验的方法,确定所述拷贝数变异的显著性。
- 如权利要求1-7中任一项所述的方法,所述待测样本选自以下组:组织样本、血液样本、唾液、胸腔积液、腹膜积液和脑脊液。
- 如权利要求2-8中任一项所述的方法,所述步骤(S2)包含步骤(S2-1):使所述参考样本分组,所述分组包含基于目标区间的所述测序数据通过聚类分析的方法使所述参考样本分组;优选地,所述聚类分析的方法包含K均值聚类和/或层次聚类;所述步骤(S2)包含步骤(S2-2):确认所述参考样本组的所述测序数据的统计值;优选地,所述确认统计值包含计算在所述目标区间上每组中所述参考样本的均值和/或标准差。
- 如权利要求2-9中任一项所述的方法,所述步骤(S3)包含基于在目标区间上所述参考样本组与所述待测样本的所述测序数据,通过计算统计距离的方法,确认所述参考样本组与所述待测样本的分布相似程度;优选地,所述分布相似程度高包含在所述目标区间上所述参考样本组与所述待测样本的所述统计距离短。
- 如权利要求10所述的方法,所述统计距离包含所述目标区间上所述参考样本组与所述待测样本的所述测序数据的差值的绝对值的p次方的统计值,所述p为1或更大,优选地,所述统计值包含求和值;优选地,所述统计距离包含闵可夫斯基距离。
- 如权利要求2-12中任一项所述的方法,所述步骤(S4)包含:确定在目标区间上待测样本的拷贝数变异的存在的概率,所述拷 贝数变异的存在的概率包含在目标区间上所述待测样本发生拷贝数扩增的概率(p a)和/或缺失的概率(p d);优选地,所述步骤(S4)包含,基于所述待测样本在目标区间i上的测序数据、以及在相应目标区间上所述与待测样本最接近的参考样本组的测序数据的均值与标准差,通过概率分布的方法确认所述拷贝数变异的存在的概率;优选地,所述概率分布包含正态性概率分布。
- 如权利要求2-13中任一项所述的方法,所述步骤(S4)包含:确定在所述目标基因上所述待测样本的显著性拷贝数扩增或缺失的存在的比例sigRatio;优选地,使所述目标基因上发生显著性拷贝数变异的目标区间数量除以所述目标基因上所有目标区间数量,得到所述sigRatio;所述发生显著性拷贝数变异的目标区间包含所述拷贝数变异的比例为约30%或更高的目标区间;所述步骤(S4)还包含:确定在所述目标基因上所述待测样本的拷贝数变异的存在的统计检验参数;优选地,基于在所述目标基因上所述待测样本的所述目标区间的数量、在所述目标基因上所述待测样本的各个目标区间的测序数据、在所述目标基因上所述待测样本的各个目标区间的测序数据的标准差、以及在相应目标基因上述与待测样本最接近的参考样本组的目标区间上的测序数据的均值和标准差,通过T检验的方法确认p值p ttest。
- 如权利要求14所述的方法,所述步骤(S4)通过以下内容确定所述待测样本的目标基因的拷贝数状态:当CN g≥CN thA,sigRatio≥sigRatio th,且p ttest≤p th时,确认所述待测样本的目标基因发生拷贝数扩增;当CN g≤CN thD,sigRatio≥sigRatio th,且p ttest≤p th时,确认所述待测样本的目标基因发生拷贝数缺失;当CN thA<CN g<CN thD,或sigRatio<sigRatio th,或p ttest>p th时,确认所述待测样本的目标基因拷贝数正常,其中CN thA,CN thD,sigRatio th,和p th各自独立地为阈值;优选地,其中CN thA为约2.25至约4;优选地,其中CN thD为约1.0至约1.75;优选地,其中sigRatio th为约0.3至约1;优选地,其中p th为约0.05至约0.00001。
- 如权利要求2-15中任一项所述的方法,所述目标基因包含选自以下组基因:ABL1、ABL2、ABRAXAS1、ACVR1、ACVR1B、AKT1、AKT2、AKT3、ALK、ALOX12B、AMER1、APC、AR、ARAF、ARFRP1、ARID1A、ARID1B、ARID2、ARID5B、ASXL1、ASXL2、ASXL3、 ATG5、ATM、ATR、ATRX、AURKA、AURKB、AXIN1、AXIN2、AXL、B2M、BAP1、BARD1、BBC3、BCL10、BCL2、BCL2L1、BCL2L11、BCL2L2、BCL6、BCOR、BCORL1、BIRC3、BLM、BMPR1A、BRAF、BRCA1、BRCA2、BRD4、BRD7、BRINP3、BRIP1、BTG1、BTG2、BTK、CALR、CARD11、CASP8、CBFB、CBL、CCND1、CCND2、CCND3、CCNE1、CD274、CD28、CD58、CD74、CD79A、CD79B、CDC73、CDH1、CDH18、CDK12、CDK4、CDK6、CDK8、CDKN1A、CDKN1B、CDKN1C、CDKN2A、CDKN2B、CDKN2C、CEBPA、CENPA、CHD1、CHD2、CHD4、CHD8、CHEK1、CHEK2、CIC、CIITA、CREBBP、CRKL、CRLF2、CRYBG1、CSF1R、CSF3R、CSMD1、CSMD3、CTCF、CTLA4、CTNNA1、CTNNB1、CUL3、CUL4A、CXCR4、CYLD、CYP17A1、CYP2D6、DAXX、DCUN1D1、DDR1、DDR2、DDX3X、DICER1、DIS3、DNAJB1、DNMT1、DNMT3A、DNMT3B、DOT1L、DPYD、DTX1、DUSP22、EED、EGFR、EIF1AX、EIF4E、EMSY、EP300、EPCAM、EPHA2、EPHA3、EPHA5、EPHA7、EPHB1、EPHB4、ERBB2、ERBB3、ERBB4、ERCC1、ERCC2、ERCC3、ERCC4、ERCC5、ERG、ERRFI1、ESR1、ETV4、ETV5、ETV6、EWSR1、EZH2、EZR、FANCA、FANCC、FANCD2、FANCE、FANCF、FANCG、FANCI、FANCL、FANCM、FAS、FAT1、FAT3、FBXW7、FGF10、FGF12、FGF14、FGF19、FGF23、FGF3、FGF4、FGF6、FGF7、FGFR1、FGFR2、FGFR3、FGFR4、FH、FLCN、FLT1、FLT3、FLT4、FOXA1、FOXL2、FOXO1、FOXO3、FOXP1、FRS2、FUBP1、FYN、GABRA6、GALNT12、GATA1、GATA2、GATA3、GATA4、GATA6、GEN1、GID4、GLI1、GNA11、GNA13、GNAQ、GNAS、GPS2、GREM1、GRIN2A、GRM3、GSK3B、H3F3A、H3F3B、H3F3C、HDAC1、HDAC2、HGF、HIST1H1C、HIST1H2BD、HIST1H3A、HIST1H3B、HIST1H3C、HIST1H3D、HIST1H3E、HIST1H3G、HIST1H3H、HIST1H3I、HIST1H3J、HIST2H3D、HIST3H3、HLA-A、HLA-B、HLA-C、HNF1A、HOXB13、HRAS、HSD3B1、HSP90AA1、ICOSLG、ID3、IDH1、IDH2、IFNGR1、IGF1、IGF1R、IGF2、IGHD、IGHJ、IGHV、IKBKE、IKZF1、IL10、IL7R、INHA、INHBA、INPP4A、INPP4B、INSR、IRF2、IRF4、IRS1、IRS2、ITK、ITPKB、JAK1、JAK2、JAK3、JUN、KAT6A、KDM5A、KDM5C、KDM6A、KDR、KEAP1、KEL、KIR2DL4、KIR3DL2、KIT、KLF4、KLHL6、KLRC1、KLRC2、KLRK1、KMT2A、KMT2C、KMT2D、KRAS、LATS1、LATS2、LMO1、 LRP1B、LTK、LYN、MAF、MAGI2、MALT1、MAP2K1、MAP2K2、MAP2K4、MAP3K1、MAP3K13、MAP3K14、MAPK1、MAPK3、MAX、MCL1、MDC1、MDM2、MDM4、MED12、MEF2B、MEN1、MERTK、MET、MFHAS1、MGA、MIR21、MITF、MKNK1、MLH1、MLH3、MPL、MRE11、MSH2、MSH3、MSH6、MST1、MST1R、MTAP、MTOR、MUTYH、MYC、MYCL、MYCN、MYD88、MYOD1、NAV3、NBN、NCOA3、NCOR1、NCOR2、NEGR1、NF1、NF2、NFE2L2、NFKBIA、NKX2-1、NKX3-1、NOTCH1、NOTCH2、NOTCH3、NOTCH4、NPM1、NRAS、NRG1、NSD1、NSD2、NSD3、NT5C2、NTHL1、NTRK1、NTRK2、NTRK3、NUP93、NUTM1、P2RY8、PAK1、PAK3、PAK5、PALB2、PALLD、PARP1、PARP2、PARP3、PAX5、PBRM1、PCDH11X、PDCD1、PDCD1LG2、PDGFRA、PDGFRB、PDK1、PGR、PHOX2B、PIK3C2B、PIK3C2G、PIK3C3、PIK3CA、PIK3CB、PIK3CD、PIK3CG、PIK3R1、PIK3R2、PIK3R3、PIM1、PLCG2、PLK2、PMS1、PMS2、PNRC1、POLD1、POLE、POM121L12、PPARG、PPM1D、PPP2R1A、PPP2R2A、PPP6C、PRDM1、PREX2、PRKAR1A、PRKCI、PRKDC、PRKN、PTCH1、PTEN、PTPN11、PTPRD、PTPRO、PTPRS、PTPRT、QKI、RAB35、RAC1、RAD21、RAD50、RAD51、RAD51B、RAD51C、RAD51D、RAD52、RAD54L、RAF1、RARA、RASA1、RB1、RBM10、RECQL4、REL、RET、RHEB、RHOA、RICTOR、RIT1、RNF43、ROS1、RPA1、RPS6KA4、RPS6KB2、RPTOR、RSPO2、RUNX1、RUNX1T1、SDC4、SDHA、SDHAF2、SDHB、SDHC、SDHD、SETD2、SF3B1、SGK1、SH2B3、SH2D1A、SHQ1、SLC34A2、SLIT2、SLX4、SMAD2、SMAD3、SMAD4、SMARCA4、SMARCB1、SMARCD1、SMO、SNCAIP、SOCS1、SOX10、SOX17、SOX2、SOX9、SPEN、SPI1、SPOP、SPTA1、SRC、SRSF2、STAG2、STAT3、STAT4、STAT5A、STAT5B、STAT6、STK11、STK40、SUFU、SYK、TAF1、TBX21、TBX3、TCF3、TCF7L2、TEK、TENT5C、TERC、TERT、TET1、TET2、TGFBR1、TGFBR2、TIPARP、TMEM127、TMPRSS2、TNFAIP3、TNFRSF14、TOP1、TOP2A、TP53、TP63、TP73、TRAF2、TRAF3、TRAF7、TRIM58、TRPC5、TSC1、TSC2、TSHR、TYRO3、U2AF1、UGT1A1、VEGFA、VEGFB、VEGFC、VHL、WISP3、WRN、WT1、XIAP、XPO1、XRCC2、XRCC3、YAP1、YES1、ZAP70、ZBTB16、ZBTB2、ZNF217、ZNF703和ZNRF3。
- 一种拷贝数状态分析装置,包含以下模块:接收模块,用于获取待测样本组的测序数据;确定模块,用于确定待测样本中的目标基因;判断模块,用于根据所述待测样本组的测序数据确定所述待测样本中的目标基因的拷贝数状态。
- 如权利要求17所述的拷贝数状态分析装置,包含以下模块:(M1)接收模块,用于获取待测样本的测序数据和/或多个参考样本的测序数据;(M2)处理模块,用于将所述参考样本分为两个或以上参考样本组;(M3)计算模块,用于确定与所述待测样本最接近的参考样本组;(M4)判断模块,用于基于所述与待测样本最接近的参考样本组的测序数据,确定所述待测样本的目标基因的拷贝数状态。
- 一种储存介质,其记载可以运行权利要求1-16中任一项所述的方法的程序。
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP22863369.9A EP4397773A4 (en) | 2021-08-30 | 2022-08-29 | METHOD FOR DETECTING VARIATION IN COPY NUMBER AND ITS APPLICATION |
| JP2024514011A JP7745090B2 (ja) | 2021-08-30 | 2022-08-29 | コピー数変異の検出方法およびその応用 |
Applications Claiming Priority (4)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202111002171.4 | 2021-08-30 | ||
| CN202111002171.4A CN113674803B (zh) | 2021-08-30 | 2021-08-30 | 一种拷贝数变异的检测方法、装置、存储介质及其应用 |
| CN202111095132.3 | 2021-09-17 | ||
| CN202111095132.3A CN113789371B (zh) | 2021-09-17 | 2021-09-17 | 一种基于批次矫正的拷贝数变异的检测方法 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2023030233A1 true WO2023030233A1 (zh) | 2023-03-09 |
Family
ID=85410853
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2022/115447 Ceased WO2023030233A1 (zh) | 2021-08-30 | 2022-08-29 | 一种拷贝数变异的检测方法及其应用 |
Country Status (3)
| Country | Link |
|---|---|
| EP (1) | EP4397773A4 (zh) |
| JP (1) | JP7745090B2 (zh) |
| WO (1) | WO2023030233A1 (zh) |
Cited By (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN117153249A (zh) * | 2023-10-26 | 2023-12-01 | 北京华宇亿康生物工程技术有限公司 | 用于检测smn基因拷贝数变异的方法、设备和介质 |
| CN117265069A (zh) * | 2023-09-21 | 2023-12-22 | 北京安智因生物技术有限公司 | 基于半导体测序平台检测brca1/2基因拷贝数变异 |
| CN117334249A (zh) * | 2023-05-30 | 2024-01-02 | 上海品峰医疗科技有限公司 | 基于扩增子测序数据检测拷贝数变异的方法、设备和介质 |
| CN120748484A (zh) * | 2025-08-15 | 2025-10-03 | 广州燃石医学检验所有限公司 | 基因拷贝数变异类型检测方法及相关产品 |
Citations (11)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20130316915A1 (en) * | 2010-10-13 | 2013-11-28 | Aaron Halpern | Methods for determining absolute genome-wide copy number variations of complex tumors |
| CN104133914A (zh) * | 2014-08-12 | 2014-11-05 | 厦门万基生物科技有限公司 | 一种消除高通量测序引入的gc偏差及对染色体拷贝数变异的检测方法 |
| US20160300013A1 (en) * | 2015-04-10 | 2016-10-13 | Agilent Technologies, Inc. | METHOD FOR SIMULTANEOUS DETECTION OF GENOME-WIDE COPY NUMBER CHANGES, cnLOH, INDELS, AND GENE MUTATIONS |
| CN106650312A (zh) * | 2016-12-29 | 2017-05-10 | 安诺优达基因科技(北京)有限公司 | 一种用于循环肿瘤dna拷贝数变异检测的装置 |
| CN106715711A (zh) * | 2014-07-04 | 2017-05-24 | 深圳华大基因股份有限公司 | 确定探针序列的方法和基因组结构变异的检测方法 |
| CN106951737A (zh) * | 2016-11-18 | 2017-07-14 | 南方医科大学 | 一种检测流产组织dna拷贝数变异和嵌合体的方法 |
| CN108256292A (zh) * | 2016-12-29 | 2018-07-06 | 安诺优达基因科技(北京)有限公司 | 一种拷贝数变异检测装置 |
| CN108427864A (zh) * | 2018-02-14 | 2018-08-21 | 南京世和基因生物技术有限公司 | 一种拷贝数变异的检测方法、装置以及计算机可读介质 |
| CN111968701A (zh) * | 2020-08-27 | 2020-11-20 | 北京吉因加科技有限公司 | 检测指定基因组区域体细胞拷贝数变异的方法和装置 |
| CN113674803A (zh) * | 2021-08-30 | 2021-11-19 | 广州燃石医学检验所有限公司 | 一种拷贝数变异的检测方法及其应用 |
| CN113789371A (zh) * | 2021-09-17 | 2021-12-14 | 广州燃石医学检验所有限公司 | 一种基于批次矫正的拷贝数变异的检测方法 |
Family Cites Families (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20140370504A1 (en) | 2011-12-31 | 2014-12-18 | Bgi Diagnosis Co., Ltd. | Method for detecting genetic variation |
| EP2826865B8 (en) | 2012-01-20 | 2017-08-16 | BGI Genomics Co., Ltd. | Method and system for determining whether copy number variation exists in sample genome, and computer readable medium |
| US10395759B2 (en) | 2015-05-18 | 2019-08-27 | Regeneron Pharmaceuticals, Inc. | Methods and systems for copy number variant detection |
-
2022
- 2022-08-29 WO PCT/CN2022/115447 patent/WO2023030233A1/zh not_active Ceased
- 2022-08-29 JP JP2024514011A patent/JP7745090B2/ja active Active
- 2022-08-29 EP EP22863369.9A patent/EP4397773A4/en active Pending
Patent Citations (11)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20130316915A1 (en) * | 2010-10-13 | 2013-11-28 | Aaron Halpern | Methods for determining absolute genome-wide copy number variations of complex tumors |
| CN106715711A (zh) * | 2014-07-04 | 2017-05-24 | 深圳华大基因股份有限公司 | 确定探针序列的方法和基因组结构变异的检测方法 |
| CN104133914A (zh) * | 2014-08-12 | 2014-11-05 | 厦门万基生物科技有限公司 | 一种消除高通量测序引入的gc偏差及对染色体拷贝数变异的检测方法 |
| US20160300013A1 (en) * | 2015-04-10 | 2016-10-13 | Agilent Technologies, Inc. | METHOD FOR SIMULTANEOUS DETECTION OF GENOME-WIDE COPY NUMBER CHANGES, cnLOH, INDELS, AND GENE MUTATIONS |
| CN106951737A (zh) * | 2016-11-18 | 2017-07-14 | 南方医科大学 | 一种检测流产组织dna拷贝数变异和嵌合体的方法 |
| CN106650312A (zh) * | 2016-12-29 | 2017-05-10 | 安诺优达基因科技(北京)有限公司 | 一种用于循环肿瘤dna拷贝数变异检测的装置 |
| CN108256292A (zh) * | 2016-12-29 | 2018-07-06 | 安诺优达基因科技(北京)有限公司 | 一种拷贝数变异检测装置 |
| CN108427864A (zh) * | 2018-02-14 | 2018-08-21 | 南京世和基因生物技术有限公司 | 一种拷贝数变异的检测方法、装置以及计算机可读介质 |
| CN111968701A (zh) * | 2020-08-27 | 2020-11-20 | 北京吉因加科技有限公司 | 检测指定基因组区域体细胞拷贝数变异的方法和装置 |
| CN113674803A (zh) * | 2021-08-30 | 2021-11-19 | 广州燃石医学检验所有限公司 | 一种拷贝数变异的检测方法及其应用 |
| CN113789371A (zh) * | 2021-09-17 | 2021-12-14 | 广州燃石医学检验所有限公司 | 一种基于批次矫正的拷贝数变异的检测方法 |
Non-Patent Citations (1)
| Title |
|---|
| See also references of EP4397773A4 * |
Cited By (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN117334249A (zh) * | 2023-05-30 | 2024-01-02 | 上海品峰医疗科技有限公司 | 基于扩增子测序数据检测拷贝数变异的方法、设备和介质 |
| CN117265069A (zh) * | 2023-09-21 | 2023-12-22 | 北京安智因生物技术有限公司 | 基于半导体测序平台检测brca1/2基因拷贝数变异 |
| CN117265069B (zh) * | 2023-09-21 | 2024-05-14 | 北京安智因生物技术有限公司 | 基于半导体测序平台检测brca1/2基因拷贝数变异 |
| CN117153249A (zh) * | 2023-10-26 | 2023-12-01 | 北京华宇亿康生物工程技术有限公司 | 用于检测smn基因拷贝数变异的方法、设备和介质 |
| CN117153249B (zh) * | 2023-10-26 | 2024-02-02 | 北京华宇亿康生物工程技术有限公司 | 用于检测smn基因拷贝数变异的方法、设备和介质 |
| CN120748484A (zh) * | 2025-08-15 | 2025-10-03 | 广州燃石医学检验所有限公司 | 基因拷贝数变异类型检测方法及相关产品 |
Also Published As
| Publication number | Publication date |
|---|---|
| JP2024532497A (ja) | 2024-09-05 |
| JP7745090B2 (ja) | 2025-09-26 |
| EP4397773A4 (en) | 2025-09-10 |
| EP4397773A1 (en) | 2024-07-10 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN113674803B (zh) | 一种拷贝数变异的检测方法、装置、存储介质及其应用 | |
| CN109880910B (zh) | 一种肿瘤突变负荷的检测位点组合、检测方法、检测试剂盒及系统 | |
| EP4397773A1 (en) | Copy number variation detection method and application thereof | |
| US20200203014A1 (en) | Methods and systems for sequencing-based variant detection | |
| CN109427412B (zh) | 用于检测肿瘤突变负荷的序列组合和其设计方法 | |
| CN111321140A (zh) | 一种基于单样本的肿瘤突变负荷检测方法和装置 | |
| US12603150B2 (en) | Calculating cell-type RNA profiles for diagnosis and treatment | |
| WO2019157791A1 (zh) | 一种拷贝数变异的检测方法、装置以及计算机可读介质 | |
| JP2022169566A (ja) | まれな変異およびコピー数多型を検出するためのシステムおよび方法 | |
| US20220399080A1 (en) | Methods and products for minimal residual disease detection | |
| CN113249483B (zh) | 一种检测肿瘤突变负荷的基因组合、系统及应用 | |
| US20220072553A1 (en) | Device and method for detecting tumor mutation burden (tmb) based on capture sequencing | |
| CN116940987A (zh) | 用于确定变体频率和监测疾病进展的方法 | |
| US12494290B2 (en) | Noise measure for copy number analysis on targeted panel sequencing data | |
| Tang et al. | Tumor mutation burden derived from small next generation sequencing targeted gene panel as an initial screening method | |
| US20230057154A1 (en) | Somatic variant cooccurrence with abnormally methylated fragments | |
| EP4600963A1 (en) | Methods and systems for determining blood tumor mutational burden in a liquid biopsy assay | |
| KR20180112727A (ko) | 생물학적 시료의 핵산 품질을 결정하는 방법 | |
| CN118979107A (zh) | 腹腔灌洗液循环肿瘤细胞与循环肿瘤dna在预测胃癌根治术后异时性腹膜转移中的应用 | |
| KR102491485B1 (ko) | 순환 종양 핵산의 복제수 변이 분석 방법 | |
| US20250095775A1 (en) | Methods for determining variant frequency and monitoring disease progression | |
| JP2025507673A (ja) | 液体生検アッセイのためのプローブセット | |
| HK40057523B (zh) | 一种拷贝数变异的检测方法、装置、存储介质及其应用 | |
| EP4381512A1 (en) | Somatic variant cooccurrence with abnormally methylated fragments | |
| CN114908163A (zh) | 预测肺癌免疫检查点抑制剂疗效的标志物及其应用 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 22863369 Country of ref document: EP Kind code of ref document: A1 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2024514011 Country of ref document: JP |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2022863369 Country of ref document: EP |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| ENP | Entry into the national phase |
Ref document number: 2022863369 Country of ref document: EP Effective date: 20240402 |















