EP2332082A1 - Procédé de caractérisation de séquences à partir d'échantillons de matériaux génétiques - Google Patents
Procédé de caractérisation de séquences à partir d'échantillons de matériaux génétiquesInfo
- Publication number
- EP2332082A1 EP2332082A1 EP09790738A EP09790738A EP2332082A1 EP 2332082 A1 EP2332082 A1 EP 2332082A1 EP 09790738 A EP09790738 A EP 09790738A EP 09790738 A EP09790738 A EP 09790738A EP 2332082 A1 EP2332082 A1 EP 2332082A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- genetic material
- snp
- sample
- int
- snps
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Withdrawn
Links
- 108090000623 proteins and genes Proteins 0.000 title claims abstract description 305
- 102000004169 proteins and genes Human genes 0.000 title claims abstract description 304
- 238000000034 method Methods 0.000 title claims abstract description 234
- 239000000203 mixture Substances 0.000 claims abstract description 346
- 239000002773 nucleotide Substances 0.000 claims abstract description 17
- 125000003729 nucleotide group Chemical group 0.000 claims abstract description 17
- 239000000523 sample Substances 0.000 claims description 398
- 238000012360 testing method Methods 0.000 claims description 179
- 108700028369 Alleles Proteins 0.000 claims description 176
- 238000004458 analytical method Methods 0.000 claims description 75
- 230000008569 process Effects 0.000 claims description 47
- 239000013074 reference sample Substances 0.000 claims description 26
- 102000054765 polymorphisms of proteins Human genes 0.000 claims description 10
- 238000011109 contamination Methods 0.000 claims description 3
- 230000001580 bacterial effect Effects 0.000 claims description 2
- 238000012512 characterization method Methods 0.000 claims 8
- 239000013308 plastic optical fiber Substances 0.000 claims 2
- 239000013309 porous organic framework Substances 0.000 claims 2
- 238000012544 monitoring process Methods 0.000 claims 1
- 238000002493 microarray Methods 0.000 abstract description 30
- 238000003205 genotyping method Methods 0.000 abstract description 29
- 108020004414 DNA Proteins 0.000 description 95
- 210000004027 cell Anatomy 0.000 description 65
- 238000009826 distribution Methods 0.000 description 62
- 101150040459 RAS gene Proteins 0.000 description 39
- 101150076031 RAS1 gene Proteins 0.000 description 39
- 239000011800 void material Substances 0.000 description 35
- 241000282414 Homo sapiens Species 0.000 description 30
- 230000006870 function Effects 0.000 description 28
- 108020005196 Mitochondrial DNA Proteins 0.000 description 27
- 108091092878 Microsatellite Proteins 0.000 description 24
- 108020004707 nucleic acids Proteins 0.000 description 20
- 102000039446 nucleic acids Human genes 0.000 description 20
- 150000007523 nucleic acids Chemical class 0.000 description 20
- 238000012549 training Methods 0.000 description 19
- 238000013459 approach Methods 0.000 description 18
- 238000004088 simulation Methods 0.000 description 18
- 239000011324 bead Substances 0.000 description 17
- 230000002068 genetic effect Effects 0.000 description 17
- 206010028980 Neoplasm Diseases 0.000 description 14
- 210000001671 embryonic stem cell Anatomy 0.000 description 14
- 238000012545 processing Methods 0.000 description 11
- 238000003491 array Methods 0.000 description 10
- 201000011510 cancer Diseases 0.000 description 10
- 238000004590 computer program Methods 0.000 description 10
- 238000005259 measurement Methods 0.000 description 10
- 241000283690 Bos taurus Species 0.000 description 9
- 238000001574 biopsy Methods 0.000 description 9
- 239000000463 material Substances 0.000 description 9
- 238000003556 assay Methods 0.000 description 8
- 210000000349 chromosome Anatomy 0.000 description 8
- 230000003068 static effect Effects 0.000 description 8
- 210000001519 tissue Anatomy 0.000 description 8
- 238000004422 calculation algorithm Methods 0.000 description 7
- 230000001186 cumulative effect Effects 0.000 description 7
- 238000004374 forensic analysis Methods 0.000 description 7
- 108091033319 polynucleotide Proteins 0.000 description 7
- 102000040430 polynucleotide Human genes 0.000 description 7
- 239000002157 polynucleotide Substances 0.000 description 7
- 238000004364 calculation method Methods 0.000 description 6
- 238000002474 experimental method Methods 0.000 description 6
- 239000000284 extract Substances 0.000 description 6
- DWCZIOOZPIDHAB-UHFFFAOYSA-L methyl green Chemical compound [Cl-].[Cl-].C1=CC(N(C)C)=CC=C1C(C=1C=CC(=CC=1)[N+](C)(C)C)=C1C=CC(=[N+](C)C)C=C1 DWCZIOOZPIDHAB-UHFFFAOYSA-L 0.000 description 6
- 238000010606 normalization Methods 0.000 description 6
- 238000003860 storage Methods 0.000 description 6
- 108091032973 (ribonucleotides)n+m Proteins 0.000 description 5
- 101001024425 Mus musculus Ig gamma-2A chain C region secreted form Proteins 0.000 description 5
- 208000037265 diseases, disorders, signs and symptoms Diseases 0.000 description 5
- 238000009396 hybridization Methods 0.000 description 5
- 230000003211 malignant effect Effects 0.000 description 5
- 230000002441 reversible effect Effects 0.000 description 5
- 238000013476 bayesian approach Methods 0.000 description 4
- 238000006243 chemical reaction Methods 0.000 description 4
- 230000002596 correlated effect Effects 0.000 description 4
- 238000013461 design Methods 0.000 description 4
- 238000000605 extraction Methods 0.000 description 4
- 238000012775 microarray technology Methods 0.000 description 4
- 230000002438 mitochondrial effect Effects 0.000 description 4
- 238000005192 partition Methods 0.000 description 4
- 238000011176 pooling Methods 0.000 description 4
- 238000011160 research Methods 0.000 description 4
- 230000000717 retained effect Effects 0.000 description 4
- 238000000926 separation method Methods 0.000 description 4
- 101100316841 Escherichia phage lambda bet gene Proteins 0.000 description 3
- 102000010029 Homer Scaffolding Proteins Human genes 0.000 description 3
- 108091034117 Oligonucleotide Proteins 0.000 description 3
- 210000002593 Y chromosome Anatomy 0.000 description 3
- 230000001174 ascending effect Effects 0.000 description 3
- 230000008901 benefit Effects 0.000 description 3
- 230000008859 change Effects 0.000 description 3
- 238000003501 co-culture Methods 0.000 description 3
- 208000035475 disorder Diseases 0.000 description 3
- 230000000694 effects Effects 0.000 description 3
- 238000001962 electrophoresis Methods 0.000 description 3
- 239000004615 ingredient Substances 0.000 description 3
- 108020004999 messenger RNA Proteins 0.000 description 3
- 238000012986 modification Methods 0.000 description 3
- 230000004048 modification Effects 0.000 description 3
- 238000002360 preparation method Methods 0.000 description 3
- 238000005070 sampling Methods 0.000 description 3
- 230000009466 transformation Effects 0.000 description 3
- 208000026817 47,XYY syndrome Diseases 0.000 description 2
- 206010005949 Bone cancer Diseases 0.000 description 2
- 208000018084 Bone neoplasm Diseases 0.000 description 2
- 102000004190 Enzymes Human genes 0.000 description 2
- 108090000790 Enzymes Proteins 0.000 description 2
- 241000282412 Homo Species 0.000 description 2
- 241000124008 Mammalia Species 0.000 description 2
- 108091093037 Peptide nucleic acid Proteins 0.000 description 2
- 108091028664 Ribonucleotide Proteins 0.000 description 2
- IQFYYKKMVGJFEH-XLPZGREQSA-N Thymidine Chemical compound O=C1NC(=O)C(C)=CN1[C@@H]1O[C@H](CO)[C@@H](O)C1 IQFYYKKMVGJFEH-XLPZGREQSA-N 0.000 description 2
- 241000700605 Viruses Species 0.000 description 2
- 230000009471 action Effects 0.000 description 2
- OIRDTQYFTABQOQ-KQYNXXCUSA-N adenosine Chemical compound C1=NC=2C(N)=NC=NC=2N1[C@@H]1O[C@H](CO)[C@@H](O)[C@H]1O OIRDTQYFTABQOQ-KQYNXXCUSA-N 0.000 description 2
- 230000003321 amplification Effects 0.000 description 2
- 238000013398 bayesian method Methods 0.000 description 2
- 150000001875 compounds Chemical class 0.000 description 2
- 238000010276 construction Methods 0.000 description 2
- 238000012937 correction Methods 0.000 description 2
- 230000000875 corresponding effect Effects 0.000 description 2
- 238000012258 culturing Methods 0.000 description 2
- 238000013075 data extraction Methods 0.000 description 2
- 239000005547 deoxyribonucleotide Substances 0.000 description 2
- 201000010099 disease Diseases 0.000 description 2
- 230000000670 limiting effect Effects 0.000 description 2
- 108091070501 miRNA Proteins 0.000 description 2
- 239000002679 microRNA Substances 0.000 description 2
- 238000003199 nucleic acid amplification method Methods 0.000 description 2
- 238000012898 one-sample t-test Methods 0.000 description 2
- 238000007639 printing Methods 0.000 description 2
- 238000000746 purification Methods 0.000 description 2
- 239000002336 ribonucleotide Substances 0.000 description 2
- 125000002652 ribonucleotide group Chemical group 0.000 description 2
- 238000012216 screening Methods 0.000 description 2
- 238000012163 sequencing technique Methods 0.000 description 2
- 238000003201 single nucleotide polymorphism genotyping Methods 0.000 description 2
- 238000010200 validation analysis Methods 0.000 description 2
- YKBGVTZYEHREMT-KVQBGUIXSA-N 2'-deoxyguanosine Chemical compound C1=NC=2C(=O)NC(N)=NC=2N1[C@H]1C[C@H](O)[C@@H](CO)O1 YKBGVTZYEHREMT-KVQBGUIXSA-N 0.000 description 1
- CKTSBUTUHBMZGZ-ULQXZJNLSA-N 4-amino-1-[(2r,4s,5r)-4-hydroxy-5-(hydroxymethyl)oxolan-2-yl]-5-tritiopyrimidin-2-one Chemical compound O=C1N=C(N)C([3H])=CN1[C@@H]1O[C@H](CO)[C@@H](O)C1 CKTSBUTUHBMZGZ-ULQXZJNLSA-N 0.000 description 1
- 238000012935 Averaging Methods 0.000 description 1
- 241000271566 Aves Species 0.000 description 1
- 241000894006 Bacteria Species 0.000 description 1
- DWRXFEITVBNRMK-UHFFFAOYSA-N Beta-D-1-Arabinofuranosylthymine Natural products O=C1NC(=O)C(C)=CN1C1C(O)C(O)C(CO)O1 DWRXFEITVBNRMK-UHFFFAOYSA-N 0.000 description 1
- 208000005243 Chondrosarcoma Diseases 0.000 description 1
- 241000938605 Crocodylia Species 0.000 description 1
- 241000195493 Cryptophyta Species 0.000 description 1
- 102000053602 DNA Human genes 0.000 description 1
- 101100012466 Drosophila melanogaster Sras gene Proteins 0.000 description 1
- 241000196324 Embryophyta Species 0.000 description 1
- 101100126165 Escherichia coli (strain K12) intA gene Proteins 0.000 description 1
- 208000006168 Ewing Sarcoma Diseases 0.000 description 1
- 241001465754 Metazoa Species 0.000 description 1
- 241000237852 Mollusca Species 0.000 description 1
- 208000034578 Multiple myelomas Diseases 0.000 description 1
- 108091007491 NSP3 Papain-like protease domains Proteins 0.000 description 1
- 238000012408 PCR amplification Methods 0.000 description 1
- 238000010222 PCR analysis Methods 0.000 description 1
- 206010034016 Paronychia Diseases 0.000 description 1
- ZYFVNVRFVHJEIU-UHFFFAOYSA-N PicoGreen Chemical compound CN(C)CCCN(CCCN(C)C)C1=CC(=CC2=[N+](C3=CC=CC=C3S2)C)C2=CC=CC=C2N1C1=CC=CC=C1 ZYFVNVRFVHJEIU-UHFFFAOYSA-N 0.000 description 1
- 206010035226 Plasma cell myeloma Diseases 0.000 description 1
- 201000010273 Porphyria Cutanea Tarda Diseases 0.000 description 1
- 241000282860 Procaviidae Species 0.000 description 1
- 229940124878 RotaTeq Drugs 0.000 description 1
- 101000767160 Saccharomyces cerevisiae (strain ATCC 204508 / S288c) Intracellular protein transport protein USO1 Proteins 0.000 description 1
- 108020004459 Small interfering RNA Proteins 0.000 description 1
- 238000000692 Student's t-test Methods 0.000 description 1
- 108020005202 Viral DNA Proteins 0.000 description 1
- JLCPHMBAVCMARE-UHFFFAOYSA-N [3-[[3-[[3-[[3-[[3-[[3-[[3-[[3-[[3-[[3-[[3-[[5-(2-amino-6-oxo-1H-purin-9-yl)-3-[[3-[[3-[[3-[[3-[[3-[[5-(2-amino-6-oxo-1H-purin-9-yl)-3-[[5-(2-amino-6-oxo-1H-purin-9-yl)-3-hydroxyoxolan-2-yl]methoxy-hydroxyphosphoryl]oxyoxolan-2-yl]methoxy-hydroxyphosphoryl]oxy-5-(5-methyl-2,4-dioxopyrimidin-1-yl)oxolan-2-yl]methoxy-hydroxyphosphoryl]oxy-5-(6-aminopurin-9-yl)oxolan-2-yl]methoxy-hydroxyphosphoryl]oxy-5-(6-aminopurin-9-yl)oxolan-2-yl]methoxy-hydroxyphosphoryl]oxy-5-(6-aminopurin-9-yl)oxolan-2-yl]methoxy-hydroxyphosphoryl]oxy-5-(6-aminopurin-9-yl)oxolan-2-yl]methoxy-hydroxyphosphoryl]oxyoxolan-2-yl]methoxy-hydroxyphosphoryl]oxy-5-(5-methyl-2,4-dioxopyrimidin-1-yl)oxolan-2-yl]methoxy-hydroxyphosphoryl]oxy-5-(4-amino-2-oxopyrimidin-1-yl)oxolan-2-yl]methoxy-hydroxyphosphoryl]oxy-5-(5-methyl-2,4-dioxopyrimidin-1-yl)oxolan-2-yl]methoxy-hydroxyphosphoryl]oxy-5-(5-methyl-2,4-dioxopyrimidin-1-yl)oxolan-2-yl]methoxy-hydroxyphosphoryl]oxy-5-(6-aminopurin-9-yl)oxolan-2-yl]methoxy-hydroxyphosphoryl]oxy-5-(6-aminopurin-9-yl)oxolan-2-yl]methoxy-hydroxyphosphoryl]oxy-5-(4-amino-2-oxopyrimidin-1-yl)oxolan-2-yl]methoxy-hydroxyphosphoryl]oxy-5-(4-amino-2-oxopyrimidin-1-yl)oxolan-2-yl]methoxy-hydroxyphosphoryl]oxy-5-(4-amino-2-oxopyrimidin-1-yl)oxolan-2-yl]methoxy-hydroxyphosphoryl]oxy-5-(6-aminopurin-9-yl)oxolan-2-yl]methoxy-hydroxyphosphoryl]oxy-5-(4-amino-2-oxopyrimidin-1-yl)oxolan-2-yl]methyl [5-(6-aminopurin-9-yl)-2-(hydroxymethyl)oxolan-3-yl] hydrogen phosphate Polymers Cc1cn(C2CC(OP(O)(=O)OCC3OC(CC3OP(O)(=O)OCC3OC(CC3O)n3cnc4c3nc(N)[nH]c4=O)n3cnc4c3nc(N)[nH]c4=O)C(COP(O)(=O)OC3CC(OC3COP(O)(=O)OC3CC(OC3COP(O)(=O)OC3CC(OC3COP(O)(=O)OC3CC(OC3COP(O)(=O)OC3CC(OC3COP(O)(=O)OC3CC(OC3COP(O)(=O)OC3CC(OC3COP(O)(=O)OC3CC(OC3COP(O)(=O)OC3CC(OC3COP(O)(=O)OC3CC(OC3COP(O)(=O)OC3CC(OC3COP(O)(=O)OC3CC(OC3COP(O)(=O)OC3CC(OC3COP(O)(=O)OC3CC(OC3COP(O)(=O)OC3CC(OC3COP(O)(=O)OC3CC(OC3COP(O)(=O)OC3CC(OC3CO)n3cnc4c(N)ncnc34)n3ccc(N)nc3=O)n3cnc4c(N)ncnc34)n3ccc(N)nc3=O)n3ccc(N)nc3=O)n3ccc(N)nc3=O)n3cnc4c(N)ncnc34)n3cnc4c(N)ncnc34)n3cc(C)c(=O)[nH]c3=O)n3cc(C)c(=O)[nH]c3=O)n3ccc(N)nc3=O)n3cc(C)c(=O)[nH]c3=O)n3cnc4c3nc(N)[nH]c4=O)n3cnc4c(N)ncnc34)n3cnc4c(N)ncnc34)n3cnc4c(N)ncnc34)n3cnc4c(N)ncnc34)O2)c(=O)[nH]c1=O JLCPHMBAVCMARE-UHFFFAOYSA-N 0.000 description 1
- 238000003149 assay kit Methods 0.000 description 1
- 230000036621 balding Effects 0.000 description 1
- IQFYYKKMVGJFEH-UHFFFAOYSA-N beta-L-thymidine Natural products O=C1NC(=O)C(C)=CN1C1OC(CO)C(O)C1 IQFYYKKMVGJFEH-UHFFFAOYSA-N 0.000 description 1
- 201000007327 bone benign neoplasm Diseases 0.000 description 1
- -1 but not limited to Chemical class 0.000 description 1
- 238000005266 casting Methods 0.000 description 1
- 238000004140 cleaning Methods 0.000 description 1
- 238000004891 communication Methods 0.000 description 1
- 239000002131 composite material Substances 0.000 description 1
- 238000000205 computational method Methods 0.000 description 1
- 230000001143 conditioned effect Effects 0.000 description 1
- 238000007796 conventional method Methods 0.000 description 1
- 238000013500 data storage Methods 0.000 description 1
- 238000013501 data transformation Methods 0.000 description 1
- 230000003247 decreasing effect Effects 0.000 description 1
- 125000002637 deoxyribonucleotide group Chemical group 0.000 description 1
- 230000001419 dependent effect Effects 0.000 description 1
- 238000001514 detection method Methods 0.000 description 1
- 238000011161 development Methods 0.000 description 1
- 235000016693 dipotassium tartrate Nutrition 0.000 description 1
- 238000005516 engineering process Methods 0.000 description 1
- 230000007613 environmental effect Effects 0.000 description 1
- 238000006911 enzymatic reaction Methods 0.000 description 1
- CNOILWYCLUVYET-UHFFFAOYSA-N ethyl n-[4-[benzyl(2-phenylethyl)amino]-2-(3-methylphenyl)-1h-imidazo[4,5-c]pyridin-6-yl]carbamate Chemical compound N=1C(NC(=O)OCC)=CC=2NC(C=3C=C(C)C=CC=3)=NC=2C=1N(CC=1C=CC=CC=1)CCC1=CC=CC=C1 CNOILWYCLUVYET-UHFFFAOYSA-N 0.000 description 1
- 238000011156 evaluation Methods 0.000 description 1
- 230000007717 exclusion Effects 0.000 description 1
- 210000003754 fetus Anatomy 0.000 description 1
- 238000007429 general method Methods 0.000 description 1
- 102000054766 genetic haplotypes Human genes 0.000 description 1
- 230000007614 genetic variation Effects 0.000 description 1
- 230000036541 health Effects 0.000 description 1
- 210000003494 hepatocyte Anatomy 0.000 description 1
- 239000012535 impurity Substances 0.000 description 1
- 238000010348 incorporation Methods 0.000 description 1
- 238000011835 investigation Methods 0.000 description 1
- 150000002500 ions Chemical class 0.000 description 1
- 238000004519 manufacturing process Methods 0.000 description 1
- 230000001404 mediated effect Effects 0.000 description 1
- 210000004925 microvascular endothelial cell Anatomy 0.000 description 1
- 210000003470 mitochondria Anatomy 0.000 description 1
- 238000010369 molecular cloning Methods 0.000 description 1
- 238000001668 nucleic acid synthesis Methods 0.000 description 1
- 238000002515 oligonucleotide synthesis Methods 0.000 description 1
- 201000008968 osteosarcoma Diseases 0.000 description 1
- NUSQOFAKCBLANB-UHFFFAOYSA-N phthalocyanine tetrasulfonic acid Chemical compound C12=CC(S(=O)(=O)O)=CC=C2C(N=C2NC(C3=CC=C(C=C32)S(O)(=O)=O)=N2)=NC1=NC([C]1C=CC(=CC1=1)S(O)(=O)=O)=NC=1N=C1[C]3C=CC(S(O)(=O)=O)=CC3=C2N1 NUSQOFAKCBLANB-UHFFFAOYSA-N 0.000 description 1
- 229920000642 polymer Polymers 0.000 description 1
- AVTYONGGKAJVTE-OLXYHTOASA-L potassium L-tartrate Chemical compound [K+].[K+].[O-]C(=O)[C@H](O)[C@@H](O)C([O-])=O AVTYONGGKAJVTE-OLXYHTOASA-L 0.000 description 1
- 230000035935 pregnancy Effects 0.000 description 1
- 230000009467 reduction Effects 0.000 description 1
- 230000002829 reductive effect Effects 0.000 description 1
- 230000004044 response Effects 0.000 description 1
- 238000007894 restriction fragment length polymorphism technique Methods 0.000 description 1
- 239000004065 semiconductor Substances 0.000 description 1
- 238000001629 sign test Methods 0.000 description 1
- 238000000638 solvent extraction Methods 0.000 description 1
- 238000010561 standard procedure Methods 0.000 description 1
- 238000007619 statistical method Methods 0.000 description 1
- 239000000126 substance Substances 0.000 description 1
- 238000006467 substitution reaction Methods 0.000 description 1
- 230000009897 systematic effect Effects 0.000 description 1
- 238000012353 t test Methods 0.000 description 1
- 229940104230 thymidine Drugs 0.000 description 1
- 230000001131 transforming effect Effects 0.000 description 1
- 125000005208 trialkylammonium group Chemical group 0.000 description 1
- 210000004881 tumor cell Anatomy 0.000 description 1
- 238000009827 uniform distribution Methods 0.000 description 1
- 210000000689 upper leg Anatomy 0.000 description 1
- UPPMZCXMQRVMME-UHFFFAOYSA-N valethamate Chemical compound CC[N+](C)(CC)CCOC(=O)C(C(C)CC)C1=CC=CC=C1 UPPMZCXMQRVMME-UHFFFAOYSA-N 0.000 description 1
- 230000035899 viability Effects 0.000 description 1
- 230000003612 virological effect Effects 0.000 description 1
- 230000000007 visual effect Effects 0.000 description 1
Classifications
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B20/00—ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B20/00—ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
- G16B20/20—Allele or variant detection, e.g. single nucleotide polymorphism [SNP] detection
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B25/00—ICT specially adapted for hybridisation; ICT specially adapted for gene or protein expression
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B40/00—ICT specially adapted for biostatistics; ICT specially adapted for bioinformatics-related machine learning or data mining, e.g. knowledge discovery or pattern finding
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B20/00—ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
- G16B20/40—Population genetics; Linkage disequilibrium
Definitions
- the present disclosure relates to systems and methods for using multiple single nucleotide polymorphisms (SNPs) for characterizing genetic material in a sample.
- SNPs single nucleotide polymorphisms
- determining whether an individual's genetic material is present within a complex mixture containing genetic material (such as DNA) from numerous individuals is of interest to multiple fields. For example, within forensics, determining whether a person contributed their genetic material to a mixture is typically a skilled process. In large part, forensically identifying whether a person is contributing less than 10% of the total genomic DNA to a mixture is not easily done, is difficult to automate, and is highly confounded with the inclusion of more individuals.
- STR short tandem repeats
- mtDNA has weaknesses, including the uniparental mode of inheritance and lower discrimination power that can be moderately mediated by using the whole mitochondrial genome or known surrounding single nucleotide polymorphisms (SNPs) (See Coble, M.D. et al. Single nucleotide polymorphisms over the entire mtDNA genome that increase the power of forensic testing in Caucasians. Int J Legal Med 118, 137-146 (2004) and Parsons, TJ. & Coble, M.D. Increasing the forensic discrimination of mitochondrial DNA testing through analysis of the entire mitochondrial DNA genome. Croat Med J 42, 304-309 (2001)).
- SNPs single nucleotide polymorphisms
- Informative SNPs have been used to help resolve problems with using mtDNA (See Coble, M.D. et al. Single nucleotide polymorphisms over the entire mtDNA genome that increase the power of forensic testing in Caucasians. Int J Legal Med 118, 137-146 (2004); Just, R.S. et al. Toward increased utility of mtDNA in forensic identifications. Forensic Sci Int 146 Suppl, S147-149 (2004); and Vallone, P.M., Just, R.S., Coble, M.D., Butler, J.M. & Parsons, T.J. A multiplex allele-specific primer extension assay for forensically informative SNPs distributed throughout the mitochondrial genome. Int J Legal Med 118, 147-157 (2004)) but have not been used wholly or separately as the discriminatory factor, or on the same scale as provided herein.
- Some of the present embodiments provide a variety of methods (and apparatuses for implementing these methods), for determining if a subject's genetic material is present in a genetic material sample (a "test genetic material sample). While there are a variety of techniques by which this can be achieved, in some embodiments, this is achieved by determining if there is a bias and/or direction of an allele occurrence and/or frequency within a collection of single nucleotide polymorphisms (SNPs) of the test genetic material sample relative to a reference and/or the subject's SNP signature or collection of SNPs genotypes.
- SNPs single nucleotide polymorphisms
- a system for determining if a subject contributed genetic material to a sample can comprise an input module configured to allow the input of one or more of a sample SNP signature, a reference SNP signature, and a subject SNP signature; a module configured to determine a bias of an allele frequency within SNPs of the sample SNP signature relative to the reference SNP signature and the subject SNP signature; and a module configured to output the bias, wherein one or more of the modules is executed on a computing device.
- a method for determining if a person of interest contributed genetic material to a test genetic material sample is provided.
- the method can comprise determining a bias of an allele frequency within SNPs of the test genetic material sample relative to a reference and a subject's SNP signature.
- a method of characterizing a test genetic material sample to determine if a person of interest's (“PCTs") genetic material is within the test genetic material sample is provided.
- the method can comprise providing a SNP analysis of the test genetic material sample; providing a SNP analysis of a reference genetic material sample; providing a SNP analysis of a POFs genetic material; in a first comparison, comparing the SNP analysis of the test genetic material sample to the SNP analysis of the POFs genetic material; in a second comparison, comparing the SNP analysis of the reference genetic material to the SNP analysis of the POFs genetic material; and comparing the first and second comparisons, thereby determining if the POFs genetic material is likely in the test genetic material sample.
- a method of characterizing a test genetic material sample can comprise providing a first allele frequency for a SNP for a person of interest (POI); providing a second allele frequency for the SNP from a reference population(s) of genetic material; providing a third allele frequency for the SNP for the test genetic material sample; repeating the above processes for at least 10 different SNPs; and analyzing the first, second, and third allele frequencies to characterize the test genetic material sample.
- POI person of interest
- a method for determining a likelihood that a subject contributed genetic material to a test genetic material sample can comprise providing a test genetic material sample; performing a single nucleotide polymorphism analysis on the test genetic material sample, whereby at least 50 different single nucleotide polymorphisms in said test genetic material sample are analyzed, thereby creating a sample SNP signature; and comparing the sample SNP signature to a subject's SNP signature to determine a likelihood that the subject contributed genetic material to a test genetic material sample.
- the invention relates generally to single nucleotide polymorphism genotyping and more specifically to single nucleotide polymorphism genotyping of samples from multiple individuals and/or sources.
- the method comprises a sample SNP signature that is from a biopsy from a subject, wherein the biopsy from the subject is to be tested for the presence of a cancer.
- the sample SNP signature is created from a female who wants to determine if she is pregnant.
- the subject's SNP signature is a viral DNA signature.
- FIG. IA To give insight into the intuition behind come embodiments of the various methods, three different scenarios are presented per SNP of the possible allele frequency of the person of interest corresponding to the genotypes AA, AB, and BB.
- the allele frequencies of the reference population, person of interest (subject), and the mixture are described as M 1 (test genetic material sample), Y, (subject), and Pop, (reference population) respectively.
- M 1 test genetic material sample
- Y, (subject) reference population
- Pop reference population
- the distance measure is greater (and positive) when the Y 1 of the person of interest is closer to the M, of the mixture than to the Pop, of the reference population.
- the distance measure is smaller (and negative) when the Y 1 of the person of interest is closer to the Pop, of the reference population than to M, of the mixture
- the test statistic is then the z-score using this distance measure.
- FIG. IB is a flow chart depicting various possible processes involved in some embodiments described herein.
- FIGS. 2A - 2C depict various simulation results: Using 1423 Wellcome Trust 58C individuals, log scaled p-values were given from simulations based off of three variables: the number of SNPs (V), the fraction of the individual in the mixture (/), and the probe variance (y p y The graphs plot the relationships between the three variables with a different variable fixed in each graph. The log scaled p-values are represented by the shading of each point in the graph, as well as the z-axis on the right graphs. These simulations indicate that one can resolve mixtures where a given individual is 0.1% of the mixture if), probe variance is at most 0.01 (v p ) and the number of SNPs probed is 50,000 (s).
- FIGS. 3A - 3D provide the results from a series of experiments. Experimental validation using a series of mixtures (see Table 1, A-F) assayed on the Affymetrix GeneChip 5.0, Illumina BeadArray 550 and the Illumina 450S Duo Human BeadChip.
- the x-axis shows each individual in the CEU HapMap population, the left y- axis shows the p-value (log scaled), and the right y-axis shows the value of the test statistic.
- mixtures A, B, E and F those in the mixture are shaded light and identified and those not in the mixture are shaded darker and identified.
- mixtures C and D those individuals who are not in the mixtures are shaded darkly and identified, those individuals who are related to the 1% or 10% individuals in the mixtures are shaded lighter and identified as "1-10", those individuals who are related to the 90% or 99% are shaded lighter still and identified as "90-99", and those people in the mixture are shaded lighter than those absent from the mixture and are identified.
- An arrow denotes identification of numerous (or a cluster) of data points while a line denotes identification of a specific data point. Unless otherwise specified, an unmarked data point is part of the closest denoted cluster.
- the present disclosure provides a variety of methods (and apparatuses for implementing these methods), for determining if a subject's genetic material is present in a genetic material sample (a "test genetic material sample). While there are a variety of techniques by which this can be achieved, in some embodiments, this is achieved by determining if there is a bias and/or direction of an allele occurrence and/or frequency within SNPs of the test genetic material sample relative to a reference and/or the subject's SNP signature (e.g., SNP genotype).
- SNP Single Nucleotide Polymorphism
- SNPs and high-density SNP genotyping arrays have been around for some time, their use has been predominately been developed as tools geneticists use to identify common genetic variants that predispose an individual to disease. Some embodiments disclosed herein allow for the use of SNPs to identify the presence or absence of one or more individuals' genetic material in a sample.
- the SNP based analysis can be used for analyzing forensic mixtures. SNPs are traditionally analyzed by genotype (e.g. AA, AT, or TT) and, prior to the present disclosure, were thought to be non-ideal in resolving mixtures.
- Exclusion probabilities give a calculation based on the probability of excluding a random individual (See Chakraborty, R., Meagher, T.R. & Smouse, P. E. Parentage analysis with genetic markers in natural populations. I. The expected proportion of offspring with unambiguous paternity. Genetics 118, 527-536 (1988)). Nevertheless, many of these methods rely on assuming the number of individuals in the mixture (See Egeland, T., Dalen, I. & Mostad, P.F. Estimating the number of contributors to a DNA profile. Int J Legal Med 117, 271-275 (2003)) and have been applied only to STR markers. In some embodiments, one need not know or estimate the number of individuals that contributed to a mixture when using the methods disclosed herein.
- Likelihood ratios are commonly used when testing which hypothesis is favored by the evidence or DNA samples (See Weir, B. S. et al. Interpreting DNA mixtures. J Forensic Sci 42, 213-222 (1997)).
- the proper prior odds ratio can then be given based on the current situation or context, and then would be combined with the likelihood ratio to give a posterior odd ratio.
- one can then use SNP microarrays to determine allele frequencies or allele counts.
- the Bayesian approach includes creation of explicit hypotheses, estimation of the total fraction of the individual of interest that contributes to the mixture, inclusion of multiple ancestral backgrounds across ancestrally informative SNPs, and inclusion of the possibility that related individuals are within the mixture.
- Enzymatic reactions and purification techniques are performed according to manufacturer's specifications or as commonly accomplished in the art or as described herein.
- the techniques and procedures described herein are generally performed according to conventional methods well known in the art and as described in various general and more specific references that are cited and discussed throughout the instant specification. See, e.g., Sambrook et ah, Molecular Cloning: A Laboratory Manual (Third ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. 2000).
- the nomenclatures utilized in connection with, and the laboratory procedures and techniques of described herein are those well known and commonly used in the art.
- genetic material refers to natural nucleic acids, artificial nucleic acids, non-natural nucleic acid, orthogonal nucleotides, analogs thereof, or combinations thereof. Genetic material can also include analogs of DNA or RNA having modifications to either the bases or the backbone. For example, genetic material, as used herein, includes the use of peptide nucleic acids (PNA). The term “genetic material” also includes chimeric molecules.
- the genetic material can include, consist, or consist essentially of a nucleic acid of one or more strands of single and/or double stranded material. Genetic material from a subject is generally (unless noted otherwise) numerous strands and numerous genes, and in some embodiments, can include the entire genome of the subject. In some embodiments, genetic material comprises, consists or consists essentially of nucleic acids.
- the genetic material is from a subject that someone wishes to determine the presence or absence of in a test genetic material sample.
- Exemplary genetic materials include DNA, RNA, mRNA, and miRNA.
- the genetic material and/or the test genetic material sample comprises, consists, or consists essentially of DNA, RNA, mRNA, miRNA, and any combination thereof.
- the genetic material is contained within the test genetic material sample. In other embodiments, the genetic material is not contained within the test genetic material sample.
- the genetic material can be one or more strands.
- the target genetic material comprises a representative selection of nucleic acids. In some embodiments, the target genetic material comprises a genome wide selection of nucleic acids. Unless explicitly noted otherwise, the term "genetic material" can be singular and/or plural (that is, “genetic material” can, for example, denote genetic material from one or more sources).
- polynucleotide As used herein, the terms “polynucleotide,” “oligonucleotide,” and “nucleic acid oligomers” are used interchangeably and mean single-stranded and double- stranded polymers of nucleic acids, including, but not limited to, 2'-deoxyribonucleotides (nucleic acid) and ribonucleotides (RNA) linked by internucleotide phosphodiester bond linkages, e.g. 3'-5' and 2'-5', inverted linkages, e.g. 3'-3' and 5'-5', branched structures, or analog nucleic acids.
- nucleic acids including, but not limited to, 2'-deoxyribonucleotides (nucleic acid) and ribonucleotides (RNA) linked by internucleotide phosphodiester bond linkages, e.g. 3'-5' and 2'-5', inverted linkages, e
- Polynucleotides have associated counter ions, such as H + , NH 4 + , trialkylammonium, Mg 2+ , Na + and the like.
- a polynucleotide can be composed entirely of deoxyribonucleotides, entirely of ribonucleotides, or chimeric mixtures thereof.
- Polynucleotides can be comprised of nucleobase and sugar analogs. Polynucleotides typically range in size from a few monomeric units, e.g. 5-40 when they are more commonly frequently referred to in the art as oligonucleotides, to several thousands of monomeric nucleotide units.
- the term "reduce” denotes some decrease in amount.
- an event is reduced by 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99, 99.9, 99.99, 99.999, percent or more, including any value above any of the preceding values, as well as any range defined between any two of the preceding values.
- the term "whole genome” means “genome wide” rather than requiring that the entire genome of any organism be present. Genome wide indicates that there is a sufficient variety and selection of various nucleic acids throughout an organism's genome for the technique being performed.
- the genome wide selection can be random, throughout an organism's genome, or biased to specific areas. In some embodiments, the genome wide selection is biased to those areas with the specific SNPs to be investigated. In some embodiments it is possible that less than one copy of an entire genome is used, such as in a degraded sample or a haploid sperm cell, as long as sufficient portions of genomic nucleic acid exist at enough SNPs to discriminate between a mixture and a person. This can be as few as a 1 ,000 SNPs, noting that millions of SNPs are known within the human genome. For example, one can identify an individual using only SNPs on chromosome 1.
- test genetic material sample denotes the sample whose composition is in question. Typically, one would like to know if a specific individual contributed to the genetic material in the test genetic material sample, and/or if other people or organisms contributed to the genetic material in the test genetic material sample.
- the test genetic material sample is the sample that is to be or has been assayed for the presence or absence of various SNPs.
- the target nucleic acid is contained within the test genetic material sample. In some embodiments, the target nucleic acid is not within the test genetic material sample.
- sample SNP signature is the SNP signature for the test genetic material sample.
- SNP signature denotes one or more various SNPs and the genotype, alleles, and/or percentage thereof for a collection of SNPs to be assessed.
- a “reference signature” denotes the alleles present for the SNPs in the reference (or a population thereof).
- a “test genetic material sample signature” denotes the alleles present for the SNPs in the test genetic material sample.
- a “subject's SNP signature,” “Person of Interest's SNP Signature,” or other similar term denotes the alleles present for the SNPs in the subject or Person of Interest.
- SNP signature does not require that the entire SNP signature be used (unless the term “entire” is explicitly used).
- comparing, employing and/or using one SNP signature with or to another SNP signature can be achieved merely by comparing a subset of the frequencies of the various alleles or by other approaches described herein.
- a SNP signature can denote one or more various SNP alleles and their frequency(ies)
- a comparison of the SNP signatures encompasses any comparison of one or more SNPs from one source to one or more alleles from a second source, as such, "comparing" a first and a second SNP signature does not actually require comparing the frequency statistics for each SNP allele (unless explicitly stated), but can be achieved by comparing and/or analyzing any data or computation that relates to these frequencies.
- the comparison can also be achieved by comparing values (including raw data) that are used to derive the noted frequencies. It can also be achieved by comparing values that are subsequently derived from the noted frequencies.
- values including raw data
- One of skill in the art will appreciate how to maintain the appropriate relationships between the various SNP signatures, based upon the present disclosure.
- a "person of interest” is not limited to a human being and, unless specified, can be any subject, such as any subject that includes genetic material (human, mammal, bacterial, viral, etc.).
- the term “Person of Interest” does denote that the subject is the one whose genetic material is being examined in the test genetic material sample. While this subject can typically be human, for example in many forensics tests, it is not limited to humans, unless explicitly noted.
- the term "reference population” denotes a population of one of more reference subjects.
- the SNP signature of the reference subjects allows for a comparison between the SNP signature of the person of interest and the SNP signature of the test genetic material.
- a reference population or SNP signature of a reference population is not required for all embodiments disclosed herein.
- the reference population and reference SNP signature will have a similar ancestral make-up as that of the sample SNP signature.
- the term "similar ancestral make-up" can be defined as a genetic distance between individuals or within a population using a set of SNPs or other genetic variants. Thus it is possible for some SNPs to be reserved for assessing ancestry and some SNPs reserved for assign wither a POI is within a mixture.
- the reference population should generally match the mixture at the SNPs being interrogated at the SNPs being investigated..
- a SNP is an inherited substitution of a nucleotide (for example from A to T, A to G, or G to C) found within more than two individuals. Generally most SNPs exceed a frequency greater than 0.1%, though lower frequency genetic variants are also envisioned.
- the methods described herein are extendable to other types of genetic variants, including indels, copy number changes, and/or other structural variants.
- test-statistic there are multiple approaches to derive a test-statistic to evaluate a hypotheses that a subject's genetic material is within a mixture, and these are discussed further in herein.
- a frequentist approach is used.
- a Bayesian approach is used. Either can be used depending on the objective of the assay.
- other approaches are used without deviating from the present methods.
- FIG. IA An overview of some embodiments of the approach is provided in FIG. IA.
- this method can be summarized as the cumulative sum of allele shifts over all available SNPs, where the shift's sign is defined by whether the individual of interest is closer to a reference sample or closer to the given mixture.
- One aspect of the invention encompasses genotyping a given SNP of a single person, which addresses the original design of SNP genotyping microarrays.
- the invention can be further adapted method to mixtures and pooled data.
- Genotyping microarray technology can assay millions of SNPs. Genotypes are expected to result from an assay and data is categorical in nature, e.g. AA, AB, BB 5 or NoCaIl where A and B symbolically represent the two alleles of a biallelic SNP. However, as evident from copy number, calling algorithm, and pooling-based GWA studies (Pearson et al; Am J Hum Genet. 2007 Jan;80(l): 126-39.
- raw preprocessed data from SNP genotyping arrays is typically in the form of allele intensity measurements that are proportional to the quantity of the "A" and "B" alleles hybridized to a specific probe (or termed features) on a microarray.
- Individual probe intensity measurements can be derived from the fluorescence measurement of a single bead (e.g. Illumina), micron-scale square on a flat surface (e.g. Affymetrix) or some combination thereof.
- a genotyping array multiple probes are present per SNP at either a fixed number of copies (Affymetrix) or a variable number of copies (Illumina).
- Affymetrix arrays typically have 3 to 4 probes specific for the A allele and B allele respectively, whereas Illumina arrays have a random number of probes averaging approximately 18 probes per allele.
- 500,000+ SNPs there are millions of probes (or features) on a SNP genotyping array. While there are considerably different sample preparation chemistries prior to hybridization between SNP genotyping platforms, any of these chemistries can be used, as they should not impact various embodiments disclosed herein.
- Y ⁇ transformation approximates allele frequency, where k j is the SNP specific correction factor accounting for experimental bias and is easily calculated from individual genotyping data.
- Y 1 is an estimate of allele frequency (termed J!? ⁇ ) of each SNP.
- values of the A allele frequency (P A ) in a single individual may be 0%, 50%, or 100% for the A allele at AA, AB, or BB, respectively.
- Equivocally Y 1 will be approximately 0, 0.5, or 1, varying from these values due to measurement noise.
- AB genotype calls are expected.
- the assumptions of the genotype-calling algorithm are invalid, since only AA, AB, BB, or NoCaIl are given regardless of the number of pooled chromosomes.
- one of skill in the art given the present disclosure, will be able to extract information and meaning from the relative probe intensity data and so be able to use that data to, for example, identify if a subject contributed to the mixture.
- M mixture
- M 1 A/(A l +k l B,)
- the mean allele frequency of the reference population is also encompassed within the term reference SNP signature.
- the reference population has a similar ancestral make-up as that of the mixture. This can mean having similar population substructure, ethnicity, and/or ancestral components interchangeably, and define similar ancestral components of an individual or mixture as having similar allele frequencies across all (or substantially all) SNPs.
- Y tJ be the allele frequency estimate for the individual i and SNP 7, where Y tJ e ⁇ 0,0.5,1 ⁇ , from a SNP genotyping array.
- the allele frequency estimate for the individual is also encompassed within the term subject SNP signature.
- the first difference ⁇ Y, j - M ⁇ ⁇ (which can also be characterized as the absolute value of the sample SNP signature subtracted from the subject SNP signature) measures how the allele frequency of the mixture M 1 at SNP j differs from the allele frequency of the individual Y, j for SNPy (or, put another way, measures how the sample SNP signature differs from the subject SNP signature).
- the second difference ⁇ Y fJ - Pop j ⁇ (which can also be characterized as the absolute value of the reference SNP signature subtracted from the subject SNP signature) measures how the reference population's allele frequency P Op 1 differs from the allele frequency of the individual Y tJ for each SNP j (or, put another way, measures how the reference SNP signature differs from the subject SNP signature).
- the values for P Op 1 can be determined from an array of equimolar pooled samples or from databases containing genotype data of various populations. Taking the difference between these two differences, one obtains the distance measure used for individual Y 1 :
- test statistic By sampling numerous SNPs (e.g., 500K+ SNPs), one would generally expect D(Y 1 J to follow a normal distribution due to the central limit theorem. In some embodiments, one can take a one-sample t-test for the subject, sampled across all (or at least one or more) SNPs, and thus obtain the test statistic:
- T(Y 1 ) (mean(D(Y ⁇ j )) - ⁇ 0 ) / (sd(D(Y t J/ sqrt(s))) Equation 2
- ⁇ o is the mean of D(Y k ) over individuals Y k not in the mixture
- Sd(D(Y 1 J) is the standard deviation Of D(Y 1 J for all SNPs j and individual Y 1
- sqrt(s) is the square root of the number of SNPs.
- one can set ⁇ o at zero since a random individual Y k should be equally distant from the mixture and the mixture's reference population and so T(Y 1 ) MeUn(D(Y 1 J) / (Sd(D(Y 1 J/sqrt(s)) . Under the null hypothesis T(Y 1 ) is zero and under the alternative hypothesis T(Y 1 ) > 0. In order to account for subtle differences in ancestry between the individual, mixture, and reference populations one can normalize allele frequency estimates to a reference population.
- STR short tandem repeats
- This genomic approach does not target specific sequences, regions or small number of polymorphisms, but instead can employ multiplex experiments performed on SNP microarrays to resolve whether an individual is present in a complex mixture. In some embodiments, this method also does not rely on knowing the number of individuals in the mixture.
- SNP microarrays have been widely used in Genome-wide Association studies, and when applied to Forensics SNP microarrays over a level of multiplexing not previously found in other methods. Nevertheless, Homer et al. (and the results discussed above and in Example 1) provide a frequentist approach based on cumulative shifts of relative allele signals across all SNPs to provide a significance value for the null hypothesis, where the individual is assumed not to be in the mixture.
- two microarrays can be run, one using DNA from the individual of interest and one using the pool of DNA from the mixture. This allows one to use a reference population for comparison, allowing one to accurately identify if an individual is present in the mixture. Additionally, this can be achieved even if a relative's DNA was used as a proxy for the individual of interest. Although such an embodiment performs well for many complex mixtures, other approaches can be used and as such, a probabilistic model is presented in the following section. Bayesian
- the following section discloses a probabilistic model based on the total observations at the raw intensity level for SNP microarrays to accurately assess the likelihood that the individual of interest (e.g., subject) is or is not in the complex mixture (e.g., test genetic material sample). Additionally, a training dataset was used to estimate the probability distribution of the raw intensity level observations. Two models were compared, one where the individual of interest is assumed to be in the mixture, and another where the individual of interest is assumed not to be in the mixture, in the form of a posterior odds ratio. The likelihood of each of the two models was derived using Bayesian inference to accurately assess the probability of the observations. With this embodiment, a more robust and accurate model of the observations was created, giving a better statistical measure of evidence. As the number of SNPs available on current microarray technologies continues to increase, so will the accuracy of various embodiments of the method to identify the contribution of an individual to a highly complex mixture. Models
- the modeling is performed to identify whether or not an individual is present within a given complex mixture. Therefore one can examine the odds ratio between two competing models, one where the individual is assumed to be in the mixture (denoted ⁇ A ) and one where the individual is assumed not to be in the mixture (denoted ⁇ ).
- the observations for the individual of interest are denoted as x and the observations for the complex mixture were denoted as y for all s SNPs.
- the observation x for the individual of interest (e.g., subject) is a raw intensity value
- the observation ⁇ for the complex mixture is similarly defined.
- probe value On a given microarray there are typically multiple probes per SNP as well as pairs of intensity values per probe.
- the probe values can be combined by taking the mean probe value over all probes, and combing the pair of intensity values into a simple ratio of the two values. For example if one had the intensity pair X and Y one can use the ratio or for a more
- Ji. has been used in previous studies using complex mixtures of DNA, namely pooling-based Genome-wide Association studies (J.V. Pearson, MJ. Huentelman, R.F. Halperin, W.D. Tembe, S. Melquist, N. Homer, M. Brun, S. Szelinger, K.D. Coon, VX. Zismann, J.A. Webster, T. Beach, S.B. Sando, J. O. Aasly, R. Heun, F. Jessen, H. Kolsch, M. Tsolaki, M. Daniilidou, E.M. Reiman, A. Papassotiropoulos, M.L. Hutton, D.A. Stephan, and D. W. Craig.
- ⁇ for each SNP i and each individual in the mixture or for the individual of interest but in this case one can estimate the distribution of these probabilities by using a training dataset, from the HapMap Project (The International HapMap Project. Nature, 426:789-796, Dec 2003). From the HapMap Project one is able to obtain for a given individual both the consensus genotype calls and raw intensity values for each SNP on the Affymetrix 5.0 platform. The HapMap project has this information for 270 individuals from four distinct populations. Additionally, the genotypes for each SNP were not only derived from the corresponding raw intensity values but also from other microarray platforms and replicate experiments resulting in a consensus genotype call for each SNP. This gives one further assurance that the genotype call is correct.
- this training data set gives, for each SNP i, the population allele frequency of A denoted/?,. It is useful when selecting the training dataset population to consider the ancestry of the population since allele frequencies can vary over population, and therefore introduce systematic biases in the model. Nevertheless, if SNPs used in the likelihood calculations are chosen to be ancestrally unbiased and unlinked, one avoids an admixture problem and can treat each SNP independently.
- Pr( ⁇ i I Xi, ⁇ , Ki,-0») Pr(Pi ⁇ ⁇ ⁇ i t ⁇ ⁇ i )
- each SNP was defined to be independent one can simply examine each SNP i independently and take the product over the probabilities for each SNP so that s
- Pr(y I ⁇ , X 1 ⁇ A ) Jl Pr(yi
- Pr(Vi I K XU BA) Pr[Vi I ⁇ u ⁇ > ⁇ A )Pr ⁇ i ⁇ ⁇ > Xi> ⁇ A )
- Pr(m I ⁇ , ⁇ i, ⁇ ⁇ ) follows a binomial distribution
- I ⁇ 2)
- Pr( ⁇ i I ⁇ , Xi, 9 A ) Pr( ⁇ i
- a*) N ( ⁇ A , ⁇ A )
- HWE Hardy- Weinberg Equilibrium
- the above analysis leverages the number of SNPs on the microarrays to accurately assess the probability that an individual of interest (e.g., subject) is present within a highly complex mixture. Since the number of SNPs on microarrays is now over one-million, one is able to obtain a sufficient number of observations to determine inclusion when compared to previous methods.
- This embodiment of the method specifically computes the posterior odds ratio between two models. The first model assumes the individual of interest is not present in the mixture and the second model assumes the individual of interest is present in the mixture. One then derives a likelihood function for both models given the observations of the mixture and individual of interest. A training dataset is used to provide for each SNP probability distributions for the observed probe intensity values given the unordered genotypes.
- FIG. IB depicts a more schematic representation of how the genetic material matching techniques described herein can be employed.
- genetic material e.g., a test genetic material sample
- This SNP signature can be, for example, created by a SNP analysis of a reference population, or obtainable in data form.
- One can then, optionally, obtain a SNP signature of a subject, as shown in process 60.
- One can then determine if there is a direction or bias of an allele count and/or frequency within the sample relative to the reference and/or the subject's signature as shown in process 70.
- One can then, optionally, analyze the direction or bias to determine a likelihood that the subject's genetic material is in the sample as shown in process 80.
- One can, optionally, have any of the results from the above processes output to an end user or memory 90.
- any one of more of the processes in FIG. IB are performed by a module configured to perform the process, which, optionally, can be part of a system.
- FIG. IB also represents modules that are capable of performing the steps for optionally obtaining a sample that can (but need not) include genetic material (e.g., a test genetic material sample) as in 10; a module to optionally purify and/or amplify at least some of any genetic material within the sample as shown in 20; a module to optionally prepare the sample to be run on a SNP array as shown in 30; a module to optionally determine one or more SNPs in the sample to obtain a sample SNP signature as shown in 40; a module to obtain a SNP signature of a reference population as shown in 50; a module to optionally obtain a SNP signature of a subject, as shown in 60; a module to determine if there is a direction or bias of an allele count and/or frequency within the sample relative to the
- one also has a module to output any correlation (or lack thereof) between the subject SNP signature and the sample SNP signature and/or the reference SNP signature to an end user, display, memory, and/or computer readable storage. In some embodiments, this information is output or provided to the subject.
- the system comprises an input module, to input one or more SNP signatures; a processing module, to compare the two or more SNP signatures; and an output module, to output the comparison.
- any one or more of the above modules are executed on one or more computing devices.
- methods and functions described herein are not limited to any particular sequence, and the blocks or states relating thereto can be performed in other sequences that are appropriate. For example, described blocks or states may be performed in an order other than that specifically disclosed, or multiple blocks or states may be combined in a single block or state.
- any other way of displaying the correlation between the subject's genetic material and the test genetic material sample and/or the reference population's genetic material can also be used and output to an end user or memory.
- Appendix A is the computer programming listing appendix referred to above, which is part of this specification. It provides some embodiments of code files usable for executing some embodiments of the processes and/or modules provided herein.
- the first code in Appendix A and any other code in Appendix A are nonlimiting examples of the code that can be employed for some of the present embodiments.
- the code used in connection with the present invention need not include any or all of the code listed in Appendix A at the end of the specification. Nevertheless, in some embodiments, the computer programming comprises, consists, or consists essentially of the code listed on the first 84 pages of Appendix A.
- a method for determining likelihood that a subject contributed genetic material to a test genetic material sample is provided.
- one tests whether a POI is in the mixture by assessing the probability that the allele frequency of the mixture is biased towards the POI, as compared to one or more reference populations.
- a complex genetic material mixture (or test genetic material sample) is one that includes genetic material (such as DNA) derived from more than one source.
- a complex mixture can also contain compounds, the presence of which causes experimental noise that could mask identification in some techniques, such as STR analysis.
- the invention involves a method of rapidly and sensitively determining whether a trace amount ( ⁇ 1%) of genomic DNA from an individual source is present within a complex DNA mixture.
- test genetic material sample includes a compound that would prevent or complicate STR analysis.
- test genetic material sample includes a molecule that degrades nucleic acids.
- test genetic material sample includes proteins and/or enzymes.
- the test genetic material sample includes mRNA, RNA, siRNA, and/or DNA.
- the mixture includes, or is suspected of including genetic material/nucleic acids from more than one human, for example, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 80, 100, 150, 200, 300, 500, 1000, 10,000 humans or more, including any amount defined between any two of the preceding values or any amount greater than any one of the preceding values.
- the subject's genetic material in the test genetic material sample is, or is suspected of being the source of less than 100% of the genetic material, for example, less than 100%, 99, 98, 95, 90, 80, 70, 60, 50, 40, 30, 20, 10, 5, 1, 0.5, 0.1, 0.05, 0.01, 0.005, 0.001, 0.0005, 0.0001 percent or less of the sample's genetic material is from the subject, including any amount defined between any two of the preceding values or any amount greater than any one of the preceding values.
- a test genetic material sample need only be manipulated enough to allow for the application of the sample onto a SNP array
- a PCR reaction is performed on the genetic material (reference, subject, and/or test genetic material sample). In some embodiments, this can be a simple PCR reaction, although any method that amplifies the desired genetic material can be used.
- primers for the amplification reaction are included in or as part of a kit for the present method. The primers can be selected so as to amplify desired sections of the genetic material to selectively amplify the SNPs to be examined. In some embodiments, the same primers can be used on one or more of the samples from the reference, subject, and test genetic material sample to increase the likelihood that the same SNPs are being reviewed.
- the use of one or more the methods described herein allows one to reduce the manipulation of the sample (reference, subject, and/or test genetic material sample) prior to examining it to prepare a SNP signature. In some embodiments, impurities that would otherwise complicate a STR analysis are not removed for the SNP analysis.
- Sources can include human beings, pets, mammals, birds, reptiles, amphibians, other animals, various cell types, algae, slime mold, mollusks, plants, bacteria, viruses, and any other organism that contains genetic material, such as DNA, whether terrestrial or extraterrestrial.
- the SNP probes are selected so as to reduce any undesirable cross-hybridization.
- cross-hybridization is addressed by normalizing markers using a quantile normalization approach, and/or by direct measurement of an individual who is homozygote for a given allele.
- the probes are random probes.
- the probes are those that will hybridize to genetic material that is linked to or similar to standard STR forensics markers.
- the probes allow for examination of genetic material that would be examined via restriction fragment length polymorphism, PCR analysis, STR analysis, mitochondrial DNA analysis and/or Y-chromosome analysis.
- the probes probe genetic material related, the same as, or linked to the 13 specific STR regions for CODIS.
- the probes reveal information regarding one or more of the following STR locus: D3S1358, vWA, FGA, D8S1179, D21S11, D18S51, D5S818, D13S317, D7S820, CSFlPO, TPOX, THOl, and/or D16S539.
- SNPs that are near the above and/or other known STRs are employed.
- SNPs that track the above or other known STRs are employed.
- the number and variance of the probes is selected based upon the results presented in Example 1 , outlining probe variance, probe number, and the number of people in the mixture.
- the devices, parts, subparts, or methods described herein can be combined into a kit for practicing any of the disclosed techniques.
- any of the methods can be provide in written format (such as in a set of instructions), or on a computer readable media.
- any of the steps or processes described herein that are capable of being executed by a machine can be provided on a computer readable media.
- programming that obtains the various SNP signatures can be provided.
- programming that compares the various SNP signatures can be provided (such as executing any of the equations provided herein).
- programming that outputs a likelihood that a subject contributed to a test genetic material sample is provided. Any such programming can be on computer readable media and/or downloadable from an online source.
- the kit includes one or more primers for SNP amplification.
- the SNPs, and thus the primers are specific for regions useful in forensics.
- a large number of SNP primers are used, for example, more than 100, such as 101, 200, 500, 1000, 2000, 5000, 10,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, or more SNPs, including any amount defined between any two of the preceding values and any range greater than any one of the preceding values.
- kits include one or more reference SNP signatures.
- SNP signatures can be stored on computer readable media or downloadable from a website.
- the reference populations are identified by groups such that the appropriate reference population can be matched with the subject and/or test genetic material sample.
- the kit includes one or more subject SNP signatures.
- SNP signatures can include, for example, the SNP signatures of a selection of convicted felons.
- reference SNP signatures can include general selections from the population.
- reference SNP signatures are configured for cell selection, biopsies, or any of the other uses provided herein.
- the kit includes programming and/or software for executing any one or more of steps 10, 20, 30, 40, 50, 60, 70, 80, and/or 90 in FIG. IB.
- the programming and/or software is in a memory or on a computer readable memory.
- the programming and/or software outputs the results of any of the processes in FIG. IB. This can include outputting any correlation (or lack thereof) between the subject SNP signature and the sample SNP signature and/or the reference SNP signature to an end user, display, memory, and/or computer readable storage
- the kit includes a SNP array and ingredients for running a SNP array. In some embodiments the kit includes tools for collecting a forensics sample. In some embodiments, the kits include PCR amplification ingredients. In some embodiments, the kit includes phi-29 and/or a similar polymerase. In some embodiments, the kits do not include all or any STR analysis ingredients. VARIOUS APPLICATIONS
- any of the methods described herein can be applied to determine if a subject's genetic material, such as DNA, matches, is consistent with, or is in a test genetic material sample. In some embodiments, one provides a likelihood that the subject's genetic material is within or the source of the genetic material in the test genetic material sample.
- any of the methods described herein can be applied to determine whether or not a subject is pregnant. In some embodiments, any of the methods described herein can be applied to determine if a male is the father of an unborn child. In some embodiments, the methods described herein can be applied to determine (including simply determining if the child's genetic material is consistent with) paternity or maternity of a child in comparison to one or more candidate parents. In some embodiments, any of the methods described herein can be applied to determine if there is an unknown person present in the test genetic material sample (in other words, if someone other than or in addition to the subject contributed to the test genetic material sample).
- any of the methods described herein can be applied to determine if someone contributed to the test genetic material sample without having to assume or factor in the number of people that may have contributed to the test genetic material sample. In some embodiments, one performs the analysis of the test genetic material sample ignoring and/or without the knowledge and/or without estimating the number of individuals that contributed to a test genetic material sample. In some embodiments, any of the methods described herein can be applied to forensics. In some embodiments, any of the methods described herein can be applied to determine a percentage or a likelihood that the subject contributed genetic material (or the subject's genetic material is a match) to the test genetic material sample.
- any of the methods described herein can be applied to determine or characterize the nature of various cells in a population of cells. This can be useful for sorting or selecting some cells over other cells, or determining the purity of a sample that comprises cells.
- any of the methods described herein can be applied on various cells or tissue from a subject. For example, in some embodiments, one can use the methods on a sample from a biopsy and determine if there are malignant vs. benign cells, and/or healthy cells vs. cancerous cells, and/or the type of cancer present in the cells. In embodiments involving numerous cells types, in some embodiments, all or part of the cells can be examined together, instead of having to separate out individual cells. In some embodiments, any of the methods described herein can be applied to determine whether a test genetic material is from a human (and/or which human) in comparison to other nonhuman organisms.
- the subject SNP signature includes genetic material from (or data representing) multiple individuals. In some embodiments, this can allow for the comparison or screening of multiple individuals against a test genetic material. Thus in some embodiments, the subject SNP is actually one or more subjects to allow for screening one or more subjects against the test genetic material sample.
- the invention involves a method of identifying trace amounts of an individual's DNA within highly complex mixtures in forensic applications. Such applications include, for example, a situation in which the presence of DNA from numerous other individuals hampers the ability to identify the presence of any single individual.
- any of the methods provided herein can be used to analyze genetic material that is degraded or from the mitochondria. The large number of assayed SNPs can allow the partitioning of sets of SNPs for different analyses, such that a small subset of SNPs becomes reserved for detecting these and other artifacts.
- the test genetic material sample includes, or is assumed or believed to include genetic material from at least 2 subjects, for example, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 16, 18, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 500, 1000, or more subjects, including any range defined between any two of the preceding values and any range above any one of the preceding values
- one or more advantages of the invention include a focus on the ratio of intensity measures from common biallelic SNPs and more robust scaling in DNA quantity or quality at any given SNP. Additionally, in some embodiments, there is no need to assume a known number of individuals present in the mixture or have equal amounts of DNA from each individual present within the mixture. Furthermore, in some embodiments, it is easy to discern whether the mixture is closer to a population or towards the individual by utilizing a cumulative distance measure. Whereas few conclusions can be drawn by a SNP measurement that is slightly biased (less than 1%) towards an individual's genotype, considerable confidence can be gained by statistical analysis of the cumulative aggregate of all measurements across hundreds to millions of SNPs.
- 1,000-100,000 SNPs are used, including the range of 2,000 to 20,000, and 3,000 to 10,000 and approximately 5,000.
- using the genotypes of a given individual it is possible to detect an individual's presence or absence in any study with available summary statistics.
- each SNP signature comprises a collection of information about various SNPs (such as, for example, allele frequencies).
- the SNP signature is a collection of SNP information regarding the subject, reference population, or test genetic material sample.
- the information is expressed as a percentage.
- the information is expressed in absolutes (e.g., presence or absence of a specific allele).
- the SNP signature is expressed in terms of raw data that represents the alleles at the SNP.
- the SNP signature can be a fluorescence readout from a SNP array, which indicates which SNPs are present.
- the size of a SNP signature can vary based on how it is to be used. In some embodiments, where one is looking to see if an unknown person contributed to a test genetic material sample, relatively few SNPs are employed as any single unknown SNP present in the test genetic material sample can indicate the presence of an unknown person. In addition, in embodiments in which a lower number of people contributed (or may have contributed) to the genetic material in the test genetic material sample, fewer SNPs will be used than in situations in which a large number of people contributed to the TGMS (test genetic material sample).
- the number of SNPs used in any one signature can also determine the degree of certainty that one has that the subject contributed to the TGMS. Thus, in embodiments, where a high degree of certainty is not required, fewer SNPs can be used. In embodiments where a higher degree of certainty is desired, more SNPs can be employed in the SNP signatures.
- SNP small neurotrophic factor
- all of the SNPs in a subject are used.
- all the SNPs across multiple subjects are used.
- SNPs from various organisms or cells are used.
- the SNPs used in the various SNP signatures should overlap (that is the same SNPs should be in the sample SNP signature, the reference SNP signature and the subject's SNP signature), not all of the SNPs need to be present in all of the signatures.
- the number and identity of SNPs can be different across the different signatures. In some embodiments, the lowest number of SNPs is found in the subject's SNP signature.
- the SNP signature is at least one SNP.
- the SNP signature includes more than one SNP, for example 1, 5, 10, 15, 20, 100, 200, 300, 500, 1000, 2000, 3000, 5000, 9,000, 10,000, 15,000, 20,000, 30,000, 40,000, 50,000, 80,000, 90,000, 100,000 SNPs or more, including any amount defined between any two of the preceding values and any amount greater than any one of the preceding numbers.
- a SNP signature can include one or more genotypes of one or more organisms (or cell types, etc.) across any number of individuals. As noted above, some SNP signatures include SNP information for 50,000 or more SNPs for tens, hundreds or more people. Other SNP signatures only include SNP information for a single person, across numerous SNPs, while yet other SNP signatures include SNP information for a single person and as little as a single SNP. Unless noted otherwise, any of the SNP signatures (sample SNP signature, reference SNP signature, subject's SNP signature) can vary in the manner noted above.
- the SNP signature does not have to be a compilation of mathematical values of the allele frequencies in all embodiments. For example, raw data showing intensity values for the various SNP probes (and thus representing what alleles are present) can be used. Similarly, the frequencies can be examined one at a time, and thus, a massive table of frequencies need not be compared to another massive table of frequencies. In some embodiments, the SNP signature merely represents or correlates to the allele information such that comparisons (mathematical, visual, or otherwise), can be consistently made between the subject and the sample and/or the reference population. Of course, in embodiments that do not employ SNPs, the consistency of the SNP is not relevant, but the consistency of the other item being monitored will be.
- the invention involves the use of any analytical methods that can be used to resolve complex mixtures.
- the analytical method used can depend on the objective of the analysis. Non-limiting examples include an assumption that the SNPs on the array are independent from one another, an assumption that multiple SNPs on the arrays are correlated and are not independent (especially in the case of increasing microarray density).
- Further examples include using population databases such as from the HapMap Project to select a subset of independent markers to be used in the analysis, the use of haplotype-based methods or Linkage Disequilibrium (LD) methods to combine information from correlated SNPs, the use of a Bayesian method to select the most informative SNPs derived from a training dataset, and the use of explicit redundancy in correlated markers.
- population databases such as from the HapMap Project to select a subset of independent markers to be used in the analysis
- haplotype-based methods or Linkage Disequilibrium (LD) methods to combine information from correlated SNPs
- LD Linkage Disequilibrium
- any method that allows for using numerous (e.g., thousands of) low-information content markers to make a cumulative decision about whether a person is, or is not, (or an unknown person is) in a mixture can be employed.
- any method that allows for using hundreds to thousands of measurements of genetic variants can be employed for the methods described herein.
- SNP signatures are not required for all of the embodiments described herein, when they are used, they can be compared in a variety of ways. In some embodiments, any comparison, as long as it allows one to determine direction or bias of an allele count and/or frequency within the test genetic material sample relative to an allele count and/or frequency of the reference and an allele count and/or frequency in a subject, can be used. In some embodiments, any of the computational methods disclosed herein can be employed for this.
- the SNP signature is shown in terms of raw data or a data readout (such as a fluorescence readout on a SNP array)
- a data readout such as a fluorescence readout on a SNP array
- allele frequencies expressed as percentages can be used in some embodiments, in some embodiments, the SNP data itself is used in the comparisons.
- Some embodiments of the invention further encompasses software that implements any of the methods and/or steps and/or processes described herein.
- Precompiled UNIX binaries are available for a software implementation of some embodiments of the method and can be found in the attached Appendix A.
- the software can run its analysis using raw data from either Affymetrix or Illumina or by using genotype calls.
- the software is also able to normalize the test statistic using the reference population and/or adjust the mean test statistic using a specified individual.
- the user can restrict the SNPs considered to a subset of the total available SNPs. For raw input data one can match the distribution of signal intensities for each raw data file to that of the mixture input file (see platform specific analysis).
- test statistics and distance calculations are implemented including the noted test statistic, Pearson correlation, Spearman rank correlation and/or Wilcoxon sign test.
- the software is configured to determine direction or bias of an allele count and/or frequency within the test genetic material sample relative to an allele count and/or frequency of the reference and an allele count and/or frequency in a subject.
- ancestry and Reference Populations are ancestry and Reference Populations.
- the reference population (and reference SNP signature) should either (a) accurately matched in terms of ancestral composition to the mixture and person of interest or (b) be limited to analysis of SNPs with minimal (or known) bias towards ancestry.
- ancestry of the reference population could be determined by analysis of a small subset of SNPs, followed by analysis of a person's contribution to the mixture with a separate set of SNPs (recognizing that nearly 500,000 SNPs are assayed).
- mismatching ancestry can be accounted for by normalizing the test-statistic using a second reference population matched to the individual of interest obtaining the normalized test-statistic S(Y 1 ). If the reference population of the mixture is mismatched, the reference population of the individual of interest will nonetheless normalize the results. Unlike the reference population of the mixture, the individual of interest's reference population is matched to the individual of interest's ancestry or population substructure and thus serves as an anchor for the distribution of T(Y 1 ). Thus one can compute a p-value for observing the result Y 1 or more extreme for individual Y 1 , assuming the reference populations for both the mixture and individual of interest are inferred correctly.
- the mean reference population test-statistic mean mean(T pop ) when matching a reference population to the individual of interest, one can choose the mean reference population test-statistic mean mean(T pop ) as a close relative to normalize for interesting familial relationships or other considerations, one could also choose to estimate the subject's reference population test-statistic standard deviation sd(T pop ) from a heterogeneous population to give a conservative overestimate of the true standard deviation of the test statistic T(Y 1 ).
- the reference population matched to the subject accounts for error in selecting the reference population of the mixture.
- the reference population is ascertained by using ancestral informative markers that are non-redundant with markers used for detecting if a person is in a mixture. In some embodiments, the reference population is ascertained by using multiple reference groups to ascertain a genetic distance. In some embodiments, the reference population is ascertained by adding individuals selected from a database of SNP calls for many individuals to effectively make a 'reference population' matched to ancestrally informative markers. In some embodiments, the reference population is obtained by collecting the SNPs of various suspects, which can optionally include the person of interest. In some embodiments, the reference population is obtained from an individual, such as a cancer patient or candidate that desires to see if she is pregnant.
- the reference population is a family or part thereof. In some embodiments, the reference population has no bias. In some embodiments, the reference population has a minimal bias measured by a genetic distance, genomic control, and which can be obtained using a subset of the SNPs not utilized for resolving within the mixture and not in linkage disequilibrium with any SNPs used in the analysis. In some embodiments, the reference population has a bias, but it is a known bias.
- the reference population is generally matched to the mixture at the SNPs being interrogated.
- High- information content SNPs can be used because they will be sensitive to different ancestral populations.
- these SNPs are independent of those SNPs used to identify a person, and thus could be restricted to one particular population.
- multiple references can be used and built into an overall likelihood statistic where a posterior probability is calculated.
- a large number of SNPs can have a correlation between each other, forcing the distribution to deviate from a normal distribution.
- additional methods such as using correction for these correlations, can also be used, such as linkage disequilibrium measurements as obtained through the HapMap project.
- the reference population comprises genetic material from one or more organisms, viruses, cell types, etc.
- the reference population can include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10,000, 50,000, 100,000, 500,000, 1,000,000, 5,000,000, 10,000,000, 100,000,000, 1,000,000,000, 5,000,000,000 or more different sources of genetic material.
- more than one reference and/or reference population and/or reference population signature can be employed by extending to a multiple dimensional test-statistic or distance measure.
- the device is a computer with relevant software to perform one or more of the processes outlined herein.
- the steps and processes disclosed herein can be implemented using combinations of one or more computing devices, such as webservers or peer-to-peer clients.
- the steps or processes can be performed on a single computing device, or, alternatively, a single step or process, such as 70 or combination of steps or processes, such as 10-90, 10-70, 20-70, 30-70, 40-70, 50-70, 60 & 70, 70 & 40, 70 & 60, and/or, 70 & 90 can be implemented on a computing device in communication with other computing devices that perform other steps or combinations of steps.
- a single step or process such as 70 or combination of steps or processes, such as 10-90, 10-70, 20-70, 30-70, 40-70, 50-70, 60 & 70, 70 & 40, 70 & 60, and/or, 70 & 90 can be implemented on a computing device in communication with other computing devices that perform other steps or combinations of steps.
- a system embodying these techniques can include appropriate input and output components, a computer processor, and a computer program product tangibly embodied in a machine-readable storage component or medium for execution by a programmable processor.
- a process embodying these techniques can be performed by a programmable processor executing a program of instructions to perform desired functions by operating on input data and generating appropriate output.
- the techniques can advantageously be implemented in one or more computer programs that are executable on a programmable system including at least one programmable processor coupled to receive data and instructions from, and to transmit data and instructions to, a data storage system, at least one input component, and at least one output component.
- Each computer program can be implemented in a high-level procedural or object-oriented programming language, or in assembly or machine language if desired; and in any case, the language can be a compiled or interpreted language.
- Suitable processors include, by way of example, both general and special purpose microprocessors. Generally, a processor will receive instructions and data from a read-only memory and/or a random access memory.
- Storage components suitable for tangibly embodying computer program instructions and data include all forms of non-volatile memory, including by way of example semiconductor memory components, such as Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), and flash memory components; magnetic disks such as internal hard disks and removable disks; magneto-optical disks; and Compact Disc Read-Only Memory (CD- ROM disks). Any of the foregoing can be supplemented by, or incorporated in, specially- designed ASICs (application-specific integrated circuits).
- EPROM Erasable Programmable Read-Only Memory
- EEPROM Electrically Erasable Programmable Read-Only Memory
- CD- ROM disks Compact Disc Read-Only Memory
- the entire process, from SNP analysis to final output of a likelihood that a subject's genetic material is in a test genetic material sample is automated and/or computerized.
- any of the results from steps 10-90 are output to an end user and/or a memory.
- any 1, 2, 3, 4, 5, 6, 7, 8 or 9 processes outlined in FIG. IB are performed and/or output via a computer.
- a computer prepares one or more SNP signatures and a person can make the comparison between the SNP signatures.
- a first computer can prepare one or more of the SNP signatures
- a second computer can prepare a different SNP signature
- a third computer can compare the different SNP signatures.
- the SNP signatures are standardized and contained in a memory system, cd, dvd, or other storage device. In some embodiments, such stored or standardized SNP signatures are for reference SNP signatures, subject SNP signatures, and/or sample SNP signatures. In some embodiments, the software and/or hardware is configured to detect various markers of various SNPs, develop the various SNP signatures (e.g., subject's SNP signature, test genetic material SNP signature and reference population SNP signature) and compare the SNP signatures.
- programming allows for the analysis of a SNP array.
- the analysis comprises data regarding fluorescence at various locations on the array of fluorescence generally.
- the programming allows for the comparison of a first SNP array (such as a subject SNP signature array) with a) second SNP array (such as a reference SNP signature array) and/or b) a third SNP array (such as a sample SNP signature array).
- one or more of the steps in FIG. IB are performed by different users and/or devices.
- the computer, device, memory, etc. comprises programming to allow for direction or bias of an allele count or frequency within a mixture relative to a reference and an in individual of interest to be determined.
- the computer, device, memory, etc. employs one or more of the formulas provided herein.
- the systems and methods described herein can advantageously be implemented using computer software, hardware, firmware, or any combination of software, hardware, and firmware.
- the system is implemented as a number of software modules that comprise computer executable code for performing the functions described herein.
- the computer- executable code is executed on one or more general purpose computers.
- any module that can be implemented using software to be executed on a general purpose computer can also be implemented using a different combination of hardware, software or firmware.
- such a module can be implemented completely in hardware using a combination of integrated circuits.
- such a module can be implemented completely or partially using specialized computers designed to perform the particular functions described herein rather than by general purpose computers.
- These computer program instructions can be stored in a computer- readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means which implement the acts specified herein.
- the computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions that execute on the computer or other programmable apparatus provide steps for implementing the acts specified herein.
- the invention further encompasses the use of a library of Y t arithmetic means derived from AA, AB, and BB to map genotype calls to expected Y t values to each SNP from individually genotyped samples.
- the method comprises the construction of a series of simulations to evaluate the theoretical limits of resolving an individual within a mixture using the described analytical framework and given characteristics of current generation SNP genotyping microarrays.
- the method further comprises experimentally testing the feasibility of detecting if an individual is contributing trace amounts of DNA to highly complex mixtures.
- particular focus was given (for some of the embodiments) on complex mixtures - those containing hundreds or thousands of individuals. Such approaches have utility in resolving a mixture of DNA from common surfaces where many individuals have left DNA.
- the invention involves a cumulative analysis of shifts in allele probe intensities in the direction of the individual's genotype. In some embodiments, the invention involves a method of measuring the difference between the distance of the individual from a reference population and the distance of an individual from the mixture. In some embodiments, one advantage the invention holds over other methods in field is that the method does not require knowledge of the number of individuals in the mixture and is capable of discriminating an individual source from a mixture comprising over one thousand sources.
- Example 1 provides an explanation of some of the embodiments with modifications in response to various factors including homogeneity of the mixture and accuracy of the reference populations.
- EXAMPLE 1 [0144] Complex Mixture Constructions. A total of 8 complex mixtures were constructed (See Table 1). Concentrations of all DNA samples were checked in triplicates using the Quant-iT PicoGreen dsDNA Assay Kit by Invitrogen (Carlsbad, CA). An eight point standard curve was prepared using Human Genomic DNA from Roche Diagnostics (Cat#: 11691112001, Indianapolis, IN). The median concentrations were calculated for each individual DNA sample.
- Mixtures Al, A2, Bl, and B2 Equimolar mixtures of HapMap individuals. Shown in Table 1, two main mixtures (mixtures A and B) were composed in duplicates resulting in a total of 4 mixtures. Mixture A was composed of 41 HapMap CEU individuals (14 trios minus one individual) and mixture B was composed of 47 HapMap CEU individuals (16 trios minus one individual).
- Mixture Cl 90% NA12752 and 10% NA07048. Two CEU males were combined in a single mixture so that one individual (NA 12752) contributed 90% (675ng) of the DNA in the mixture, while the other individual (NA07048) contributed 10% (75ng) DNA into the mixture by concentration.
- Mixture C2 90% NA10839 and 10% NA07048. Two CEU individuals, a female and a male, were combined in a single mixture so that one individual (NAl 0839) contributed 90% (675ng) of the DNA in the mixture, while the other individual (NA07048) contributed 10% (75ng) DNA into the mixture by concentration.
- Mixture Dl 99% NA12752 and 1% NA07048.
- Two CEU males were combined in a single mixture so that one individual (NA 12752) contributed 99% (742.5ng) of the DNA in the mixture, while the other individual (NA07048) contributed 1% (7.5ng) DNA into the mixture by concentration.
- Mixture D2 99% NA10839 and 1% NA7048. Two CEU individuals, a female and a male, were combined in a single mixture so that one individual (NAl 0839) contributed 99% (742.5ng) of the DNA in the mixture, while the other individual (NA07048) contributed 1% (7.5ng) DNA into the mixture by concentration.
- Mixture E 50% Mixture Al and 50% Mixture of 184 equimolar Caucasians. Two mixtures were combined into a single mixture so that each of the original mixtures contributed the same amount of genomic DNA by volume into the final mixture.
- CAU2 mixture contained 184 Caucasian control individuals obtained from the Coriell Cell Repository.
- Mixture Al was constructed as above and contained 41 CEU individuals.
- Mixture F 50% Mixture B2 and 50% Mixture of 184 equimolar Caucasians. Two mixtures were combined into a single mixture so that each mixture contributed the same amount of genomic DNA by volume into the final mixture.
- CAU3 mixture contained 184 Caucasian control individuals obtained from the Coriell Cell Repository.
- Mixture B2 was constructed as above.
- Mixture G 5% Mixture A2 and 95% Mixture of 184 equimolar Caucasians. Two mixtures were combined into a single mixture with Mixture A2 comprising of 5% of the mixture and the CAU3 comprising of 95% of the mixture. CAU3 mixture contained 184 Caucasian control individuals obtained from the Coriell Cell Repository. Mixture A2 was constructed as above.
- Mixture H 5% Mixture Bl and 95 % Mixture of 184 equimolar Caucasians. Two mixtures were combined into a single mixture with Mixture Bl comprising of 5% of the mixture and the CAU2 comprising of 95% of the mixture. CAU2 mixture contained 184 Caucasian control individuals obtained from the Coriell Cell Repository. Mixture Bl was constructed as above.
- Genotyping Four cohorts were assayed on the Illumina (San Diego, CA) HumanHap550 Genotyping BeadChip v3, one cohort was assayed on the Illumina (San Diego) HumanHap450S Duo, and three cohorts were assayed on the Affymetrix (Emeryville, CA) Genome- Wide Human SNP 5.0 array, with each cohort being assayed on a single chip.
- Probe intensity values were extracted for analysis from the file folders generated by the BeadScan software for the Illumina platform, and from Affymetrix GTYPE 4.008 software for the Affymetrix data, as described in previous studies (See Pearson, J.V. et al. Identification of the genetic basis for complex disorders by use of pooling-based genomewide single-nucleotide-polymorphism association studies. Am J Hum Genet 80, 126-139 (2007)).
- this type of adjustment is the preferred type of normalization method when raw data is available for the mixture, person of interest, and reference population.
- the genotypes from the HapMap dataset (See The International HapMap Project. Nature 426, 789-796 (2003)) were used of both the person of interest and the reference populations instead of raw intensity values as had been done with the Affymetrix platform. With the mixture the raw intensity values were used. This set of data mimics the case where raw data may not be available but genotype calls are available. Reduction in errors between different microarrays was achieved by normalizing each microarray by dividing by the mean channel intensity from each respective channel. This was performed on the raw data from the mixture. This platform specific adjustment may not be needed when the raw data of a person's genotype is present on the same platform. In the Illumina specific example, the calls from the HapMap were utilized without having platform specific genotype data.
- Simulation was used to test the efficacy of using high- density SNP genotyping data in resolving mixtures.
- the relevant variables of the simulation are: the number of SNPs s, the fraction /of the total DNA mixture contributed by the person of interest Y t , and the variance or noise inherent to assay probes v p .
- theoretical mixtures were composed by randomly sampling individuals from the 58C Wellcome Trust Case-Control Consortium (WTCCC) dataset (See Genome-wide association study of 14,000 cases of seven common diseases and 3,000 shared controls. Nature 447, 661-678 (2007)). After removing duplicates, relatives and other data anomalies, a total of 1423 individuals remained.
- WTCCC Wellcome Trust Case-Control Consortium
- the genotype calls for these individuals were provided from the WTCCC and were previously genotyped on the Affymetrix 500K platform. Within each simulation, N individuals were randomly chosen to be equally represented in the mixture and then computed the mean allele frequency (Y 1 ) of the mixture for each SNP. SNPs j with an observed Yy below 0.05 or above 0.95 in the reference population were removed due to their potential for having false positives and low inherent information content.
- a microarray was simulated that would contain a mean of 16 probes for simplicity, approximating the mean number of probes found on the Illumina 550K, Illumina 450S Duo and Affymetrix 5.0 platforms (18.5, 14.5 and 4 respectively).
- the Y g of each probe was added to a Gaussian noise based off the previously measured probe variance.
- probe variance was set to 0.006 when simulating Affymetrix 5.0 arrays, and to 0.001 for both Illumina 550K and Illumina 450S Duo arrays. The allele frequency of the mixture was then calculated to be the mean of these probe values.
- Equimolar mixtures ranging from 10 individuals to 1 ,000 individuals were tested. Using this design, each individual was tested for their presence where they contributed between 10% and 0.1% genomic DNA to the total mixture. To obtain significance levels (p-values) to test the null hypothesis, the normal distribution was sampled. There were not enough samples to test the tail of the distribution and therefore the p-values are not completely accurate (e.g. below 10 '6 ). Nonetheless, p-values are expected to be sufficiently accurate to qualitatively assess the limits of the method.
- the probe variance is below 0.001 one can easily resolve an individual whose DNA comprises 10% to 0.1% of the mixture. Even with increasing noise, one is still able to resolve mixtures where the person of interest contributes less than 2.5% with a p- value of less than 10 " .
- the probe variance does not have a large impact on the p- value, and in this case the fraction of the mixture is the important factor when the number of SNPs is fixed.
- the shading on the pvalues for FIG. 2b is noted in the bar beneath the graph. Dark grey is present primarily on the lower and right-hand side, followed by a band of white (as on moves left and upward across the graph), followed by an area of grey.
- the probe variance has little effect on the significance of the test. Consequently, it would be sufficient to use 50,000 SNPs, even with very high levels of noise to resolve mixtures of sizes up to 100.
- the number of probes is fixed to be 16, and thus the noise does not affect the allele frequency estimate, as would be the case with arrays using 4 probes.
- the shading on the pvalues for FIG. 2c is noted in the bar beneath the graph. Dark grey is present primarily on the left-hand side, followed by a band of white (as one moves to the right), followed by an area of grey.
- FIG. 3 shows the test-statistic for each individual within each mixture. Both individuals in the mixture and not in the mixture were tested for presence within the mixture.
- the left y-axis represents the -log p-value
- the right y-axis represents the normalized test-statistic S(Y tj )
- the bottom axis represents each individual.
- Each experiment was performed more than once and thus there are multiples of 86 individuals indexed on the bottom axis. For mixtures A, B, E, F, G and H, those in the mixture are shaded lightly and identified and those not in the mixture are shaded darker and identified. All individuals in the mixtures composed of more than 40 individuals were identified with zero false positives
- This example demonstrates a method to detect the presence of an individual's genetic material (nucleic acid) in a complex mixture of genetic material from multiple subjects.
- a reference sample of genetic material is created to provide an estimate of the mean allele frequencies of SNPs in the population represented by the reference sample (to obtain a reference SNP signature).
- the reference sample can be constructed by obtaining samples of genetic material from a commercial provider, such as the Coriel Cell Repository (Coriel Institute for Medical Research, Camden, NJ).
- the reference sample is composed of genetic material from one hundred individuals of Caucasian descent.
- the genetic material for the reference sample is available from the Coriel Cell Repository, Catalog number HDlOOCAU.
- the specific SNPs to be included in the analysis are selected.
- the allele frequencies of all selected SNPs in the reference sample are measured. Once measured, SNPs with a mean allele frequency less than 0.05 or greater than 0.95 are eliminated from consideration. All remaining SNPs are selected for use in the subsequent analysis, and the mean allele frequencies from those remaining SNPs are recorded. Alternatively, the allele frequencies of the selected SNPs can be obtained from a database that has previously measured the allele frequencies of the selected SNPs in a comparable reference population.
- DNA is taken from a person of interest (or subject). This DNA is analyzed to determine the allele frequencies of the selected SNPs in the DNA from the person of interest.
- the data obtained from the SNPs of the person of interest is compared with the data obtained from the reference sample and the data from the mixture to determine the source of the unknown sample. This process is repeated for a sufficient number of the selected SNPs to obtain the degree of certainty desired for establishing the match of the person of interest's DNA to the DNA in the complex mixture.
- the results from each SNP are combined and the output indicates the likelihood that the genetic material in the complex mixture belongs to the individual of interest.
- a reference sample of genetic material is assembled to provide an estimate of the mean allele frequencies of the SNPs to be analyzed in a given human population.
- the reference sample is constructed by obtaining samples of human genetic material from a commercial provider such as the Coriel Cell Repository (Coriel Institute for Medical Research, Camden, NJ). Genetic material from various human populations is available from the Coriel Cell Repository, including panels of individuals of Caucasian, African American, Middle Eastern, Asian, and other ethnic descents. In this example, reference samples representing panels of 10 or more individuals of Caucasian, African American, Middle Eastern, and Asian descent are obtained from the Coriel Cell Repository and combined to form the reference sample.
- the reference sample is then tested to determine the mean allele frequencies of all available SNPs and create a reference SNP signature.
- the mean allele frequencies of the SNPs to be analyzed can be obtained from a commercial database (thereby obtaining the reference SNP signature). SNPs returning a frequency value below 0.05 or above 0.95 can optionally be eliminated from consideration.
- a subject SNP signature is created by obtaining genetic material from the individual who is suspected of contributing genetic material to a sample obtained at a crime scene.
- the allele frequencies of the selected SNPs are measured for a genetic material sample from the subject to obtain the subject SNP signature.
- test genetic material sample the sample of genetic material from the crime scene (test genetic material sample) is analyzed.
- the test genetic material sample is analyzed and the mean allele frequencies of the selected SNPs are obtained and recorded, thereby providing the sample SNP signature.
- each of the signatures is compared to determine whether the unknown sample taken from the crime scene belongs to the subject.
- the subject SNP signature e.g., the allele frequency of each SNP for the subject
- the reference SNP signature e.g., the mean allele frequency of the same SNP in the reference
- the sample SNP signature the mean allele frequency in the test genetic material sample
- the output can be expressed in terms of the likelihood that the subject contributed to the test genetic material sample.
- the methods in the current disclosure are used to conduct a forensic analysis of a sample that has been degraded as a result of exposure to environmental or other factors.
- a reference sample of genetic material is assembled to provide an estimate of the mean allele frequencies of the SNPs to be analyzed in a given human population, and thereby provide a reference SNP signature.
- Genetic material from various human populations is available from the Coriel Cell Repository, including panels of individuals of Caucasian, African American, Middle Eastern, Asian, and other ethnic descents.
- Genetic material samples representing panels of 10 or more individuals of Caucasian, African American, Middle Eastern, and Asian descent are obtained from the Coriel Cell Repository and combined to form the reference sample.
- the reference sample is then tested to determine the allele frequencies of all available SNPs, thereby creating a reference SNP signature.
- SNPs returning a frequency value below 0.05 or above 0.95 are eliminated from consideration.
- a subject's genetic material is then collected from one or more individuals that are suspected of contributing genetic material to a test genetic material sample.
- genetic material is collected from 10 different suspects who had access to the location of the test genetic material sample.
- the genetic material from all 10 individuals is combined to form a mixture sample, and the allele frequencies of the selected SNPs are measured, thereby forming a subject SNP signature.
- the degraded sample of genetic material is analyzed.
- the allele frequencies of the selected SNPs are measured and recorded, creating a sample SNP signature.
- the signatures (or at least a part thereof) obtained from each sample are compared to determine whether the degraded sample belongs to one of the 10 individuals who contributed genetic material to the test genetic material sample.
- the allele frequency of at least some of the SNPs in the degraded sample is compared to the mean allele frequency of the same SNPs in both the reference sample and the mixture sample. This process is repeated as many times as necessary for the selected SNPs. One thereby obtains enough SNP comparisons to determine if one of the 10 subjects contributed to the genetic material in the test genetic material sample.
- the methods of the current disclosure are used to determine whether a human female is pregnant.
- a suitable sample (a sample that can contain genetic material from a fetus in the host) is taken from the female host for analysis.
- the genetic material in the sample is isolated and a sample SNP signature is prepared from the genetic material.
- a subject SNP signature is then prepared by using a sample from the female subject.
- sample SNP signature is compared to the subject SNP signature, and if the comparison reveals that another person's genetic material is present, such as through additional SNPs, one concludes that the host is pregnant.
- a further reference SNP signature can be used from an appropriate reference population, and the comparison can be between a) the subject SNP signature and each of b) the reference SNP signature and the sample SNP signature.
- the methods of the current disclosure are used to determine the paternity of an unborn child.
- a suitable sample is taken from a pregnant female for analysis.
- the sample will include genetic material from the unborn child.
- the SNPs in the sample are determined and a sample SNP signature is obtained from the unborn child.
- the sample can optionally include the mother's genetic material.
- the SNP signature of the potential father can be compared to the sample SNP signature, and when the sample SNP signature only includes genetic material from the child, the likelihood that the potential father is the father of the child can be determined.
- a reference SNP signature can be prepared and the SNP signature of the potential father can be compared to each of the reference SNP signature and the sample SNP signature to determine if the potential father contributed to DNA of the unborn child.
- a method is used to determine whether unknown tissue remains are of bovine or human origin.
- a reference sample is created by obtaining a sample of bovine genetic material.
- the bovine genetic material can be obtained from a donor bovine animal, or can be obtained from a commercial provider, such as the Coriel Cell Repository.
- the sample of bovine genetic material is prepared and analyzed to determine the mean allele frequencies of 1,000 SNPs. Remaining SNPs are selected for analysis and their values are recorded.
- the human genetic material can be obtained from a human donor, or can be obtained from a commercial provider, such as the Coriel Cell Repository.
- the human genetic material is analyzed, using the methods in the current disclosure, to determine the mean allele frequencies of the selected SNPs. Once obtained, the values are recorded.
- the data obtained from each sample are compared to determine the source of the unknown sample.
- the mean allele frequency of each SNP in the unknown tissue remains sample is compared to the mean allele frequency of the same SNPs in each of the bovine sample and the human sample. If the SNP frequencies of the unknown sample are more similar to the bovine allele frequencies, it will indicate a lower chance that the sample is human and if the SNP frequencies of the unknown sample are more similar to the human allele frequencies, it will indicate a lower chance that the sample is bovine.
- the results from each SNP are combined and summed, and the output indicates whether the unknown tissue remains are of bovine or human origin.
- an embryonic stem cell line is cultured in co-culture with several different mouse embryonic feeder cells for several passages. After culturing the embryonic stem cells for several passages, the embryonic stem cells are isolated from the mouse embryonic feeder cells. The methods of the current disclosure are then used as described below.
- a reference sample is created by combining genetic material from the several different feeder cell lines that are used to culture the embryonic stem cell line of interest. The mean allele frequencies of numerous available SNPs in the reference sample are measured and the values are recorded.
- the cell line of interest is a human embryonic stem cell line that is available from the NIH.
- a sample of this cell line is obtained, and the allele frequencies of the selected SNPs are measured and recorded.
- the embryonic stem cells of interest are isolated from the feeder cells.
- a sample of isolated embryonic stem cells is collected and the genetic material from the cells is prepared for analysis. The mean allele frequencies of the selected SNPs in the sample are obtained and recorded.
- the data obtained from the sample of isolated embryonic stem cells are compared to the data obtained from each of the embryonic stem cell sample and the feeder cell mixture sample.
- the allele frequency of each SNP in the isolated embryonic stem cell sample is compared to the mean allele frequency of the same SNP in each of the embryonic stem cell sample and feeder cell mixture sample. This process is repeated for all of the selected SNPs. The results from each SNP are combined and the output indicates whether the isolated embryonic stem cell sample is free of feeder cells.
- cells from the tumor are typically analyzed to determine whether the cells are malignant or benign.
- the methods in the current disclosure can be used to analyze cells from a tumor biopsy and determine whether those cells are malignant or benign.
- a benign tumor sample is created by combining genetic material from several different known benign tumor cells and/or healthy cells.
- several different known forms of benign bone tumors are used to create the sample.
- the mean allele frequencies of all available SNPs in the benign tumor sample are measured and the values are recorded.
- a malignant tumor sample is created to represent the different types of malignant bone cancers.
- several different known forms of malignant bone tumors are used to create the sample.
- Genetic material from malignant tumors classified as multiple myeloma, osteosarcoma, Ewing's sarcoma, and chondrosarcoma are combined to create the malignant tumor sample.
- the mean allele frequencies of the selected SNPs in the malignant tumor sample are measured and the values are recorded.
- a tissue biopsy is obtained from an unknown bone tumor and cells are isolated from the biopsied tissue using methods that are well known in the art. The genetic material from the cells is isolated and the mean allele frequencies of the selected SNPs are measured and recorded.
- the data obtained from the tumor biopsy sample are compared to the data obtained from each of the benign tumor sample and the malignant tumor sample.
- the mean allele frequency of each SNP in the unknown tumor biopsy sample is compared to the mean allele frequency of the same SNP in each of the benign tumor sample and the malignant tumor sample. This process is repeated for a sufficient number of the selected SNPs. The results from each SNP are combined, and the output indicates whether the tumor is composed of benign or malignant cells.
- This example demonstrates one method of comparing allele frequencies for a SNP.
- a first set of SNP data are identified as the reference population, and a second set of SNP data are identified as the mixture population.
- the allele frequency values of the data in the reference population are averaged to provide a mean allele frequency value for each SNP in the reference population (thereby providing a reference SNP signature).
- This process is repeated with the mixture population, providing a mean allele frequency value for each SNP in the mixture population (thereby providing a sample SNP signature).
- the value of the allele frequency at each subject's SNP is compared to the mean allele frequency value of the same SNP in both the reference population and the sample SNPs from the mixture.
- the mean allele frequency of the SNP in the mixture is subtracted from the SNP allele frequency value of the subject, and the absolute value of this difference is stored.
- the mean allele frequency of the SNP in the reference population is subtracted from the SNP allele frequency value of the subject, and the absolute value of this difference is stored.
- a value is obtained for the individual SNP by subtracting the absolute value of the first value from the second value.
- a negative value denotes that the subject is likely to be in the reference population.
- a positive value denotes that the subject is likely to be in the mixture, and a value of 0 denotes that the subject is equally likely to be in the mixture and the reference population.
- the above process can be repeated across all SNPs to be included in the analysis, and the value YiJ obtained for each SNP is summed as follows:
- the summation result is used to determine whether the subject is a member of the mixture population, a member of the reference population, or neither. Additionally, a one-sample t-test for individual i can be taken and used to obtain a test statistic as follows:
- a reference population for use with the methods of the current disclosure. Such a reference population can be used to manage the effect of ancestry on the allele frequencies observed across many samples.
- the subject's population is identified. If the subject is of Caucasian ancestry, a reference sample is created based on a Caucasian population.
- the reference sample can typically include samples from ten or more individuals who are members of the target population. Ideally, the individuals represent typical members of the target population.
- the samples used to create the reference sample can include both female and male Caucasian individuals.
- the reference population sample is constructed by obtaining representative samples of genetic material from members of the target population.
- the reference population sample can be constructed by obtaining samples of genetic material from individual donors. Ten Caucasian donors are chosen to create the reference population sample. Five of the donors are Caucasian females and five of the donors are Caucasian males.
- Samples of genetic material are obtained from each reference donor.
- the allele frequencies of each SNP are measured in each sample, and the results are recorded.
- the values obtained for each SNP are summed across all ten of the donor samples and the mean allele frequency value is determined.
- the mean allele frequency value of each SNP (e.g., a reference SNP signature) can then be used in subsequent analyses as the mean allele frequency value of the reference population.
- a sample of genetic material is obtained from a subject.
- the sample is analyzed and the allele frequencies of the SNPs in the sample are determined (providing a subject SNP signature).
- genetic material is isolated from the forensic sample.
- the sample is analyzed and the allele frequencies of the SNPs in the sample are determined (providing a sample SNP signature).
- the comparison can also include a reference SNP signature, where the subject's genetic material is also represented in the reference SNP signature, and the comparison can be between a) the subject SNP signature and the reference SNP signature, and b) the subject SNP signature and the sample SNP signature, in order to demonstrate that the subject is more likely to have contributed to the reference population than to the forensic sample.
- a forensic sample can contain genetic material from one or more unknown individuals. This example demonstrates how the currently disclosed methods can be used to determine whether a complex sample contains genetic material from one or more unknown subjects.
- the three SNP signatures are compared and the results indicate that the subject is not likely to have contributed to the genetic material in the forensic sample or that, while the subject did contribute to the forensic sample, at least one other subject, with a SNP signature difference from the subject's SNP signature, also contributed to the forensic sample.
- This example demonstrates one method of determining if any one of a number of subjects contributed to a test genetic material sample.
- Genetic material from a forensic sample is isolated and characterized to obtain a sample SNP signature.
- Genetic material from 100 subjects is isolated and characterized to obtain a subject SNP signature.
- the subject SNP signature includes the mean frequencies of the various SNPs across the 100 subjects.
- the three SNP signatures are compared, as described herein.
- the results demonstrate that at least one of the 100 subjects contributed to the test genetic material sample.
- additional individual comparisons can be made to determine which of the 100 subjects contributed to the test genetic material sample.
- This Example outlines how one can analyze SNP signatures.
- Each of the signatures includes the intensity levels from SNP microarrays from one of the microarrays of a reference sample, a subject sample, or a test genetic material sample.
- QuerySnpCode (int) (strtod(QuerySnp.c_str(), NULL) );
- Terminating ! ⁇ n IlluminaFiles. c_str ⁇ ); exitC ⁇ ); ⁇ assertCNoOfFiles>0) ;
- InputFileNames[i] Cchar *)mallocCsizeofCchar)*FILE_NAME_LENGTH); ⁇
- NumEntries GetNumEntries(InputFileNames, NoOfFiles);
- GreenMean- (GreenMin-l);
- RedMean- (RedMin-l) ;
- GreenFilterPercent- (GreenMin-l);
- RedFilterPercent- (RedMin-l) ;
- temp_double_green Cgreen_values[j] / GreenMean
- temp_double_red (red_values[j] / RedMean) ;
- V temp[Z*(k+l)3 (uintl6_t)red_values[j] ; /* Only keep the values if at least one is greater than zero */ if(temp[2*(k)+l] > 0 I I temp[2*(k+l)] > 0)
- temp_snp_code ( ⁇ int32_t)prev_snp_code
- temp_snp_code htonl(temp_snp_code)
- cur_green_val (*(Entries[i])).
- SortEntries (Entries, 0, NumEntries-1, GREEN);
- SortEntries (Entries, 0, NumEntries-1, RED);
- temp_entries[ctr] Entries[start_upper] ; start_upper++;
- variable may or may not have a "printable” string */ if(VariableName) fprintf(stderr, "VariableName: %s ", VariableName); fprintfCstderr, " ⁇ n”); switchCAction) ⁇ case Exit: fprintfCstderr, " ***** Exiting due to errors ***** ⁇ n "); exit(0); break; /* Not necessary actually! */ case Warn: return; break; default: fprintfCstderr, "Trouble! ! ! ⁇ n”);
- PopulationMean (double*)malloc(sizeof(double)*NumSnps); fprintfCstderr, "%s” , BREAK_LINE) ; fprintf(stderr, "Getting Population MeanVT);
- MeanPeopleTestStatisties Cdouble*)mallocCsizeofCdouble)*NumMeanPeople); fprintfCstderr, "%s", BREAK_LINE); fprintfCstderr, "Computing Mean People Test StatisticsXn”); GetTestStatisticsC&MixtureEntries, NumMeanPeople, MeanPeopleFileNames , &MeanPeopleChipTypes , NumSnps, &SnpNames , &PopulationMean , TestStatistic, &MeanPeopleTestStatistics, DistanceMeasure, CorrelationXDistance, CorrelationYDistance,
- Header ntohl(Header) ;
- ChipType as the third byte */ fread(&((*h).ChipType), sizeof(unsigned int), 1, fp);
- ChipType ntohl((*h) . ChipType) ;
- ProcessMMFlag ntohl((*h) .
- ProcessMMFlag ;
- Avearge PMA as the seventh byte */ fread(&((*h).AverageChannell), sizeof(unsigned int), 1, fp);
- curNumSnps Re ⁇ dGenotypesIntoEntriesC&PeopleEntries, SnpN ⁇ mes, NumSnps, PeopleFileN ⁇ mes[i]); ⁇ ssert(curNumSnps > 0);
- MixtureEntriesOrder[i] i; ⁇ SortSnpEntries(MixtureEntries, SMixtureEntriesOrder, 0, NumSnps-1);
- PeopleEntriesOrder[i] i; ⁇ SortSnpEntries(PeopleEntries, &PeopleEntriesOrder, 0, NumSnps-1);
- SortSnpEntries (SnpEntry ** Entries, int ** EntriesOrder, int low, int high)
- SortSnpEntries Entries, EntriesOrder, low, mid
- SortSnpEntries Entries, EntriesOrder, mid+1, high
- temp_entries (SnpEntry *)malloc(sizeof(SnpEntry)*(high-low+l)
- temp_order (int*)malloc(sizeof(int)*Chigh-low+l))
- TotalNoOfProbes GetMeanAndMin(CelFileName, CdfFileName, QuerySnp, &PMAMean, SPMBMean, &PMAMin, SPMBMin);
- pmA - (floatXPMAMin-1); inten[j].
- pmB - (float ⁇ PMBMin-1); if(FilterPercent > 0 && FilterPercent ⁇ 100 (inten[j] .pmA ⁇
- NoOfValues (u Lntl6_t)(4*No0fProbesTrue) ;
- NoOfV ⁇ lues htons(NoOfValues) ; fwrUe(SNoOfValues, sizeof(uintl6_t), 1, CurrentOutputFile);
Landscapes
- Health & Medical Sciences (AREA)
- Life Sciences & Earth Sciences (AREA)
- Physics & Mathematics (AREA)
- Engineering & Computer Science (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Medical Informatics (AREA)
- Theoretical Computer Science (AREA)
- Bioinformatics & Computational Biology (AREA)
- Biotechnology (AREA)
- Evolutionary Biology (AREA)
- General Health & Medical Sciences (AREA)
- Spectroscopy & Molecular Physics (AREA)
- Biophysics (AREA)
- Molecular Biology (AREA)
- Genetics & Genomics (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Analytical Chemistry (AREA)
- Chemical & Material Sciences (AREA)
- Artificial Intelligence (AREA)
- Bioethics (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Data Mining & Analysis (AREA)
- Databases & Information Systems (AREA)
- Epidemiology (AREA)
- Evolutionary Computation (AREA)
- Public Health (AREA)
- Software Systems (AREA)
- Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
Abstract
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US8291208P | 2008-07-23 | 2008-07-23 | |
| PCT/US2009/051441 WO2010011776A1 (fr) | 2008-07-23 | 2009-07-22 | Procédé de caractérisation de séquences à partir d'échantillons de matériaux génétiques |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP2332082A1 true EP2332082A1 (fr) | 2011-06-15 |
Family
ID=41129339
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP09790738A Withdrawn EP2332082A1 (fr) | 2008-07-23 | 2009-07-22 | Procédé de caractérisation de séquences à partir d'échantillons de matériaux génétiques |
Country Status (8)
| Country | Link |
|---|---|
| US (2) | US20100086926A1 (fr) |
| EP (1) | EP2332082A1 (fr) |
| CN (1) | CN102165456B (fr) |
| AU (1) | AU2009274031A1 (fr) |
| BR (1) | BRPI0915619A2 (fr) |
| CA (1) | CA2731830A1 (fr) |
| CO (1) | CO6351830A2 (fr) |
| WO (1) | WO2010011776A1 (fr) |
Families Citing this family (28)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| EP2425240A4 (fr) | 2009-04-30 | 2012-12-12 | Good Start Genetics Inc | Procédés et compositions d'évaluation de marqueurs génétiques |
| US12129514B2 (en) | 2009-04-30 | 2024-10-29 | Molecular Loop Biosolutions, Llc | Methods and compositions for evaluating genetic markers |
| US9447474B2 (en) | 2009-12-03 | 2016-09-20 | Yissum Research Development Company Of The Hebrew University Of Jerusalem, Ltd. | System and method for analyzing DNA mixtures |
| US9163281B2 (en) | 2010-12-23 | 2015-10-20 | Good Start Genetics, Inc. | Methods for maintaining the integrity and identification of a nucleic acid template in a multiplex sequencing reaction |
| CA2852665A1 (fr) | 2011-10-17 | 2013-04-25 | Good Start Genetics, Inc. | Methodes d'identification de mutations associees a des maladies |
| US8209130B1 (en) | 2012-04-04 | 2012-06-26 | Good Start Genetics, Inc. | Sequence assembly |
| US10227635B2 (en) | 2012-04-16 | 2019-03-12 | Molecular Loop Biosolutions, Llc | Capture reactions |
| EP2917368A1 (fr) * | 2012-11-07 | 2015-09-16 | Good Start Genetics, Inc. | Procédés et systèmes permettant d'identifier une contamination dans des échantillons |
| EP2971159B1 (fr) | 2013-03-14 | 2019-05-08 | Molecular Loop Biosolutions, LLC | Procédés d'analyse d'acides nucléiques |
| US10851414B2 (en) | 2013-10-18 | 2020-12-01 | Good Start Genetics, Inc. | Methods for determining carrier status |
| US11053548B2 (en) | 2014-05-12 | 2021-07-06 | Good Start Genetics, Inc. | Methods for detecting aneuploidy |
| AU2015277198B2 (en) | 2014-06-18 | 2018-11-08 | The Regents Of The University Of California | Method for determining relatedness of genomic samples using partial sequence information |
| WO2016025818A1 (fr) | 2014-08-15 | 2016-02-18 | Good Start Genetics, Inc. | Systèmes et procédés pour une analyse génétique |
| JP2016051461A (ja) * | 2014-08-29 | 2016-04-11 | 日本コントロールシステム株式会社 | クラスタリング装置、クラスタリング方法、およびプログラム |
| WO2016040446A1 (fr) | 2014-09-10 | 2016-03-17 | Good Start Genetics, Inc. | Procédés permettant la suppression sélective de séquences non cibles |
| JP2017536087A (ja) | 2014-09-24 | 2017-12-07 | グッド スタート ジェネティクス, インコーポレイテッド | 遺伝子アッセイのロバストネスを増大させるためのプロセス制御 |
| WO2016112073A1 (fr) | 2015-01-06 | 2016-07-14 | Good Start Genetics, Inc. | Criblage de variants structuraux |
| US10854316B2 (en) * | 2015-12-03 | 2020-12-01 | Syracuse University | Methods and systems for prediction of a DNA profile mixture ratio |
| CN105463116B (zh) * | 2016-01-15 | 2018-08-28 | 中南大学 | 一种基于20个三等位基因snp遗传标记的法医学复合检测试剂盒及检测方法 |
| WO2018144135A1 (fr) | 2017-01-31 | 2018-08-09 | Counsyl, Inc. | Systèmes et procédés permettant d'inférer l'ascendance génétique à partir de données génomiques à faible taux de couverture |
| US20200104285A1 (en) * | 2017-03-29 | 2020-04-02 | Nantomics, Llc | Signature-hash for multi-sequence files |
| CN108823296B (zh) * | 2017-05-05 | 2021-12-21 | 深圳华大基因股份有限公司 | 一种检测核酸样本污染的方法、试剂盒及应用 |
| AU2018317875A1 (en) * | 2017-08-17 | 2020-03-05 | Tai Diagnostics, Inc. | Methods of determining donor cell-free DNA without donor genotype |
| US12084720B2 (en) | 2017-12-14 | 2024-09-10 | Natera, Inc. | Assessing graft suitability for transplantation |
| US11931674B2 (en) | 2019-04-04 | 2024-03-19 | Natera, Inc. | Materials and methods for processing blood samples |
| CN111575386B (zh) * | 2020-05-27 | 2023-10-03 | 广州市刑事科学技术研究所 | 一种检测人y-snp基因座的荧光复合扩增试剂盒及应用 |
| WO2021251834A1 (fr) * | 2020-06-10 | 2021-12-16 | Institute Of Environmental Science And Research Limited | Procédés et systèmes d'identification d'acides nucléiques |
| CN120641985A (zh) * | 2023-01-30 | 2025-09-12 | 基因识别有限公司 | 元素溯源系统和方法 |
-
2009
- 2009-07-22 EP EP09790738A patent/EP2332082A1/fr not_active Withdrawn
- 2009-07-22 WO PCT/US2009/051441 patent/WO2010011776A1/fr not_active Ceased
- 2009-07-22 BR BRPI0915619A patent/BRPI0915619A2/pt not_active IP Right Cessation
- 2009-07-22 AU AU2009274031A patent/AU2009274031A1/en not_active Abandoned
- 2009-07-22 CN CN200980137391.9A patent/CN102165456B/zh not_active Expired - Fee Related
- 2009-07-22 US US12/507,695 patent/US20100086926A1/en not_active Abandoned
- 2009-07-22 CA CA2731830A patent/CA2731830A1/fr not_active Abandoned
-
2011
- 2011-02-23 CO CO11021798A patent/CO6351830A2/es not_active Application Discontinuation
-
2017
- 2017-04-03 US US15/477,808 patent/US10679728B2/en active Active
Non-Patent Citations (1)
| Title |
|---|
| See references of WO2010011776A1 * |
Also Published As
| Publication number | Publication date |
|---|---|
| CO6351830A2 (es) | 2011-12-20 |
| US20100086926A1 (en) | 2010-04-08 |
| CN102165456B (zh) | 2014-07-23 |
| BRPI0915619A2 (pt) | 2016-11-01 |
| AU2009274031A1 (en) | 2010-01-28 |
| WO2010011776A1 (fr) | 2010-01-28 |
| US20170206311A1 (en) | 2017-07-20 |
| CA2731830A1 (fr) | 2010-01-28 |
| US10679728B2 (en) | 2020-06-09 |
| CN102165456A (zh) | 2011-08-24 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2010011776A1 (fr) | Procédé de caractérisation de séquences à partir d'échantillons de matériaux génétiques | |
| EP4070318B1 (fr) | Systèmes et procédés d'automatisation d'appels d'expression d'arn dans un pipeline de prédiction de cancer | |
| AU2018375008B2 (en) | Methods and systems for determining somatic mutation clonality | |
| EP3571615B1 (fr) | Procédés d'évaluation non invasive d'alterations genetique | |
| EP3642747B1 (fr) | Procédés et systèmes de décomposition et de quantification de mélanges d'adn provenant de multiples contributeurs ayant des génotypes connus ou inconnus | |
| AU2014205038B2 (en) | Noninvasive prenatal molecular karyotyping from maternal plasma | |
| Harjanto et al. | RNA editing generates cellular subsets with diverse sequence within populations | |
| US20130324417A1 (en) | Determining the clinical significance of variant sequences | |
| Stoler et al. | Streamlined analysis of duplex sequencing data with Du Novo | |
| US20260011403A1 (en) | Detecting and genotyping variable number tandem repeats | |
| US12073921B2 (en) | System for increasing the accuracy of non invasive prenatal diagnostics and liquid biopsy by observed loci bias correction at single base resolution | |
| WO2025250322A1 (fr) | Génotypage pour répétitions en tandem | |
| HK40100599A (en) | Noninvasive prenatal molecular karyotyping from maternal plasma | |
| HK40080479A (en) | Noninvasive prenatal molecular karyotyping from maternal plasma | |
| Liu et al. | Transcriptomic Approaches for Muscle Biology and Disorders | |
| Sharma | Novel Algorithms to Estimate Genome Coverage Using High Throughput Sequencing Data | |
| Guo | Statistical Modelling of Mutations in Cancer Genomes | |
| Irizarry et al. | Model-Based Quality Assessment and Base-Calling for Second-Generation Sequencing Data | |
| HK40019712A (en) | Methods and systems for decomposition and quantification of dna mixtures from multiple contributors of known or unknown genotypes | |
| HK40019712B (en) | Methods and systems for decomposition and quantification of dna mixtures from multiple contributors of known or unknown genotypes | |
| NZ759848B2 (en) | Liquid sample loading | |
| NZ759848A (en) | Method and apparatuses for screening | |
| Shen et al. | Gene Expression Analysis Using RNA-Seq from Organisms Lacking Substantial Genomic Resources | |
| HK40012482B (en) | Methods for non-invasive assessment of genetic alterations | |
| HK40012482A (en) | Methods for non-invasive assessment of genetic alterations |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| 17P | Request for examination filed |
Effective date: 20110222 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC MK MT NL NO PL PT RO SE SI SK SM TR |
|
| AX | Request for extension of the european patent |
Extension state: AL BA RS |
|
| RAX | Requested extension states of the european patent have changed |
Extension state: RS Payment date: 20110222 |
|
| RIN1 | Information on inventor provided before grant (corrected) |
Inventor name: HOMER, NILS Inventor name: CRAIG, DAVID |
|
| 17Q | First examination report despatched |
Effective date: 20131219 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: EXAMINATION IS IN PROGRESS |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE APPLICATION IS DEEMED TO BE WITHDRAWN |
|
| 18D | Application deemed to be withdrawn |
Effective date: 20190201 |