Output Files
The following section describes the outputs produced by DRAGEN Array.
PGx CNV VCF File
DRAGEN Array produces one PGx CNV variant call file (VCF) (*.cnv.vcf) per sample to report the CN status on the gene and sub gene level, along with the CN events for PGx targets.
The PGx CNV VCF output file follows the standard VCF format. The QUAL field in the VCF file measures the CNV call quality. The CNV call quality is a Phred-scaled score capped at 60 and the minimal value is 0. Low quality calls (QUAL<7) are flagged by the Q7 filter. Low quality samples with LogRDev greater than a threshold 0.2 are flagged with the SampleQuality flag.
The PGx CNV VCF files are by default bgzipped (Block GZIP) and have the “.gz” extension. The compression saves storage space and facilitates efficient lookup when indexed with the TBI Index File. To view these files as plain text, they can be uncompressed with bgzip from Samtools or other third-party tools. The CNV VCF must be bgzipped and indexed to be used in downstream DRAGEN Array commands, such as star allele calling.
The PGx CNV VCF output file includes the following content.
##fileformat=VCFv4.1
##source=dragena 1.3.0
##genomeBuild=38
##reference=file:///hg38_with_alt/hg38_nochr_MT.fa
##FORMAT=<ID=CN,Number=1,Type=Integer,Description="Copy number genotype for imprecise events. CN=5 indicates 5 or 5+">
##FORMAT=<ID=NR,Number=1,Type=Float,Description="Aggregated normalized intensity">
##ALT=<ID=CNV,Description="Copy number variant region">
##FILTER=<ID=Q7,Description="Quality below 7">
##FILTER=<ID=SampleQuality,Description="Sample was flagged as potentially low-quality due to high noise levels.">
##INFO=<ID=CNVLEN,Number=1,Type=Integer,Description="Number of bases in CNV hotspot">
##INFO=<ID=PROBE,Number=1,Type=Integer,Description="Number of probes assayed for CNV hotspot">
##INFO=<ID=END,Number=1,Type=Integer,Description="End position of CNV hotspot">
##INFO=<ID=SVTYPE,Number=1,Type=String,Description="Structural Variant Type">
##OverallPloidy=1.8
##GCCorrect=True
##contig=<ID=1,length=248956422>
##contig=<ID=4,length=190214555>
##contig=<ID=10,length=133797422>
##contig=<ID=16,length=90338345>
##contig=<ID=19,length=58617616>
##contig=<ID=22,length=50818468>
##contig=<ID=22_KI270879v1_alt,length=304135>
#CHROM POS ID REF ALT QUAL FILTER INFO FORMAT 204619760001_R01C01
1 109687842 CNV:GSTM1:chr1:109687842:109693526 N <CNV> 60 PASS CNVLEN=5685;PROBE=124;END=109693526;SVTYPE=CNV CN:NR 2:0.966631132771593
4 68537222 CNV:UGT2B17:chr4:68537222:68568499 N <CNV> 60 PASS CNVLEN=31278;PROBE=383;END=68568499;SVTYPE=CNV CN:NR 0:0.376696837881692
10 133527374 CNV:CYP2E1:chr10:133527374:133539096 N <CNV> 60 PASS CNVLEN=11723;PROBE=194;END=133539096;SVTYPE=CNV CN:NR 2:0.980059731860893
16 28615068 CNV:SULT1A1:chr16:28603587:28613544 N <CNV> 57 PASS CNVLEN=8315;PROBE=164;END=28623382;SVTYPE=CNV CN:NR 2:0.980552325552963
19 40844791 CNV:CYP2A6.intron.7:chr19:40844791:40845293 N <CNV> 60 PASS CNVLEN=503;PROBE=38;END=40845293;SVTYPE=CNV CN:NR 2:0.9663775484762
19 40850267 CNV:CYP2A6.exon.1:chr19:40850267:40850414 N <CNV> 60 PASS CNVLEN=148;PROBE=21;END=40850414;SVTYPE=CNV CN:NR 2:0.9663775484762
22 42126498 CNV:CYP2D6.exon.9:chr22:42126498:42126752 N <CNV> 48 PASS CNVLEN=255;PROBE=370;END=42126752;SVTYPE=CNV CN:NR 2:0.981703411438716
22 42129188 CNV:CYP2D6.intron.2:chr22:42129188:42129734 N <CNV> 10 PASS CNVLEN=547;PROBE=333;END=42129734;SVTYPE=CNV CN:NR 2:0.965498002434641
22 42130886 CNV:CYP2D6.p5:chr22:42130886:42131379 N <CNV> 60 PASS CNVLEN=494;PROBE=172;END=42131379;SVTYPE=CNV CN:NR 2:0.970341562236357
22_KI270879v1_alt 270316 CNV:GSTT1:chr22_KI270879v1_alt:270316:278477 N <CNV> 60 PASS CNVLEN=8162;PROBE=91;END=278477;SVTYPE=CNV CN:NR 2:1.01191145130511
Cytogenetics VCF File
DRAGEN Array produces one cytogenetics Variant Call File (VCF) (*.cnv.vcf) per sample to report the CN and LOH status of the detected variants.
The cytogenetics CNV VCF output file follows the standard VCF format. The QUAL field in the VCF file measures the CNV/LOH call quality. The CNV/LOH call quality is a Phred-scaled score capped at 60 and the minimal value is 0. Low quality calls (QUAL<10) are flagged by the Q10 filter. Low quality samples with LogRDev greater than a threshold 0.2 are flagged with the SampleQuality flag.
The cytogenetics CNV VCF files are by default bgzipped (Block GZIP) and have the “.gz” extension. The compression saves storage space and facilitates efficient lookup when indexed with the TBI Index File. To view these files as plain text, they can be uncompressed with bgzip from Samtools or other third-party tools. The CNV VCF must be bgzipped and indexed to be used in downstream DRAGEN Array commands, such as cyto annotate.
One example file can be found below:
##fileformat=VCFv4.1
##source=dragena 1.3.0 Cyto
##genomeBuild=37
##product=GDACyto-8v1-0_A
##reference=file://genome.fa
##FORMAT=<ID=GT,Number=1,Type=String,Description="Genotype">
##FORMAT=<ID=CN,Number=1,Type=Integer,Description="Copy number genotype. CN=4 indicates 4 or 4+">
##FORMAT=<ID=NR,Number=1,Type=Float,Description="Aggregated normalized intensity">
##FORMAT=<ID=LRD,Number=1,Type=Float,Description="Standard deviation of logR ratios">
##platform=cytoplatform
##ALT=<ID=DEL,Description="Copy number loss region">
##ALT=<ID=DUP,Description="Copy number gain heterozygous region">
##ALT=<ID=LOH,Description="AOH/LOH/ROH, absence of heterozygosity region, or, loss of heterozygosity region">
##FILTER=<ID=Q10,Description="Quality below 10">
##FILTER=<ID=SampleQuality,Description="Sample was flagged as potentially low-quality due to high noise levels.">
##INFO=<ID=SVLEN,Number=1,Type=Integer,Description="Number of bases in CNV/LOH region">
##INFO=<ID=PROBE,Number=1,Type=Integer,Description="Number of probes assayed for CNV/LOH region">
##INFO=<ID=END,Number=1,Type=Integer,Description="End position of CNV/LOH region">
##INFO=<ID=LOHTYPE,Number=A,Type=String,Description="Type of LOH (Loss/absence of heterozygosity). Valid values are AOH (germline, copy number neutral or gain LOH), CNLOH (somatic, copy number neutral LOH), GAINLOH (somatic, copy number gain LOH)">
##OverallPloidy=1.9
##GCCorrect=True
##contig=<ID=1,length=249250621>
##contig=<ID=2,length=243199373>
##contig=<ID=3,length=198022430>
##contig=<ID=4,length=191154276>
##contig=<ID=5,length=180915260>
##contig=<ID=6,length=171115067>
##contig=<ID=7,length=159138663>
##contig=<ID=8,length=146364022>
##contig=<ID=9,length=141213431>
##contig=<ID=10,length=135534747>
##contig=<ID=11,length=135006516>
##contig=<ID=12,length=133851895>
##contig=<ID=13,length=115169878>
##contig=<ID=14,length=107349540>
##contig=<ID=15,length=102531392>
##contig=<ID=16,length=90354753>
##contig=<ID=17,length=81195210>
##contig=<ID=18,length=78077248>
##contig=<ID=19,length=59128983>
##contig=<ID=20,length=63025520>
##contig=<ID=21,length=48129895>
##contig=<ID=22,length=51304566>
##contig=<ID=X,length=155270560>
##contig=<ID=Y,length=59373566>
#CHROM POS ID REF ALT QUAL FILTER INFO FORMAT 208588190001_R02C01
1 109687842 DEL:chr1:109687842:109693526 N <DEL> 60 PASS SVLEN=5685;PROBE=99;END=109693526 GT:CN:NR:LRD 1/1:1:0.8860:0.21
16 28603587 DUP:chr16:28603587:28613544 N <DUP> 60 PASS SVLEN=9958;PROBE=197;END=28613544 GT:CN:NR:LRD 1/1:3:1.1666:0.11
22 42129188 AOH:chr22:42129188:42129734 N <LOH> 37 PASS SVLEN=547;PROBE=198;END=42129734;LOHTYPE=AOH GT:CN:NR:LRD 1/1:2:1.0208:0.25
SNV VCF File
The software produces one genotyping variant call file (*.snv.vcf) file per sample, covering single nucleotide variants (SNV) and indels for the sample. It reports GenCall score (GS), B Allele Frequency (BAF), and Log R Ratio (LRR) per variant. The VCF file output follows VCF4.1 format.
Some additional details:
The FILTER column is hardcoded to
PASSand is not dependent on theGTvalue. It does not reflect the underlying quality of the call. Refer to theGSvalue for quality information.Genotypes are adjusted to reflect the sample ploidy. Calls are haploid for loci on Y, MT, and non-PAR chromosome X for males.
Multiple SNPs in the input manifest which are mapped to the same chromosomal coordinate (e.g. tri-allelic loci or duplicated sites) are collapsed into one VCF entry and a combined genotype generated. To produce the combined genotype, the set of all possible genotypes is enumerated based on the queried alleles. Genotypes which are not possible based on called alleles and assay design limitations (e.g. Infinium II designs cannot distinguish between A/T and C/G calls) are filtered. If only one consistent genotype remains after the filtering process, then the site is assigned this genotype. Otherwise, the genotype is ambiguous (more than 1) or inconsistent (less than 1) and a no-call is returned.
Certain SNV and indel calls will be skipped when reported in the VCF. Skipped data can include unmapped loci (i.e.,
Chris0in the manifest), intensity-only probes used for CNV identification, and indels that do not map back to the genome. See Warning/Error Messages and Logs for messages that may be seen with DRAGEN Array Local related to the skipped data.The BAF and LRR are oriented with Ref as A and Alt as B relative to the reference genome, while GS is agnostic to the reference genome. Users familiar with GenomeStudio may observe BAF and LRR reported in the VCF as 1 minus the value reported in GenomeStudio depending on the Ref Alt allele orientation with the reference genome. GenomeStudio reports these values based on the information in the manifest without knowledge of the reference genome.
The SNV VCF files are by default bgzipped (Block GZIP) and have the “.gz” extension. The compression saves storage space and facilitates efficient lookup when indexed with the TBI Index File. To view these files as plain text, they can be uncompressed with bgzip from Samtools or other third-party tools. The SNV VCF must be bgzipped and indexed to be used in downstream DRAGEN Array commands, such as star allele calling.
The SNV VCF output file includes the following content. The last row shows an example of variant call.
##fileformat=VCFv4.1
##source=dragena 1.3.0
##genomeBuild=38
##reference=file:///genomes/38/genome.fa
##FORMAT=<ID=GT,Number=1,Type=String,Description="Genotype">
##FORMAT=<ID=GS,Number=1,Type=Float,Description="GenCall score. For merged multi-assay or multi-allelic records, min GenCall score is reported.">
##FORMAT=<ID=BAF,Number=1,Type=Float,Description="B Allele Frequency">
##FORMAT=<ID=LRR,Number=1,Type=Float,Description="LogR ratio">
##contig=<ID=1,length=248956422>
##contig=<ID=2,length=242193529>
##contig=<ID=3,length=198295559>
##contig=<ID=4,length=190214555>
##contig=<ID=5,length=181538259>
##contig=<ID=6,length=170805979>
##contig=<ID=7,length=159345973>
##contig=<ID=8,length=145138636>
##contig=<ID=9,length=138394717>
##contig=<ID=10,length=133797422>
##contig=<ID=11,length=135086622>
##contig=<ID=12,length=133275309>
##contig=<ID=13,length=114364328>
##contig=<ID=14,length=107043718>
##contig=<ID=15,length=101991189>
##contig=<ID=16,length=90338345>
##contig=<ID=17,length=83257441>
##contig=<ID=18,length=80373285>
##contig=<ID=19,length=58617616>
##contig=<ID=20,length=64444167>
##contig=<ID=21,length=46709983>
##contig=<ID=22,length=50818468>
##contig=<ID=MT,length=16569>
##contig=<ID=X,length=156040895>
##contig=<ID=Y,length=57227415>
#CHROM POS ID REF ALT QUAL FILTER INFO FORMAT 202937470021_R06C01
1 2290399 rs878093 G A . PASS . GT:GS:BAF:LRR 0/1:0.7923:0.50724137:0.14730307
Note on Multi-Allelic Variants (MAV) calling limitations
DRAGEN Array can combine multiple assays with different target bases but the same genomic position to make MAV calls. However, Illumina Microarrays are inherently bi-allelic assays made up of Infinium I or Infinium II probe designs which require special design considerations and have some inherent limitations.
The MAV calling algorithm currently filters across all overlapping assays, retaining only genotypes whose alleles are present in the intersection of all assays. If multiple genotypes remain after filtering, the result is considered ambiguous and reported as a NoCall to avoid false positives. This ambiguity often arises when one assay is a NoCall due to presumed probe failure rather than missing signal. In such cases, its potential genotypes are not excluded, contributing to ambiguity. When DRAGEN Array outputs a NoCall because of the described behaviors, they are logged as warnings (e.g., Failed to combine genotypes due to ambiguity...).
Overall, the current algorithm errs on the side of caution to ensure quality calls, but produces some idiosyncratic behavior and potential false NoCalls when genotypes are biologically consistent but differ due to probe designs. We hope to improve this behavior in future versions of DRAGEN Array.
Some illustrative examples are below to help understand the current limitations:
Inf II [A/G] -> AA + Inf I [T/G] -> TT
AT
NoCall
The Inf II assay cannot differentiate A versus T alleles, hence AA call for the Inf II assay is consistent with AA, AT, or TT genotypes. The Inf I assay TT call, however, is consistent with AT or TT genotypes. Combining the two assays results in a NoCall due to the AT or TT ambiguity.
Inf I [T/A] -> AA + Inf II [A/G] -> NC
AA
NoCall
NoCall for Inf II probe leaves possibility of AG. Ambiguity between AG or AA leads to NoCall.
Note on delimiters in the "ID" field
By default, when multiple probes are present for a given variant, all probe names are included in the "ID" field of the resulting VCF file.
For SNP entries, probe names are separated by commas (,).
For Indel entries, probe names are separated by semicolons (;).
Note on REF/ALT "flipping" for INDELs
Expected REF and ALTs for INDELS may not match the dbSNP annotations in rare cases. E.g., an expected "Deletion" with the REF = "ATCG" and the ALT = "A" may be "flipped" to an "Insertion" variant with REF="A" and ALT="ACTG". The corresponding genotype output will take this into account so the actual VCF is still correct. This is simply a notation issue in some of the manifest files.
Note on PLINK compatibility
It is possible to make DRAGEN Array genotype VCF files compatible for conversion to PED/MAP format with preprocessing using tabix (v1.19.1), BCFtools (v1.21) and PLINK (v1.9). The following three commands demonstrate the basic process.
bcftools merge -l vcf_list.txt -Oz -o merged.vcf.gzcreates a single compressed VCFs from individual sample VCFs listed in thevcf_list.txttext file.tabix -p vcf merged.vcf.gzcreates a binary index for the merged file.plink --vcf merged.vcf.gz --recode --out mergedcreates .ped and .map files with the prefix provided to--out.
Some optional arguments may be provided to PLINK depending on the content of the VCFs to be converted and the downstream analysis.
For VCFs containing non-standard human chromosomes (e.g. haplotype chromosomes or unplaced contigs), the
--allow-extra-chrflag can be used.If using non-human data, refer to the PLINK documentation for the
--chr-setargument and supported options.By default, PLINK will only consider the most common ALT allele for multi-allelic variants. The
--biallelic-onlyargument can be provided to exclude multi-allelic variants altogether. As an alternative, usingbcftools norm -m - in.vcf.gz -Oz -o out.vcf.gz, can be used upstream of PLINK to split multi-allelic variants into bi-allelic records to retain them for downstream processing.
For more info on the options described and others, refer the PLINK VCF conversion documentation.
Genotype Call (GTC) File
The genotype call algorithm produces one genotype call file (.gtc) per sample analyzed. The Genotype Call (GTC) file contains the small variant (SNV and indel) genotype for each marker specified by the product and sample quality metrics. The sample marker location is not included and must be extracted from the manifest file. Binary proprietary format can be parsed using the Illumina open-source tool BeadArray Library File Parser.
Note on lack of i18n: GTCs are binary/fixed format files built designed before modern internationalization and localization tools. There is a related known issue that makes the GTCs unable to be used in downstream analyses. Refer to the same issue to see a workaround.
Note on legacy GTCs: Other Illumina software (such as AutoConvert and Beeline) also product GTC files.
These "legacy GTC" files will work in DRAGEN Array genotyping commands such as genotype gtc-to-vcf but they will not work with all other downstream analyses such as Cytogenetics analysis and PGx. We recommend using DRAGEN Array end-to-end starting from IDATs for these analyses.
BedGraph Files
The BedGraph files contains the Log R Ratios (LRR.bedgraph) and B-Allele Frequencies (BAF.bedgraph) from the genotyping algorithm for use in visual tools.
Star Allele CSV File
The Star Allele CSV file is an intermediate file generated by the pgx star-allele call command and serves as the input to the pgx star-allele annotate command. It contains all the star allele calls for all samples in a run. Each row in the file provides either a star allele diplotype or simple variant call for a PGx-related gene. Star allele diplotype calls for a sample and a gene may span multiple lines where alternative solutions can be listed.
The Star Allele CSV file also contains meta information marked by # at the top of the file for the genome build and PGx database used for the star allele calling.
The star_allele.csv file contains the following details per sample:
Sample
Sentrix barcode and position of the sample.
Rank
Rank of a single star allele solution for a gene. The top solution based on quality score is ranked as 1 with the alternative solutions ranked lower.
Gene or Variant
The gene symbol, or gene symbol plus rsID for variants.
Type
‘Haplotype’ (star allele) or ‘Variant’ PGx calling type.
Solution
Star allele or variant solution. If diploid, variant solutions have the format of Allele1/Allele2.
Solution Long
Long format solution for star alleles. The field has the following format: Structural Variant Type: Underlying Star allele.
An example of a long solution is: Complete: CYP2D64, Complete: CYP2D610, CYP2D668: CYP2D64 where there are two complete alleles that have CYP2D64 and CYP2D610 haplotypes and one CYP2D668 structural variant that has a CYP2D64 haplotype configuration.
Supporting Variants
All variants present in the array that support the star allele solution. The field has the following format: Long Solution Star Allele: (Supporting Variants).
Each supporting variant is listed with essential information extracted from the SNV VCF to assist with troubleshooting, including Chromosome, Location, Reference allele, Alternative allele, Genotype, GenCall score (GS), and B-allele frequency (BAF).
Missing/Masked Core Variants
All variants not present in the array or not called in the SNV VCF file for the star allele. The field has the following format: Long Solution Star-Allele: (Missing Variants).
All Missing Variants in Array
All core definition variants that are not on the array or are not called in the SNV VCF along with the associated star alleles that are impacted. The field has the following format: Missing Variant: (List of impacted star alleles).
Collapsed Star-Alleles
Star alleles that cannot be distinguished from the solution star allele given the input array’s content. The field has the following format: Long Solution Star-Allele: (List of collapsed star alleles).
The most frequent star allele based on the population frequency of PGx alleles will be the star allele in the solution.
Score
Quality score of the solution including the population frequency of PGx alleles. The score ranges from 0 to 1.
Raw Score:
Raw quality score of the solution without including the population frequency of PGx alleles. The score ranges from 0 to 1.
Copy Number Solution
Estimated copy number for each gene region. The field has the following format: Gene Region: Copy Number.
Below is an example of the first 4 columns from a star allele CSV file:
Sample,Rank,Gene or Variant,Type,Solution
204650490282_R02C01,1,CYP2C9,Haplotype,*9/*11
204650490282_R02C01,1,CYP2C19,Haplotype,*2/*10
Genotype Summary Files
The software produces genotype summary files (gt_sample_summary.csv and gt_sample_summary.json) that contains the following details per sample:
Sample ID
Sample Name
Sample Folder
Autosomal Call Rate
Call Rate
Log R Ratio Std Dev
Sex Estimate
TGA_Ctrl_5716 Norm R
(Optional) User defined fields from the samplesheet
The TGA_Ctrl_5716 Norm R field is specific to PGx products (e.g., Global Diversity Array with enhanced PGx). The field value is the Normalized R value of one probe and is meant as an assay control where < 1 indicates the sample failed in the TGA (Targeted Gene Amplification) process. If the product does not have this probe, it is not included in the gt_sample_summary.
The user defined fields from the samplesheet will appear as-is in the gt_sample_summary files. e.g. for the given samplesheet:
It would produce something like the following gt_sample_summary.csv:
And something like the following gt_sample_summary.json:
Note: As of v1.3, samples that fail during genotyping will still be present in this file. See the details in the release notes.
Final Report
DRAGEN Array Cloud produces a Final Report (gtc_final_report.csv) per analysis batch similar to the one available in GenomeStudio. It contains the following details per locus per sample:
SNP Name
SNP identifier.
SNP
SNP alleles as reported by assay probes. Alleles on the Design strand (the ILMN strand) are listed in order of Allele A/B.
Sample ID
Sample identifier.
Allele 1 – Top
Allele 1 corresponds to Allele A and are reported on the Top strand.
Allele 2 – Top
Allele 2 corresponds to Allele B and are reported on the Top strand.
Allele 1 – Forward
Allele 1 corresponds to Allele A and are reported on the Forward strand.
Allele 2 – Forward
Allele 2 corresponds to Allele B and are reported on the Forward strand.
Allele 1 – Plus
Allele 1 corresponds to Allele A and are reported on the Plus strand.
Allele 2 – Plus
Allele 2 corresponds to Allele B and are reported on the Plus strand.
GC Score
Quality metric calculated for each genotype (data point), and ranges from 0 to 1.
GT Score
The SNP cluster quality. Score for a SNP from the GenTrain clustering algorithm.
Log R Ratio
Base-2 log of the normalized R value over the expected R value for the theta value (interpolated from the R-values of the clusters). For loci categorized as intensity only; the value is adjusted so that the expected R value is the mean of the cluster.
B Allele Freq
B allele frequency for this sample as interpolated from known B allele frequencies of 3 canonical clusters: 0, 0.5 and 1 if it is equal to or greater than the theta mean of the BB cluster. B Allele Freq is between 0 and 1, or set to NaN for loci categorized as intensity only.
Chr
Chromosome containing the SNP.
Position
SNP chromosomal position.
Note: Analyses on products with large numbers of loci (>1 Million) and large numbers of samples (>100) yield a large (50+ Gigabyte) Final Report that are difficult to download and review. It’s recommended to create analysis configurations that do not produce this report if large batches are desired.
For more information on interpreting DNA strand and allele information, see Illumina Knowledge article How to interpret DNA strand and allele information for Infinium genotyping array data.
Locus Summary
DRAGEN Array Cloud produces a Locus Summary (locus_summary.csv) per analysis batch similar to the one available in GenomeStudio. It contains the following details per locus:
Locus_Name
Locus name from the manifest file.
Illumicode_Name
Locus ID from the manifest file.
#No_Calls
Number of loci with GenCall scores below the call region threshold.
#Calls
Number of loci with GenCall scores above the call region threshold.
Call_Freq
Call frequency or call rate calculated as follows: #Calls/(#No_Calls + #Calls)
A/A_Freq
Frequency of homozygote allele A calls.
A/B_Freq
Frequency of heterozygote calls.
B/B_Freq
Frequency of homozygote allele B calls.
Minor_Freq
Frequency of the minor allele.
Gentrain_Score
Quality score for samples clustered for this locus.
50%_GC_Score
50th percentile GenCall score for all samples.
10%_GC_Score
10th percentile GenCall score for all samples.
Het_Excess_Freq
Heterozygote excess frequency, calculated as (Observed -Expected)/Expected for the heterozygote class. If $f_{ab}$ is the heterozygote frequency observed at a locus, and p and q are the major and minor allele frequencies, then het excess calculation is the following: $(f_{ab} - 2pq)/(2pq + \varepsilon)$
ChiTest_P100
Hardy-Weinberg p-value estimate calculated using genotype frequency. The value is calculated with 1 degree of freedom and is normalized to 100 individuals.
Cluster_Sep
Cluster separation score.
AA_T_Mean
Normalized theta angles mean for the AA genotype.
AA_T_Std
Normalized theta angles standard deviation for the AA genotype.
AB_T_Mean
Normalized theta angles mean for the AB genotype.
AB_T_Std
Standard deviation of the normalized theta angles for the AB genotype.
BB_T_Mean
Normalized theta angles mean for the BB genotypes.
BB_T_Std
Standard deviation of the normalized theta angles for the BB genotypes.
AA_R_Mean
Normalized R value mean for the AA genotypes.
AA_R_Std
Standard deviation of the normalized R value for the AA genotypes.
AB_R_Mean
Normalized R value mean for the AB genotypes.
AB_R_Std
Standard deviation of the normalized R value for the AB genotypes.
BB_R_Mean
Normalized R value mean for the BB genotypes.
BB_R_Std
Standard deviation of the normalized R value for the BB genotypes.
Plus/Minus Strand
Designated "+" or "-" with respect to the reference genome strand. "U" designates unknown.
CN Summary File
The sample summary contains per sample key stats for each sample in a batch that contains the following details per sample:
Sample ID
Sample Name
Sample Folder
Copy Number Batch File
The copy number batch summary file (cn_batch_summary.csv) shows the total copy number gain, loss, and neutral (CN=2) values for each target region across all the samples in the analysis.
Example copy number batch summary file content:
Target Region,Total CN gain,Total CN loss,Total CN neutral
CYP2A6.exon.1,0,1,47
CYP2A6.intron.7,0,1,47
CYP2D6.exon.9,2,4,42
CYP2D6.intron.2,7,2,39
CYP2D6.p5,13,2,33
CYP2E1,2,0,46
GSTM1,0,42,6
GSTT1,0,33,15
SULT1A1,0,0,48
UGT2B17,0,34,14
All Target Regions,24,119,337
Warning/Error Messages and Logs
The following scenarios result in a warning or error message:
Manifest file used to generate GTC is not the same as the manifest file used to generate the CN model.
FASTA files and FASTA index files do not match.
For the following scenarios, the software reports messages to the terminal output (as either a warning or an error):
Indel processing for GTC to VCF conversion failed.
The input folder does not contain the required input files.
An input file is corrupt.
Examples of such notifications can include the following:
Failed to normalize and gencall sample: {sample_id}, it will be skipped. Error: The given key '{loci_id}' was not present in the dictionary.
Warning
This generally occurs because of a mismatch between the manifest (bpm) and cluster file (egt) (i.e., the cluster file was generated via a different manifest). To remedy the issue, use the manifest and cluster files intended for use together.
Reference allele is not queried for locus: {identifier}
Warning
True reference allele does not match any alleles in the manifest. The error is common for MNVs and will be addressed in future versions of the software.
Skipping non-mapped locus: {identifier}
Warning
Locus has no chromosome position (usually 0) These loci may be used for quality purposes or CNV calling only.
Skipping intensity only locus: {identifier}
Warning
Similar to non-mapped loci, intensity only probes have applications outside creating variants for SNV VCFs such as CNV calling.
Skipping indel: {identifier}
Warning
Indel context (deletion/insertion) could not be determined.
Failed to process entry for record: {identifier}
Warning
Unable to determine reference allele for indel.
Incomplete match of source sequence to genome for indel: {identifier}
Warning
Indel not properly mapped to the reference genome.
Failed to combine genotypes due to ambiguity - exm1068284 (InfiniumII): TT, ilmnseq_rs1131690890_mnv (InfiniumII): AA, rs1131690890_mnv (InfiniumII): AA
Warning
Detailed information about a NoCall ("./.”) in the VCF as a result of combining multiple probes that assay the same variant with conflicting results. The example here is two probes with homozygous REF genotypes (AA) and one probe with homozygous ALT probe (TT)
Cluster file ({GTC.egt}) is not the same as CN Model Cluster file ({CN_Model.egt}).
Warning
Cluster file used to generated GTCs used for copy number calling is not the same as was used for the GTCs used during copy number training that created the input CN model. Though CNV model is robust to minor cluster file updates, CNV training should be considered when there are significant updates in the cluster file. To remove the warning, copy number training needs to be re-run with the new GTCs generated via the new cluster file during genotyping, a different CN model with the expected cluster file needs to be used, or different GTCs should be used for copy number calling that were generated using the same cluster file as was used during the generation of the input CN model.
{numPassingSamples} sample(s) passed QC.
Requires at least {minPassingSamples} samples to proceed.
Error
CNV calling is batch dependent and requires a certain number of samples with high-quality to make accurate calls. More high-quality samples need to be added to analysis batch to resolve error.
Invalid manifest file path {manifestPath}
Error
Application could not find manifest file provided or user error.
Failed to load cluster file: {e.Message}
Error
Corrupted file or unsupported version.
System.IO.EndOfStreamException: Unable to read beyond the end of the stream.
Error
Likely failure to read a GTC file, see this known issue for more details on root cause and a workaround
Star allele JSON File
The star allele JSON file is produced per sample. It contains the fields present in the star allele CSV file as well as additional meta data and annotations.
Fields included in the star allele JSON header are described below.
softwareVersion
DRAGEN Array software version, e.g. dragena 1.0.0.
genomeBuild
Genome build, e.g hg38.
starAlleleDatabaseSources
Public databases with versions used as the sources of the star allele definitions and population frequencies.
phenotypeDatabaseSources
Public databases with versions used as the sources of the star allele phenotypes.
mappingFile
The PGx database file used for the star allele calling.
pgxGuideline
The PGx guidelines used for metabolizer status/phenotype annotations, e.g. CPIC or DPWG
sampleId
Sentrix barcode and position of the sample.
locusAnnotations
The star allele call information.
Fields included in the star allele call (locusAnnotations) information are described below.
gene
The gene symbol.
callType
‘Star Allele’ or ‘Variant’ PGx calling type.
genotype
Most likely star allele or variant solution. If diploid, variant solutions have the format of Allele1/Allele2. More than one solution meeting threshold requirements can be reported. Multiple top solutions are separated by a semi-colon.
activityScore
Activity score annotation of the determined genotype of the gene determined based on public PGx guidelines CPIC or DPWG.
phenotypeDatabaseAnnotation
Metabolizer status and function annotations of the determined genotype of the gene based on lookup into public PGx guidelines CPIC or DPWG per user choice.
qualityScore
Quality score of the solution including the population frequency of PGx alleles. The score ranges from 0 to 1.
rawScore
Raw quality score of the solution without including the population frequency of PGx alleles. The score ranges from 0 to 1.
supportingVariants
All variants present in the array that support the star allele solution. The field provides an array (list) of supporting Variants.
Each supporting variant is listed with essential information extracted from the SNV VCF to assist with troubleshooting, including Chromosome (chrom), Location (pos), Reference allele (ref), Alternative allele (alt), Genotype (gt), GenCall score (gs), B-allele frequency (baf), the variant ID (id), and the associated star allele IDs (alleleIds).
candidateSolutions
The set of alternative star allele calling solutions, this is only relevant for genes of the ‘Star Allele’ call type.
missingVariantSites
All core variants that are not available (e.g. not on the array, or no calls in the SNV VCF) for star allele calling for this gene. For star alleles, the field provides an array (list) of variant "id" and impacted "alleleIds" pairs
allelesTested
Alleles that are covered by the star allele caller. The capability to call star alleles is also dependent on array content coverage and data quality. This field is defined by the array's content and will be the same across all samples.
Fields included in the candidateSolution section, only available for star allele call type, are described below.
rank
Rank of a single star allele solution for a gene. The top solution based on quality score is ranked as 1 with the alternative solutions ranked lower.
genotype
Star allele or variant solution. If diploid, variant solutions have the format of Allele1/Allele2.
activityScore
Activity score annotation of the determined genotype of the gene determined based on public PGx guidelines CPIC or DPWG.
phenotype
Metabolizer status and function annotations of the determined genotype of the gene based on lookup into public PGx guidelines CPIC or DPWG per user choice.
qualityScore
Quality score of the solution including the population frequency of PGx alleles. The score ranges from 0 to 1.
rawScore
Raw quality score of the solution without including the population frequency of PGx alleles. The score ranges from 0 to 1.
alleles
The composite alleles of the candidate genotype solution.
solutionLong
Long format solution for star alleles. The field has the following format: Structural Variant Type: Underlying Star allele.
An example of a long solution is: Complete: CYP2D64, Complete: CYP2D610, CYP2D668: CYP2D64 where there are two complete alleles that have CYP2D64 and CYP2D610 haplotypes and one CYP2D668 structural variant that has a CYP2D64 haplotype configuration.
supportingVariants
All variants present in the array that support the star allele solution. The field provides an array (list) of supporting Variants.
Each supporting variant is listed with essential information extracted from the SNV VCF to assist with troubleshooting, including Chromosome (chrom), Location (pos), Reference allele (ref), Alternative allele (alt), Genotype (gt), GenCall score (gs), B-allele frequency (baf), and the variant ID (id).
missingVariantSites
All variants not present in the array or not called in the SNV VCF file for the star allele solution. The field provides an array (list) of missing variants.
collapsedAlleles
Star alleles that cannot be distinguished from the solution star allele given the input array’s content. The field has the following format: Long Solution Star-Allele: (List of collapsed star alleles).
The most frequent star allele based on the population frequency of PGx alleles will be the star allele in the solution.
copyNumberRegions
Gene regions for the copy numbers listed in CopyNumberSolution.
copyNumberSolution
Estimated copy number for each gene region listed in CopyNumberRegions
Example of JSON file content:
Guidance on alternative star-allele results
Typically, the star allele solution with highest quality score is accepted as the final genotype (i.e. star allele diplotype) for the PGx locus. In rare cases, there are lower ranked star allele solutions with quality scores no less than 50% of the highest quality score, these lower ranked solutions are considered feasible and they are all listed in the genotype field of the locus annotation of the PGx gene in the PGx JSON file. Alternative solutions should also be considered if there are supporting variants for those solutions with low (less than 0.15) GS scores. The clustering of low GS scoring supporting variants should also be evaluated for cluster quality and any potential cluster shift.
Cytogenetics Annotation JSON File
DRAGEN Array produces one cytogenetics annotation JSON (*.json) per sample to report more sample-level, chromosome-level, and event-level metrics and annotations.
Example of JSON file content:
The fields in the annotation JSON for each sample are described as follows.
annotateDb
File name of the database annotation file.
softwareVersion
Version of DRAGEN Array used for the analysis.
referenceGenome
File name of the reference genome.
annotationType
Integer representing annotation methodology where 0=Constitutional and 1=Oncology.
genomeBuild
Genome build e.g. hg19, hg38.
databaseSources
Release versions of annotation data.
iscnVersion
Release date of ISCN formatting used.
sampleId
ID string assigned to the sample.
gcCorrect
Boolean indicating whether GC correction was enabled.
minDelProbes
Deletions must contain this many probes to be reported.
minDupProbes
Duplications must contain this many probes to be reported.
minLOHProbes
LOH variants must contain this many probes to be reported.
minDelSize
Minimum length filter for reporting deletions in kilobases (kb).
minDupSize
Minimum length filter for reporting a duplication in kb.
minLOHSize
Minimum length filter for reporting a loss-of-heterzygozity (LOH) variant in kb.
minQual
Minimum quality score filter for reporting a variant.
overallPloidy
Arithmetic mean of the ploidy across the genome. This value accounts for the length of all variant calls. The baseline ploidy value without any variants will differ by sex.
callRate
Frequency of expected calls i.e. #Calls/(#No_Calls + #Calls).
logRDev
Standard deviation of the Log R ratio values for all probes.
bafDev
Standard deviation of the B allele frequency values for each assigned genotype (AA/AB/BB).
numLOHOver1M
Count of LOH variants detected > 1 Mbp in length.
numLOHOver8M
Count of LOH variants detected > 8 Mbp in length.
totalSizeLOHOver1M
Cumulative length of all detected LOH variants > 1 Mbp in length.
copyNumberMedian
Length-normalized genome-wide median copy number value. Copy number values are assigned to contiguous segements of variable size in the genome by the algorithm. The length-weighted copy numbers of each variant are aggregated to calculate the overall median.
percentLOH
Percent of the genome comprised of LOH variants.
sexEstimate
Detected sex of the sample.
traditionalNomenclature
Simplified ISCN format designation for all detected variants in the sample.
microarrayNomenclature
ISCN format designation for all detected variants in the sample.
chromosomeAnnotations
Counts of each type of variant detected per chromosome, including mosaic calls.
locusAnnotations
Locus level statistics (see additional table for locus-level statistics).
The fields for each chromosome under the chromosomeAnnotations field of the Cyto annotation JSON are described below.
id
Chromosome name.
size
Chromosome size.
percentHet
Percent of probes in the chromosome called as heterzygous i.e. AB.
hasMosaicism
Boolean indicating presence of any mosaic variant on the chromosome.
lrrMedian
Median log R ratio value of the probes within the chromosome.
lrrDev
Standard deviation of the log R ratio values of the probes within the chromosome.
numLOHOver1M
Count of LOH variants detected > 1 Mbp in length.
numLOHOver8M
Count of LOH variants detected > 8 Mbp in length.
totalSizeLOHOver1M
Cumulative length of all detected LOH variants > 1 Mbp in length.
percentLOH
Percent of the genome comprised of LOH variants.
copyNumberMedian
Copy number median of the chromosome.
copyNumberMean
Copy number mean of the chromosome.
minLogRRatio
Minimum Log R ratio of the chromosome.
maxLogRRatio
Maximum Log R ratio of the chromosome.
medianMosaicFraction
Median mosaic fraction of mosaic events on the chromosome.
numberDel
Number of deletion events on the chromosome.
numberDup
Number of duplication events on the chromosome.
numberLOH
Number of LOH events on the chromosome.
numberMosaic
Number of mosaic events on the chromosome.
The fields within each variant (CNV/LOH event) under the locusAnnotations field of the Cyto annotation JSON are described below.
id
Unique variant ID containing variant type, chromosome and the start and end positions.
chrom
Chromosome of the variant.
start
Variant start position.
end
Variant end position.
callType
Variant class (DEL, DUP, LOH).
mosaicState
Boolean indicating whether the locus is a mosaic variant.
copyNumber
Copy number of the locus.
qualityScore
Phred-scaled score of the variant call quality.
size
Length of the variant.
effectiveSize
Gap-excluded length of the variant. A gap is defined when probe spacing is more than 150 times the median probe spacing. In that case, the gap is replaced with the median probe spacing.
probeCount
Number of probes contained in the called variant region.
percentHet
Percent of probes in the region call as heterzygous i.e. AB.
lrrMedian
Median log R ratio value of the probes within the variant.
lrrDev
Standard deviation of the log R ratio values of the probes within the variant.
bafDev
Standard deviation of the B allele frequency values of the probes within the variant.
startCytoBand
Cytoband in which the variant starts.
endCytoBand
Cytoband in which the variant ends.
traditionalNomenclature
Simplified ISCN format designation for the detected variant.
microarrayNomenclature
ISCN format designation for the detected variant.
geneCount
Count of annotated genes within the variant region.
genes
List of names of all annotated genes within the variant region.
The traditionalNomenclature field is used to describe individual and cumulative copy number variants at a coarse resolution. They show gains (dup) and losses (del) according to their chromosome number, arm (p or q), and band (e.g., p36.13). The microarrayNomenclature field follows the Comparative Genomic Hybridization or SNP array conventions. Values are prefixed with arr[] to indicate array data as the source, along with the genome build e.g. GRCh38. These data can more precisely describe the location (start_end in bp) and copy number (x1 for loss, x3 for gain, etc).
The genes list field for each variant include those with transcript coordinates that intersect with the described variant. The values included are a combination of HGNC gene symbols taken from NCBI RefSeq database (e.g. TBC1D3I), as well as the subset of Ensembl gene accessions unmatched to a gene symbol (ENSG00000278395).
TBI Index File
The TBI (TABIX) index file is associated with the bgzipped VCF files. It allows for data line lookup in VCF files for quick data retrieval. The format is a tab-delimited genome index file developed by Samtools as part of the HTSlib utilities. For more information, visit the Samtools website.
Methylation Control Probe Output File
The software produces a control probe output file ({BeadChipBarcode}_{Position}_ctrl.tsv.gz) per sample that includes the raw methylated and unmethylated values for each control probe.
Each control probe has an address, type, color channel, name, and probe ID. It also provides the raw signal for methylated green (MG), methylated red (MR), unmethylated green (UG) and unmethylated red (UR).
The file can help identify which probes are available on a given BeadChip.
Methylation CG Output File
The software produces a CG output file ({BeadChipBarcode}_{Position}_cgs.tsv.gz) per sample that includes beta values, m-values and detection p-values for each CG site.
Beta values measure methylation levels in a linear fashion for easy interpretation. Unmethylated probes are close to zero and methylated probes are close to 1.
M-values are a log transformed beta value which provides a more representative measure of methylation.
Detection p-values measure the likelihood that the signal is background noise. It is recommended that p-value >0.05 are excluded from analysis as they are likely background noise.
see High-throughput Infinium methylation array QC using DRAGEN Array Methylation QC software tech note for further detail on calculation of these metrics.
Methylation Sample QC Summary Files
The software produces methylation sample QC summary in .xlsx and .tsv file formats (sample_qc_summary.xlsx and sample_qc_summary.tsv) per analysis batch, which provides per sample QC data for all samples in the batch.
The QC summary provides details on 21 controls metrics (see tables below), which are computed in same way as in the BeadArray Controls Reporter software from Illumina. In addition, it provides average red and green raw and normalized signals, time of scanning, proportion of probes passing, overall sample pass/fail status, and the failure codes for control metrics that did not pass. The sample pass status is defined as the passing of all 21 control metrics. The QC summary .xlsx file further highlights failing parameters for easy viewing.
The QC summary files contain the following fields:
Sentrix_ID
12-digit BeadChip Barcode associated with the sample.
Sentrix_Position
Row and column on the BeadChip ie R01C01
Sample_ID
Optional field that can be indicated using IDAT Sample Sheet
User Defined Meta Data
Optional field(s) that can be indicated using IDAT Sample Sheet. Any number of fields indicated will appear in this output file.
restoration
The default threshold is 0.
If using the FFPE DNA Restore Kit, the restoration control identifies success of the FFPE restoration chemistry. Change the threshold from 0 to 1 if the FFPE DNA Restore Kit was used.
The green channel intensity is higher than Background. Therefore, the metric provided is the Green Channel Intensity/Background.
staining_green
staining_red
Staining controls are used to examine the efficiency of the staining step in both the red and green channels. These controls are independent of the hybridization and extension step.
The green channel shows a higher signal for biotin staining when compared to biotin background, whereas the red channel shows higher signal for DNP staining when compared to DNP background.
The metric provided for green is the (Biotin High value)/ (Biotin Bkg) and the metric provided for red is (DNP High value)/(DNP Bkg value)
The default threshold is 5. This threshold can be increased on some scanners.
extension_green
extension_red
Extension controls test the extension efficiency of A, T, C, and G nucleotides from a hairpin probe, and are therefore sample independent.
In the green channel, the lowest intensity for C or G is always greater than the highest intensity for A or T.
The metric provided is the (lowest of the C or G intensity)/ (highest of A or T extension) for a single sample.
The default threshold is 5. This threshold can be increased on some scanners.
hybridization_high_medium
hybridization_medium_low
Hybridization controls test the overall performance of the Infinium Assay using synthetic targets instead of amplified DNA. These synthetic targets complement the sequence on the array, allowing the probe to extend on the synthetic target as a template. Synthetic targets are present in the Hybridization Buffer at 3 levels, monitoring the response from high-concentration (5 pM), medium concentration (1 pM), and low concentration (0.2 pM) targets. All bead type IDs result in signals with various intensities, corresponding to the concentrations of the initial synthetic targets.
The value for high concentration is always higher than medium and the value for medium concentration is always higher than low.
The metric provided is the value of high/medium and the value of medium/low.
The default thresholds are 1. Do not change the default threshold.
target_removal1
target_removal2

