1 of 9

Analysis Methods

The software processes sequencing data to perform quality control, detect variants, determine tumor mutational burden (TMB), microsatellite instability (MSI) status, and genomic instability score (GIS), and report results. The following sections describe the analysis methods used in DRAGEN TruSight Oncology 500 Analysis Software.

DRAGEN TruSight Oncology 500 Analysis Software uses the following workflows to analyze sequencing data.

FASTQ Generation
DNA Analysis
- DNA Alignment and Realignment
- Read Collapsing
- Indel Realignment and Read Stitching
- Small Variant Calling
RNA Analysis
- Downsampling
- Read Trimming
- Alignment
Quality Control
- Run QC
- DNA Sample QC
- RNA Sample QC

FASTQ Generation

Sequencing data stored in BCL format are demultiplexed through a process that uses the index sequences unique to each sample to assign clusters to the library from which they originated. Each cluster contains two indexes (i7 and i5 sequences, one at each end of the library fragment). The combination of those index sequences are used to demultiplex the pooled libraries.

After demultiplexing, this process generates FASTQ files, which contain the sequencing reads for each individual sample library and the associated quality scores for each base call, excluding reads from any clusters that did not pass filter.

DNA Analysis Methods

DNA Alignment and Error Correction

DNA alignment and error correction involves aligning sequencing reads derived from DNA libraries to a reference genome and correcting errors in the sequencing reads prior to variant calling.

DRAGEN unique molecular identifier (UMI) error correction comprises three main steps:

DRAGEN UMI uses its HW accelerated mapper (based on a hash table implementation) to align DNA sequences in FASTQ files to the hg19 reference genome. These alignments are not written to a BAM.
The raw alignments are processed to remove errors, including errors introduced during FFPE preservation, PCR amplification, and sequencing. Reads from the same original DNA molecule are tagged with the same UMI during library preparation. The UMI allows DRAGEN to compare related reads, remove outlier signals, and collapse multiple reads into a single high-quality sequence. Read collapsing adds the following BAM tags:
- RX/XU—UMI.
- XV—Number of reads in the family.
DRAGEN performs a final alignment step on the UMI-collapsed reads. These final alignments are then written to a BAM file and a corresponding BAM index file is created.

DRAGEN continues to use these final alignments as input for gene amplification (copy number) calling, small variant calling (SNV, indel, MNV, delin), microsatellite instability (MSI) status determination, and DNA library quality control.

Small Variant Calling and Filtering

DRAGEN supports calling SNVs, indels, MNVs, and delins in tumor-only samples by using mapped and aligned DNA reads from a tumor sample as input. Variants are detected via both column wise pileup analysis and local de novo assembly of haplotypes. The de novo haplotypes allow the detection of much larger insertions and deletions than possible through column wise pileup analysis only. DRAGEN insertions and deletions are validated with lengths of at least 0–25 bp and more than 25 bp can be supported. In addition, DRAGEN also uses the de novo assembly to detect SNVs, insertions, and deletions that are co-phased and part of the same haplotypes. Any such co-phased variants that are within a window of 15 bp can then be reassembled into complex variants (MNVs and delins). The tumor-only pipeline produces a VCF file containing both germline and somatic variants that can be further analyzed to identify tumor mutations. Variant calling extends ± 10 bp into introns; details of the regions covered can be found in the assay manifest file. The pipeline makes no ploidy assumptions, enabling detection of low-frequency alleles.

DRAGEN small variant calling includes the following steps:

Detects regions with sufficient read coverage (callable regions).
Detects regions where the reads deviate from the reference and there is a possibility of a germline or somatic call (active regions).
Assembles de novo graph haplotypes are assembled from reads (haplotype assembly).

Copy Number Variant Calling

The DRAGEN copy number variant caller performs amplification, reference, and deletion calling for CNV targets within the assay. It counts the coverage of each target interval on the panel, uses a preprocessed panel of normal samples to normalize target counts, corrects for GC coverage bias, and calculates scores of a CNV event from observed coverage and makes copy number calls.

Exon-Level Copy Number Variant Calling

The BRCA large rearrangement step generates segmentation of the BRCA1 and BRCA2 genes for exon-level CNV detection from the BAM file. Using the same method as CNV calling, the large rearrangement component counts coverage of each target interval of the panel, performs normalization, and calculates the fold change values for each probe across the BRCA genes. Normalization includes GC bias correction, sequencing depth, and probe efficiency using a collection of normal FFPE and genomic DNA samples. Initial segmentation is performed for each gene with circular binary segmentation. The merging of segments is then determined by amplitude, noise, and variance at adjacent segments using thresholds established with in silico data. A large rearrangement is reported for genes with more than one segment. Coordinates of the exon-level CNV and the log2 mean fold change for each of the BRCA gene segments are found in the *_DragenExonCNV.json file.

Annotation

The Illumina Annotation Engine performs annotation of small variants, CNVs, and exon-level CNVs. The inputs are gVCF files and the outputs are annotated JSON files.

The Illumina Annotation Engine processes each variant entry and annotates with available information from databases such as dbSNP, gnomAD genome and exome, 1000 genomes, ClinVar, COSMIC, RefSeq, and Ensembl. The header includes version information and general details. Each annotated variant is included as a nested dictionary structure in separate lines following the header.

The following table shows version information for each annotation database:

Database

Version

Tumor Mutational Burden

DRAGEN is used to compute tumor mutational burden (TMB) in coding regions where there is sufficient coverage.

The following variants are excluded from the TMB calculation:

Non-PASS variants.
Mitochondrial variants.
MNVs.
Variants that do not meet a minimum depth threshold (50).

Variants with a population allele count ≥ 10 that are observed in either the 1000 Genomes or gnomAD databases are marked as germline. MNVs, which do not count towards TMB, may be marked as germline when all their component small variants are marked as germline. The proxy filter scans the variants surrounding a specific variant and identifies those variants with similar variant allele frequencies (VAF). If the majority of surrounding variants of similar VAF are germline, then the variant is also marked as germline.

The formula for TMB calculation is:

Outputs are captured in a _TMB_Trace.tsv file that contains information on variants used in the TMB calculation and a .tmb.json file that contains the TMB score calculation and configuration details.

Microsatellite Instability Status

DRAGEN can determine the MSI status of a sample. It uses a normal reference file, which was created from a set of normal samples. Normal reference files were generated by tabulating read counts for each microsatellite site. The normal file contains the read count distribution for each microsatellite site.

MSI calling is assessed on a predefined list of 130 A and T repeats. The first step in calculating the MSI score is determining how many sites are assessable. A site is considered assessable if it has at least 60 spanning reads. A spanning read is defined as one that extends 5 bp before and after the repeat.

Once assessable sites are identified, the distribution of repeat lengths is compared to the panel of normals. A site is classified as unstable if:

Jensen-Shannon distance ≥ 0.1, and
P-value ≤ 0.01.

After all sites are evaluated, DRAGEN reports:

The total number of sites assessed
The count of unstable sites
The percentage of unstable sites across the sample

Finally, the MSI score is calculated as:

Genomic Instability Score

Requires HRD add-on assay

Genomic instability score (GIS) is a whole genome signature for homologous recombination deficiency. The GIS is composed of the sum of three components: loss of heterozygosity, telomeric allele imbalance, and large-scale state transition. These components are estimated using the GIS algorithm contracted from Myriad Genetics, which uses an input of the b-allele frequency and coverage across a genome-wide single nucleotide panel. A panel of normal samples is used for both bias reduction and normalization prior to GIS estimation. Final GIS results can be found in the *.gis.json file.

Contamination Detection

The contamination analysis step detects foreign human DNA contamination using the SNP error file and pileup file that are generated during the small variant calling and the TMB trace file. The software determines whether a sample has foreign DNA using the contamination score. In contaminated samples, the variant allele frequencies in SNPs shift from the expected values of 0%, 50%, or 100%. The algorithm collects all positions that overlap with common SNPs that have variant allele frequencies of < 25% or > 75%. Then, the algorithm computes the likelihood that the positions are an error or a real mutation. The contamination score is the sum of all the log likelihood scores across the predefined SNP positions with minor allele frequency < 25% in the sample and are not likely due to CNV events.

The larger the contamination score, the more likely there is foreign DNA contamination. A sample is considered to be contaminated if the contamination score is above predefined quality threshold. The contamination score was found to be high in samples with highly rearranged genomes or HRD samples. 1% of HRD samples found to be above the threshold with no evidence for actual contamination.

Tumor fraction

Tumor fraction is calculated as described in the User Guide, section “HRD Metrics Report” and leverages the Myriad Genetics algorithm. Tumor fraction is output in the Logs_Intermediates/Gis/SAMPLE/SAMPLE.gis.json and Combined Variant Output file.

Ploidy

Ploidy is calculated as described in the User Guide, section “HRD Metrics Report” and leverages the Myriad Genetics algorithm. Ploidy is output in the in the Logs_Intermediates/Gis/SAMPLE/SAMPLE.gis.json and Combined Variant Output file.

Absolute Copy Number (Beta)

This is a beta feature. Beta feature results are included in the Combined Variant Output file and other files. However, disclaimers that the results are generated by beta features are only provided in the Combined Variant Output file. Requires HRD add-on assay.

Absolute copy numbers are calculated by leveraging the Myriad Genetics algorithm. The algorithm segments the entire genome using the HRD panel and provides an A and B allele estimate for each segment. After the TSO 500 pipeline determines CNV calls (using the TSO 500 panel), the segment covering the gene is identified, and the A and B allele numbers of the segment overlapping the gene are reported. If the gene is within 300 kbases from the segment boundary, the estimate is unreliable and “-1” is output. Absolute copy numbers are output in the Logs_Intermediates/Gis/SAMPLE/SAMPLE.abcn_annotated.vcf, Logs_Intermediates/Gis/SAMPLE/SAMPLE.abcn_genes.tsv and Combined Variant Output file.

Gene-Level Loss of Heterozygosity (Beta)

Gene-level loss of heterozygosity is calculated based on the minor copy number reported in the abcn_annotated.vc f. If the minor copy number is 0 then the gene is assumed to have a loss of heterozygosity. Gene-level loss of heterozygosity is output in the Logs_Intermediates/Gis/SAMPLE/SAMPLE.abcn_genes.tsv and Combined Variant Output file.

Block List

The block list represents high noise regions in the panel where false positive variant calls are likely produced. As a result, all positions in the gVCF are marked as Filter=excluded_regions to indicate variant call results are not reliable in such regions.

The block list includes the following genes:

HLA A
HLA B
HLA C
KMT2B
KMT2C
KMT2D
chrY
Any position with VAF 1% occurrence in six or more of the 60 baseline samples.

RNA Analysis Methods

Refer to RNA Output for more information.

Downsampling

Each sample is downsampled to 30 million RNA reads. This number represents the total number of single reads (eg, R1 + R2, from all lanes). When using the recommended sequencing configurations or plexity, the samples can have fewer reads than the downsampling limit. In these cases, the FASTQ files are left as-is.

Read Trimming

Reads are trimmed to 76 base pairs for further processing.

RNA Alignment and Fusion Detection

RNA alignment and fusion detection uses trimmed reads in FASTQ format as input. The outputs include a BAM file that contains duplicate-marked read alignments, an SJ.out.tab file that contains unannotated splice junctions, and a CSV file that contains fusion candidates.

DRAGEN aligns RNA reads in a transcript-aware mode using the human hg19 genome containing unplaced contigs (ie, chrUn_gl regions) and uses GENCODEv19 transcript annotations to identify splice sites. DRAGEN identifies and marks duplicate read alignments using start and end coordinates of alignments, which are adjusted for soft clipped reads.

Fusion and splice variant calling only use deduped fragments to score variants. DRAGEN identifies fusion candidates using chimeric split read alignments (pairs of primary and supplementary alignments) against multiple genes. DRAGEN scores and filters reads based on the various features of each candidate such as the number of supporting reads, mapping quality of supporting reads, and sequence homology between parent genes.

The DRAGEN RNA Fusion caller identifies gene fusions by searching for chimeric reads spanning two distinct parent genes. Based on the chimeric reads, DRAGEN first creates a list of fusion candidates, then scores the candidates to report the list of high confidence fusion calls from the candidate pool.

DRAGEN RNA Fusion caller performs the following steps:

Generates fusion candidate generation based on split read alignment.
Recruits additional evidence from fusion supporting discordant read pairs and soft-clipped reads.
Computes fusion candidate features such as gene coverage, read mapping quality, alternate allele frequency, gene homology, alignment anchor length, and breakpoint distance from exon boundary.

Splice Variant Calling

RNA splice variant calling is performed for RNA sample libraries. Candidate splice variants (junctions) from RNA Alignment are compared against a database of known transcripts and a splice variant baseline of non-tumor junctions generated from a set of normal FFPE samples from different tissue types. Any splice variants that match the database or baseline are filtered out unless they are in a set of junctions with known oncological function. If there is sufficient read support, the candidate splice variant is kept. This process also identifies candidate RNA fusions.

RNA Fusion Merging

Fusions identified during RNA fusion calling are merged with fusions from proximal genes identified during RNA splice variant calling. These are then annotated with gene symbols or names with respect to a static database of transcripts (GENCODE Release 19). The result of this process is a set of fusion calls that are eligible for reporting

RNA Splice Variant Annotation

The Illumina Annotation Engine annotates detected RNA splice variant calls with transcript-level changes (eg, affected exons in the transcript of a gene) with respect to RefSeq. This RefSeq database is the same RefSeq database used by the small variant annotation process.

Quality Control

The software calculates several quality control metrics for runs and samples.

These metrics and guidelines apply to DRAGEN TSO 500 v2.1 and above.

Run QC

The Run Metrics section of the metrics output report provides sequencing run quality metrics along with suggested values to determine if they are within an acceptable range. The overall percentage of reads passing filter is compared to a minimum threshold. For Read 1 and Read 2, the average percentage of bases ≥ Q30, which gives a prediction of the probability of an incorrect base call (Q‑score), are also compared to a minimum threshold. The following tables show run metric and quality threshold information for different systems.

The values in the Run Metrics section are listed as NA in the following situations:

If the analysis was started from FASTQ files.
If the analysis was started from BCL files and the InterOp files are missing or corrupt.

NextSeq 500/550 or NextSeq 550Dx (RUO)

Metric

Description

Recommended Guideline Quality Threshold

Variant Class

NovaSeq 6000 or NovaSeq 6000Dx (RUO)

Metric

Description

Recommended Guideline Quality Threshold

Variant Class

NextSeq 1000/2000

Metric

Description

Recommended Guideline Quality Threshold

Variant Class

NovaSeq X

Metric

Description

Recommended Guideline Quality Threshold

Variant Class

DNA Sample QC

DRAGEN TruSight Oncology 500 uses QC metrics to assess the validity of analysis for DNA libraries that pass contamination quality control. If the library fails one or more quality metrics, then the corresponding variant type or biomarker is not reported, and the associated QC category in the report header displays FAIL. Additionally, a companion diagnostic result may not be available if it relies on QC passing for one or more of the following QC categories.

DNA library QC results are available in the MetricsOutput.tsv file.

Metric

Description

Recommended Guideline Quality Threshold

Variant Class

RNA Sample QC

The input for RNA Library QC is RNA alignment. Metrics and guideline thresholds can be found in the MetricsOutput.tsv file.

Metric

Description

Recommended Guideline Quality Threshold

Variant Class

*To avoid failing RNA samples unnecessarily, Illumina does not recommend a universal threshold to determine RNA sample quality. RNA expression varies significantly across tissue types and a small panel size (55 genes), which makes normalization challenging. Tissue-specific thresholds could be considered for normalization.

DNA Expanded Metrics

DNA expanded metrics are provided for information only. They can be informative for troubleshooting but are provided without explicit specification limits and are not directly used for sample quality control. For additional guidance, contact Illumina Technical Support.

Metric

Description

Troubleshooting

RNA Expanded Metrics

RNA expanded metrics are provided for information only. They can be informative for troubleshooting but are provided without explicit specification limits and are not directly used for sample quality control. For additional guidance, contact Illumina Technical Support.

Metric

Description

Units

PCT_CHIMERIC_READS

Percentage of reads that are aligned as two segments which map to nonconsecutive regions in the genome.

PCT_ON_TARGET_READS

Percentage of reads that cross any part of the target region versus total reads. A read that partially maps to a target region is counted as on target.

Contamination

The contamination score evaluates presence of sample-to-sample contamination. The algorithm uses common germline SNPs in the homozygous state expected to have variant allele frequencies (VAF) at 0% and 100%. In contaminated samples, the VAFs shift away from the expected values allowing the detection of sample-to-sample contamination.

The contamination score can detect sample-to-sample contamination greater than or equal to 2% (more than 2% of DNA input is coming from the contaminant)

DNA Analysis Methods

DNA Alignment and Error Correction

DNA alignment and error correction involves aligning sequencing reads derived from DNA libraries to a reference genome and correcting errors in the sequencing reads prior to variant calling.

DRAGEN unique molecular identifier (UMI) error correction comprises three main steps:

DRAGEN UMI uses its HW accelerated mapper (based on a hash table implementation) to align DNA sequences in FASTQ files to the hg19 reference genome. These alignments are not written to a BAM.
The raw alignments are processed to remove errors, including errors introduced during FFPE preservation, PCR amplification, and sequencing. Reads from the same original DNA molecule are tagged with the same UMI during library preparation. The UMI allows DRAGEN to compare related reads, remove outlier signals, and collapse multiple reads into a single high-quality sequence. Read collapsing adds the following BAM tags:
- RX/XU—UMI.
- XV—Number of reads in the family.
DRAGEN performs a final alignment step on the UMI-collapsed reads. These final alignments are then written to a BAM file and a corresponding BAM index file is created.

Small Variant Calling and Filtering

DRAGEN small variant calling includes the following steps:

Detects regions with sufficient read coverage (callable regions).
Detects regions where the reads deviate from the reference and there is a possibility of a germline or somatic call (active regions).
Assembles de novo graph haplotypes are assembled from reads (haplotype assembly).

Copy Number Variant Calling

Exon-Level Copy Number Variant Calling

Annotation

The Illumina Annotation Engine performs annotation of small variants, CNVs, and exon-level CNVs. The inputs are gVCF files and the outputs are annotated JSON files.

The following table shows version information for each annotation database:

Database

Version

Tumor Mutational Burden

DRAGEN is used to compute tumor mutational burden (TMB) in coding regions where there is sufficient coverage.

The following variants are excluded from the TMB calculation:

Non-PASS variants.
Mitochondrial variants.
MNVs.
Variants that do not meet a minimum depth threshold (50).

The formula for TMB calculation is:

Microsatellite Instability Status

Once assessable sites are identified, the distribution of repeat lengths is compared to the panel of normals. A site is classified as unstable if:

Jensen-Shannon distance ≥ 0.1, and
P-value ≤ 0.01.

After all sites are evaluated, DRAGEN reports:

The total number of sites assessed
The count of unstable sites
The percentage of unstable sites across the sample

Finally, the MSI score is calculated as:

Genomic Instability Score

Requires HRD add-on assay

Analysis Methods

FASTQ Generation

DNA Analysis Methods

hashtagDNA Alignment and Error Correction

hashtagSmall Variant Calling and Filtering

hashtagCopy Number Variant Calling

hashtagExon-Level Copy Number Variant Calling

hashtagAnnotation

hashtagTumor Mutational Burden

hashtagMicrosatellite Instability Status

hashtagGenomic Instability Score

hashtagContamination Detection

hashtagTumor fraction

hashtagPloidy

hashtagAbsolute Copy Number (Beta)

hashtagGene-Level Loss of Heterozygosity (Beta)

Block List

RNA Analysis Methods

hashtagDownsampling

hashtagRead Trimming

hashtagRNA Alignment and Fusion Detection

hashtagSplice Variant Calling

hashtagRNA Fusion Merging

hashtagRNA Splice Variant Annotation

Quality Control

hashtagRun QC

hashtagNextSeq 500/550 or NextSeq 550Dx (RUO)

hashtagNovaSeq 6000 or NovaSeq 6000Dx (RUO)

hashtagNextSeq 1000/2000

hashtagNovaSeq X

hashtagDNA Sample QC

hashtagRNA Sample QC

DNA Expanded Metrics

RNA Expanded Metrics

Contamination

FASTQ Generation

Block List

RNA Expanded Metrics

DNA Expanded Metrics

RNA Analysis Methods

hashtagDownsampling

hashtagRead Trimming

hashtagRNA Alignment and Fusion Detection

hashtagSplice Variant Calling

hashtagRNA Fusion Merging

hashtagRNA Splice Variant Annotation

Analysis Methods

Quality Control

hashtagRun QC

hashtagNextSeq 500/550 or NextSeq 550Dx (RUO)

hashtagNovaSeq 6000 or NovaSeq 6000Dx (RUO)

hashtagNextSeq 1000/2000

hashtagNovaSeq X

hashtagDNA Sample QC

hashtagRNA Sample QC

DNA Analysis Methods

hashtagDNA Alignment and Error Correction

hashtagSmall Variant Calling and Filtering

hashtagCopy Number Variant Calling

hashtagExon-Level Copy Number Variant Calling

hashtagAnnotation

hashtagTumor Mutational Burden

hashtagMicrosatellite Instability Status

hashtagGenomic Instability Score

hashtagContamination Detection

hashtagTumor fraction

hashtagPloidy

hashtagAbsolute Copy Number (Beta)

hashtagGene-Level Loss of Heterozygosity (Beta)

Contamination

hashtagContamination Score Interpretation

hashtagHow to build a VAF plot for visual examination

DNA Alignment and Error Correction

Small Variant Calling and Filtering

Copy Number Variant Calling

Exon-Level Copy Number Variant Calling

Annotation

Tumor Mutational Burden

Microsatellite Instability Status

Genomic Instability Score

Contamination Detection

Tumor fraction

Ploidy

Absolute Copy Number (Beta)

Gene-Level Loss of Heterozygosity (Beta)

Downsampling

Read Trimming

RNA Alignment and Fusion Detection

Splice Variant Calling

RNA Fusion Merging

RNA Splice Variant Annotation

Run QC

NextSeq 500/550 or NextSeq 550Dx (RUO)

NovaSeq 6000 or NovaSeq 6000Dx (RUO)

NextSeq 1000/2000

NovaSeq X

DNA Sample QC

RNA Sample QC

Downsampling

Read Trimming

RNA Alignment and Fusion Detection

Splice Variant Calling

RNA Fusion Merging

RNA Splice Variant Annotation

Run QC

NextSeq 500/550 or NextSeq 550Dx (RUO)

NovaSeq 6000 or NovaSeq 6000Dx (RUO)

NextSeq 1000/2000

NovaSeq X

DNA Sample QC

RNA Sample QC

DNA Alignment and Error Correction

Small Variant Calling and Filtering

Copy Number Variant Calling

Exon-Level Copy Number Variant Calling

Annotation

Tumor Mutational Burden

Microsatellite Instability Status

Genomic Instability Score

Contamination Detection

Tumor fraction

Ploidy

Absolute Copy Number (Beta)

Gene-Level Loss of Heterozygosity (Beta)

Contamination Score Interpretation

How to build a VAF plot for visual examination