> For the complete documentation index, see [llms.txt](https://help.connected.illumina.com/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://help.connected.illumina.com/annotation/v4.0/data-sources/splice-ai.md).

# SpliceAI

### Overview

SpliceAI is an AI annotation model developed by the Illumina Artificial Intelligence Lab to predict splice effects.

The model evaluates probability that a variant disrupts or creates acceptor and donor splice sites. Higher delta scores suggest a stronger predicted effect on RNA splicing.

Splicing defects are a major contributor to human disease, particularly in rare disease and oncology, where many pathogenic variants occur in non-coding regions.

DRAGEN Annotation uses SpliceAI 2.0 for GRCh38 and SpliceAI 1.3 for GRCh37.

For more details, refer to:

{% hint style="info" %}
**Publication**

Jaganathan, et al. Predicting splicing from primary sequence with deep learning. *Cell* (2019). <https://doi.org/10.1016/j.cell.2018.12.015>
{% endhint %}

{% hint style="warning" %}
**Professional data source**

This data source requires a Professional license. Contact `annotation_support@illumina.com` to request access.
{% endhint %}

## SpliceAI 2.0 (GRCh38)

SpliceAI 2.0 is built in two steps. First, the original per-gene files are preprocessed into per-chromosome TSVs. Those TSVs are then converted into the supplementary annotation files (`.esa`) that DRAGEN Annotation loads. See [SAUtils](/annotation/v4.0/utilities/sautils.md) for how supplementary files are produced.

### Preprocessing

The original SpliceAI 2.0 release is a set of per-gene gzipped TSVs (separate SNV and indel files) plus a gene manifest. Those files were grouped by chromosome, SNV and indel rows were merged, rows were sorted by position and alleles, and delta scores were rounded to 2 decimal places. The result is one gzipped TSV per chromosome.

Each preprocessed TSV has 23 columns: variant keys (`chrom`, `pos`, `ref`, `alt`, `strand`), two sets of donor/acceptor gain and loss scores and distances (`*_0` and `*_1`), `gene_id`, and `variant_type`.

### Parsing

#### TSV File

The snippet below includes SNVs, an insertion (`C`/`CGTT`), and a deletion (`AG`/`A`) from the preprocessed GRCh38 files:

```scss
chrom	pos	ref	alt	gene_id	donor_gain_delta_score_0	donor_loss_delta_score_0	acceptor_gain_delta_score_0	acceptor_loss_delta_score_0	donor_gain_dist_0	donor_loss_dist_0	acceptor_gain_dist_0	acceptor_loss_dist_0
chr21	10521577	G	T	ENSG00000274391	0.10	0.02	0.00	0.00	-1	118	-22103	42
chr21	10521620	T	G	ENSG00000274391	0.07	0.02	0.11	0.01	0	75	-1	-22146
chr21	10521559	C	CGTT	ENSG00000274391	0.01	0.01	0.01	0.05	19607	5854	6727	5795
chr21	10521698	AG	A	ENSG00000274391	0.11	0.15	0.04	0.03	347	-3	-7577	-79
```

The header is required. Column order does not matter. Of the 23 preprocessed columns, DRAGEN Annotation reads these 13:

* `chrom`
* `pos`
* `ref`
* `alt`
* `gene_id`
* `donor_gain_delta_score_0`
* `donor_loss_delta_score_0`
* `acceptor_gain_delta_score_0`
* `acceptor_loss_delta_score_0`
* `donor_gain_dist_0`
* `donor_loss_dist_0`
* `acceptor_gain_dist_0`
* `acceptor_loss_dist_0`

These columns are not used:

* `strand`
* all `*_1` score and distance columns
* `variant_type`

#### Filtering

Low delta scores are treated as uninformative. During ESA generation, each of the four `*_0` delta scores below `0.05` is dropped together with its paired distance. If all four selected scores are below `0.05`, the variant is omitted.

This 0.05 cutoff is how the default `2.0_filtered` files are produced. It sits below the published low / likely benign band (`x < 0.1`) described under [Interpreting scores](#interpreting-scores).

#### Left-shiftable indels

Indels that can be left-shifted — the same event can be represented at a more 5′ position — are skipped. Annotation matches left-aligned alleles, so those records would not join. Remaining alleles are trimmed (shared prefix and suffix removed) before they are stored.

This applies to both `2.0` and `2.0_filtered`.

### Available files

GRCh38 ships two SpliceAI 2.0 versions:

| Version                  | Scores         | Size    |
| ------------------------ | -------------- | ------- |
| `2.0_filtered` (default) | threshold 0.05 | 625 MB  |
| `2.0`                    | unfiltered     | 53.5 GB |

`2.0_filtered` is the default so most users can download a much smaller file while still keeping scores at or above 0.05. Use `2.0` when you need every precomputed score.

### Using the full 2.0 files

The default run config uses `2.0_filtered`. To annotate with the full scores, change the SpliceAI small-variant version in the run config under your data folder (typically `<data.directory>/runConfigs/`) from `2.0_filtered` to `2.0`:

```json
"spliceAI": {
  "SmallVariant": "2.0_filtered"
}
```

```json
"spliceAI": {
  "SmallVariant": "2.0"
}
```

Then run the downloader and start annotating. See [Command Line Parameters](/annotation/v4.0/software-functionality/command-line-parameters.md) for download. You can also override the SpliceAI small-variant version on the command line instead of editing the file.

### JSON output

{% hint style="info" %}
**Gene ID**

SpliceAI 2.0 reports `geneId` as the Ensembl gene ID from the source data. This is different from SpliceAI 1.3, which reported `hgnc` (HGNC gene symbol).
{% endhint %}

```json
"spliceAI": [
  {
    "geneId": "ENSG00000292344",
    "acceptorGainScore0": 0.33,
    "acceptorGainDistance0": 530,
    "acceptorLossScore0": 0.44,
    "acceptorLossDistance0": 8,
    "donorGainScore0": 0.11,
    "donorGainDistance0": 73,
    "donorLossScore0": 0.22,
    "donorLossDistance0": -940
  }
]
```

<table><thead><tr><th width="199.54296875">Field</th><th width="140.0703125">Type</th><th>Notes</th></tr></thead><tbody><tr><td>geneId</td><td>string</td><td>Ensembl gene ID</td></tr><tr><td>acceptorGainDistance0</td><td>int</td><td>± bp from current position</td></tr><tr><td>acceptorGainScore0</td><td>float</td><td>range: 0 - 1.0. up to 2 decimal places</td></tr><tr><td>acceptorLossDistance0</td><td>int</td><td>± bp from current position</td></tr><tr><td>acceptorLossScore0</td><td>float</td><td>range: 0 - 1.0. up to 2 decimal places</td></tr><tr><td>donorGainDistance0</td><td>int</td><td>± bp from current position</td></tr><tr><td>donorGainScore0</td><td>float</td><td>range: 0 - 1.0. up to 2 decimal places</td></tr><tr><td>donorLossDistance0</td><td>int</td><td>± bp from current position</td></tr><tr><td>donorLossScore0</td><td>float</td><td>range: 0 - 1.0. up to 2 decimal places</td></tr></tbody></table>

Score and distance keys are omitted when that score is below `0.05` or was not present.

## SpliceAI 1.3 (GRCh37)

### Parsing

#### VCF File

```scss
##fileformat=VCFv4.0
##assembly=GRCh37/hg19
##INFO=<ID=SYMBOL,Number=1,Type=String,Description="HGNC gene symbol">
##INFO=<ID=STRAND,Number=1,Type=String,Description="+ or - depending on whether the gene lies in the positive or negative strand">
##INFO=<ID=TYPE,Number=1,Type=String,Description="E or I depending on whether the variant position is exonic or intronic (GENCODE V24lift37 canonical annotation)">
##INFO=<ID=DIST,Number=1,Type=Integer,Description="Distance between the variant position and the closest splice site (GENCODE V24lift37 canonical annotation)">
##INFO=<ID=DS_AG,Number=1,Type=Float,Description="Delta score (acceptor gain)">
##INFO=<ID=DS_AL,Number=1,Type=Float,Description="Delta score (acceptor loss)">
##INFO=<ID=DS_DG,Number=1,Type=Float,Description="Delta score (donor gain)">
##INFO=<ID=DS_DL,Number=1,Type=Float,Description="Delta score (donor loss)">
##INFO=<ID=DP_AG,Number=1,Type=Integer,Description="Delta position (acceptor gain) relative to the variant position">
##INFO=<ID=DP_AL,Number=1,Type=Integer,Description="Delta position (acceptor loss) relative to the variant position">
##INFO=<ID=DP_DG,Number=1,Type=Integer,Description="Delta position (donor gain) relative to the variant position">
##INFO=<ID=DP_DL,Number=1,Type=Integer,Description="Delta position (donor loss) relative to the variant position">
#CHROM	POS	ID	REF	ALT	QUAL	FILTER	INFO
10	92946	.	C	T	.	.	SYMBOL=TUBB8;STRAND=-;TYPE=E;DIST=-53;DS_AG=0.0000;DS_AL=0.0000;DS_DG=0.0000;DS_DL=0.0000;DP_AG=-26;DP_AL=-10;DP_DG=3;DP_DL=35
10	92946	.	C	G	.	.	SYMBOL=TUBB8;STRAND=-;TYPE=E;DIST=-53;DS_AG=0.0008;DS_AL=0.0000;DS_DG=0.0003;DS_DL=0.0000;DP_AG=34;DP_AL=-27;DP_DG=35;DP_DL=1
10	92946	.	C	A	.	.	SYMBOL=TUBB8;STRAND=-;TYPE=E;DIST=-53;DS_AG=0.0004;DS_AL=0.0000;DS_DG=0.0001;DS_DL=0.0000;DP_AG=-10;DP_AL=-48;DP_DG=35;DP_DL=-21
10	92947	.	A	C	.	.	SYMBOL=TUBB8;STRAND=-;TYPE=E;DIST=-54;DS_AG=0.0002;DS_AL=0.0000;DS_DG=0.0000;DS_DL=0.0000;DP_AG=-49;DP_AL=-11;DP_DG=0;DP_DL=34
10	92947	.	A	T	.	.	SYMBOL=TUBB8;STRAND=-;TYPE=E;DIST=-54;DS_AG=0.0002;DS_AL=0.0000;DS_DG=0.0000;DS_DL=0.0000;DP_AG=33;DP_AL=-11;DP_DG=-22;DP_DL=34
10	92947	.	A	G	.	.	SYMBOL=TUBB8;STRAND=-;TYPE=E;DIST=-54;DS_AG=0.0006;DS_AL=0.0000;DS_DG=0.0001;DS_DL=0.0000;DP_AG=33;DP_AL=-11;DP_DG=34;DP_DL=32
```

DRAGEN Annotation extracts these INFO fields:

* `DS_AG` - Δ score (acceptor gain)
* `DS_AL` - Δ score (acceptor loss)
* `DS_DG` - Δ score (donor gain)
* `DS_DL` - Δ score (donor loss)
* `DP_AG` - Δ position (acceptor gain) relative to the variant position
* `DP_AL` - Δ position (acceptor loss) relative to the variant position
* `DP_DG` - Δ position (donor gain) relative to the variant position
* `DP_DL` - Δ position (donor loss) relative to the variant position

These fields report the predicted splice gain or loss and the relative position of the effect.

#### Filtering

SpliceAI provides entries across the genome. Many low-scoring entries have limited value, especially in intergenic regions. These entries increase storage requirements and slow annotation.

DRAGEN Annotation filters out low-confidence entries except within 15 bp of nascent splice sites. In those regions, low-confidence predictions can still help identify potential splice disruption.

### JSON output

```json
"spliceAI":[ 
   {
      "hgnc":"BLCAP",
      "acceptorGainDistance":-3,
      "acceptorGainScore":0.3,
      "donorLossDistance":7,
      "donorLossScore":0.9
   },
   { 
      "hgnc":"NNAT",
      "acceptorGainDistance":-1,
      "acceptorGainScore":0.2,
      "donorGainDistance":-2,
      "donorGainScore":0.3
   }
]
```

<table><thead><tr><th width="199.54296875">Field</th><th width="140.0703125">Type</th><th>Notes</th></tr></thead><tbody><tr><td>hgnc</td><td>string</td><td>HGNC gene symbol</td></tr><tr><td>acceptorGainDistance</td><td>int</td><td>± bp from current position</td></tr><tr><td>acceptorGainScore</td><td>float</td><td>range: 0 - 1.0. 1 decimal place</td></tr><tr><td>acceptorLossDistance</td><td>int</td><td>± bp from current position</td></tr><tr><td>acceptorLossScore</td><td>float</td><td>range: 0 - 1.0. 1 decimal place</td></tr><tr><td>donorGainDistance</td><td>int</td><td>± bp from current position</td></tr><tr><td>donorGainScore</td><td>float</td><td>range: 0 - 1.0. 1 decimal place</td></tr><tr><td>donorLossDistance</td><td>int</td><td>± bp from current position</td></tr><tr><td>donorLossScore</td><td>float</td><td>range: 0 - 1.0. 1 decimal place</td></tr></tbody></table>

## Interpreting scores

SpliceAI delta scores range from `0` to `1`.

The SpliceAI team suggests this interpretation:

| Range         | Confidence | Pathogenicity     |
| ------------- | ---------- | ----------------- |
| 0 ≤ x < 0.1   | low        | likely benign     |
| 0.1 ≤ x ≤ 0.5 | medium     | likely pathogenic |
| x > 0.5       | high       | pathogenic        |

## Resources

* [SpliceAI GitHub](https://github.com/Illumina/spliceAI)
* [Download URL](https://basespace.illumina.com/s/5u6ThOblecrh)
* Related paper:
  * Rowlands, et al. Comparison of in silico strategies to prioritize rare genomic variants impacting RNA splicing for the diagnosis of genomic disorders. *Scientific Reports* (2021). <https://doi.org/10.1038/s41598-021-99747-2>


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://help.connected.illumina.com/annotation/v4.0/data-sources/splice-ai.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
