> For the complete documentation index, see [llms.txt](https://help.connected.illumina.com/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://help.connected.illumina.com/annotation/v4.0/software-functionality/parallel-processing.md).

# Parallel Processing

### Introduction

DRAGEN Annotation can annotate variants using multiple worker processes in parallel to reduce end-to-end runtime on large inputs. Parallel processing is intended for large annotation jobs where the per-variant work dominates the overhead of spawning and coordinating workers.

{% hint style="info" %}
Parallel processing is supported for both `json` and `vcf` output formats — the only two output formats available.
{% endhint %}

### When Parallel Processing Is Used

DRAGEN Annotation automatically decides between single-worker and multi-worker execution based on the number of variant positions in the input file:

* If the number of positions is **less than or equal to** `singleWorkerPositionsCountThreshold` (default: `500000`), the annotator runs in **single-worker mode**, regardless of the value of `workers`.
* If the number of positions is **greater than** the threshold, the annotator runs in **parallel mode** using the configured number of workers.

This behavior avoids the situation where the overhead of coordinating workers (process startup, data distribution, output merging) makes parallel execution slower than a single worker. For small inputs, single-worker mode is typically the fastest option.

The annotator also falls back to **single-worker mode** regardless of `workers` or the position threshold when:

* **Methylation annotation** is enabled (`annotationOptions.enableMethylationAnnotation`).
* The input VCF is **plain gzip** (not block-gzipped or uncompressed). Multi-worker mode requires random access into the file; plain gzip cannot be seeked.

### Configuration

Parallel processing is controlled by two parameters that can be set in the run configuration or on the command line.

#### Run configuration

```json
"parallel": {
  "workers": 2,
  "singleWorkerPositionsCountThreshold": 500000
}
```

#### Command line

```bash
--parallel.workers 2
--parallel.singleWorkerPositionsCountThreshold 500000
```

#### Parameters

| Parameter                                      | Default                                              | Description                                                                                                      |
| ---------------------------------------------- | ---------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------- |
| `parallel.workers`                             | `2` in setup-generated configs; CPU count if omitted | Number of worker processes to use when the input exceeds the single-worker threshold.                            |
| `parallel.singleWorkerPositionsCountThreshold` | `500000`                                             | If the input has fewer positions than this value, the annotator runs with a single worker and ignores `workers`. |

{% hint style="info" %}
`download --parallel.workers` controls the number of **concurrent downloads**, not annotation worker processes.
{% endhint %}

### Choosing the Number of Workers

Selecting a value for `workers` is a trade-off between runtime, CPU utilization, and memory usage. A few practical guidelines:

* **More workers reduce runtime up to a point.** Runtime improves as workers are added, but the gains taper off well before you reach the number of physical CPU cores. Beyond that point, additional workers mostly add coordination overhead without reducing runtime, and can even make the job slower.
* **Memory usage grows with the number of workers.** Each worker loads its own annotation state, so peak RAM increases as `workers` increases. Make sure your machine has enough memory to comfortably fit the expected peak (see the example below) with headroom for the operating system and other processes.
* **Do not oversubscribe cores.** Setting `workers` higher than the number of physical cores rarely helps and typically increases both runtime and memory pressure.
* **Small inputs stay single-worker.** For inputs below `singleWorkerPositionsCountThreshold` (default `500000` positions), increasing `workers` has no effect — the annotator will still run with a single worker.

If you are unsure where to start, the `2` workers written by `setup` is a safe choice for most environments. For large VCFs on a machine with ample RAM and many cores, `8` workers typically offers the best balance between speed and memory usage — you capture most of the achievable speedup (roughly 3x–3.5x) at a small fraction of the memory cost of higher worker counts. Values up to the physical core count (e.g., `16`–`24` on the machine used for the example below) can shave off a bit more runtime, but at significantly higher peak memory.

### Example: Scaling Behavior on Real-World VCFs

The tables below show how runtime, CPU utilization, and peak memory scale with `parallel.workers` across three whole-genome VCFs of varying size and characteristics, all annotated with the full set of supplementary annotation (SA) data. Each configuration was run twice (`n_reps = 2`) and the median runtime is reported.

{% hint style="info" %}
These numbers are illustrative and hardware-specific. Absolute runtimes, CPU percentages, and memory footprints will vary between machines, but the overall shape of the trends — diminishing runtime improvements past a certain worker count and steadily growing memory usage — is representative of what to expect.
{% endhint %}

**Test machine:** Linux `x86_64`, 24 physical cores / 48 logical cores, 376 GB RAM.

**Columns:**

* **Configuration** — Annotator version and worker count.
* **Time (s)** — Median wall-clock runtime.
* **Speedup** — Runtime relative to the DRAGEN Annotation 4.0 single-worker baseline for the same VCF. Values above `1.0x` are faster than the baseline.
* **p90 CPU (%)** — 90th-percentile CPU utilization across the run, expressed as a percentage of a single core (so `100%` ≈ 1 core fully utilized, `4800%` ≈ 48 cores).
* **Peak Memory (GB)** — Maximum resident set size observed during the run.

Nirvana 3.27 rows are included for comparison against the previous major version.

#### VCF 1: `dragen_pedigree_wg`

Whole-genome pedigree VCF produced by DRAGEN with approximately 7 million variants

| Configuration                      | Time (s) | Speedup | p90 CPU (%) | Peak Memory (GB) |
| ---------------------------------- | -------: | ------: | ----------: | ---------------: |
| Nirvana 3.27                       |    799.2 |   0.99x |         408 |             13.3 |
| DRAGEN Annotation 4.0 — 1 worker   |    789.4 |   1.00x |         482 |             16.1 |
| DRAGEN Annotation 4.0 — 2 workers  |    527.2 |   1.50x |         904 |             25.3 |
| DRAGEN Annotation 4.0 — 4 workers  |    316.8 |   2.49x |        1736 |             42.1 |
| DRAGEN Annotation 4.0 — 8 workers  |    227.9 |   3.46x |        3133 |             72.6 |
| DRAGEN Annotation 4.0 — 16 workers |    212.6 |   3.71x |        4168 |            122.0 |
| DRAGEN Annotation 4.0 — 20 workers |    202.7 |   3.89x |        4314 |            138.1 |
| DRAGEN Annotation 4.0 — 24 workers |    208.5 |   3.79x |        4429 |            164.8 |
| DRAGEN Annotation 4.0 — 32 workers |    213.5 |   3.70x |        4479 |            165.7 |
| DRAGEN Annotation 4.0 — 48 workers |    228.7 |   3.45x |        4568 |            202.2 |
| DRAGEN Annotation 4.0 — 64 workers |    242.8 |   3.25x |        4667 |            259.0 |

#### VCF 2: `dragen_HG004_NSX_10B_Wave3_NIST_35x`

Whole-genome VCF for HG004 (NIST Ashkenazi son) sequenced at \~35x on NovaSeq X with approxmiately 5.1 million variants

| Configuration                      | Time (s) | Speedup | p90 CPU (%) | Peak Memory (GB) |
| ---------------------------------- | -------: | ------: | ----------: | ---------------: |
| Nirvana 3.27                       |    564.9 |   0.94x |         432 |             13.4 |
| DRAGEN Annotation 4.0 — 1 worker   |    532.8 |   1.00x |         507 |             15.6 |
| DRAGEN Annotation 4.0 — 2 workers  |    347.3 |   1.53x |         975 |             25.4 |
| DRAGEN Annotation 4.0 — 4 workers  |    219.4 |   2.43x |        1841 |             42.1 |
| DRAGEN Annotation 4.0 — 8 workers  |    159.5 |   3.34x |        3294 |             71.8 |
| DRAGEN Annotation 4.0 — 16 workers |    155.4 |   3.43x |        4283 |            119.6 |
| DRAGEN Annotation 4.0 — 20 workers |    150.6 |   3.54x |        4406 |            147.1 |
| DRAGEN Annotation 4.0 — 24 workers |    155.6 |   3.42x |        4478 |            165.5 |
| DRAGEN Annotation 4.0 — 32 workers |    160.6 |   3.32x |        4546 |            166.7 |
| DRAGEN Annotation 4.0 — 48 workers |    172.6 |   3.09x |        4634 |            215.3 |
| DRAGEN Annotation 4.0 — 64 workers |    187.9 |   2.84x |        4703 |            266.1 |

#### VCF 3: `veritas_NS26001675`

Whole-genome VCF from a Veritas Genetics NovaSeq run with approximately 5.6 million variants.

| Configuration                      | Time (s) | Speedup | p90 CPU (%) | Peak Memory (GB) |
| ---------------------------------- | -------: | ------: | ----------: | ---------------: |
| Nirvana 3.27                       |    464.0 |   1.08x |         467 |              7.2 |
| DRAGEN Annotation 4.0 — 1 worker   |    500.9 |   1.00x |         532 |             10.1 |
| DRAGEN Annotation 4.0 — 2 workers  |    359.0 |   1.40x |         999 |             17.8 |
| DRAGEN Annotation 4.0 — 4 workers  |    198.6 |   2.52x |        1930 |             30.5 |
| DRAGEN Annotation 4.0 — 8 workers  |    150.7 |   3.32x |        3457 |             51.5 |
| DRAGEN Annotation 4.0 — 16 workers |    159.6 |   3.14x |        4257 |             78.2 |
| DRAGEN Annotation 4.0 — 20 workers |    151.7 |   3.30x |        4317 |             89.5 |
| DRAGEN Annotation 4.0 — 24 workers |    147.0 |   3.41x |        4449 |            103.7 |
| DRAGEN Annotation 4.0 — 32 workers |    156.8 |   3.19x |        4527 |            129.3 |
| DRAGEN Annotation 4.0 — 48 workers |    171.8 |   2.92x |        4715 |            174.4 |
| DRAGEN Annotation 4.0 — 64 workers |    186.1 |   2.69x |        4832 |            244.3 |

#### Key observations

* **Diminishing returns past \~8 workers.** Across all three VCFs, going from 1 to 8 workers produces roughly a 3.3x–3.5x speedup. Adding more workers beyond that yields only marginal improvements, and after \~24 workers runtime typically gets *worse* due to coordination overhead.
* **Optimal worker count is close to the physical core count.** The fastest runtime for each VCF is achieved between 20 and 24 workers, which is at or just under the 24 physical cores of the test machine.
* **Peak memory grows steeply with worker count.** From 1 → 64 workers, peak RSS grows by roughly 15x–25x. At 64 workers, peak memory can exceed 250 GB — provision memory based on this peak, not the average.
* **Nirvana 3.27 vs. DRAGEN Annotation 4.0 (1 worker) is roughly comparable.** With a single worker, DRAGEN Annotation 4.0 runs within about ±10% of Nirvana 3.27. The runtime improvements in 4.0 come primarily from enabling parallel processing.

### Recommendations

* Start with the setup default (`workers: 2`) and increase only if runtime is a bottleneck.
* Do not set `workers` higher than the number of physical CPU cores on the machine.
* Ensure available RAM comfortably exceeds the expected peak memory for the chosen `workers` value.
* Leave `singleWorkerPositionsCountThreshold` at its default unless you have a specific reason to change it; the default avoids the overhead of parallel execution on small inputs.
* Use parallel processing only for `json` or `vcf` outputs.


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://help.connected.illumina.com/annotation/v4.0/software-functionality/parallel-processing.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
