Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
The Emedgene platform utilizes the Okta Identity Management solution to control user access. This improves user management, enhances access and authentication security, and allows organizations to implement single sign-on for their users.
The Cases tab provides an overview of genomic sequencing cases submitted by the organization, as well as individual case details.
—displays a list of cases along with key details
—enables customization of the table view, including grouping and filtering of cases
—opens when a case is selected, providing additional information
The Family tree tab includes the following information:
Pedigree diagram. Pedigree legend can be found .
Sample details for each family member:
Phenotypes. For family members other than the test subject, phenotypes are categorized as:
You can sort cases by Creation date, Due date, Quality, or Resolution.
Hover over the column header and click the up or down arrow to sort in ascending or descending order.
Alternatively, click the column name and select Sort ascending or Sort descending from the dropdown menu
The current sort direction is indicated by a single arrow icon next to the column name.
Build a pedigree via the visual tool.
It is ideal that a proband selected for case analysis is affected and has disease phenotype(s).
You can add a Father, a Mother, a Sibling, or a Child to any family member, starting with the Proband. To do this, choose their icon, then click on the Add family member button in the bottom right corner of the pedigree builder to select a family member.
More information about the pedigree symbols can be found .
To delete a family member, choose their icon, then click on the Delete Subject button in the top right corner of the Add patient information panel.
The Emedgene pipeline prioritizes variant annotations based on the calling methodology rank order. The first appearance of a variant is annotated according to the following hierarchy:
TARGETED
STAR_ALLELE
STR_REPEAT_EXPANSION
The header on the Individual Case page shows the Case ID and current .
Change the
Reanalyze the case
and write interpretation notes
Related—directly match one of the proband’s phenotypes
Unrelated—do not match any of the proband’s phenotypes
Medical Condition – Indicates whether the individual is considered Healthy or Affected in the case
Sex. Specified by the user
Age. Automatically calculated in years based on the provided date of birth
Maternal and Paternal ethnicity—ethnic background of the proband’s parents
BAM file location. Shown where relevant


Every case is annotated with the attached table of resources, including proprietary Illumina prediction scores PrimateAI-3D and SpliceAI. All annotations are versioned, and versions recorded in a Versions tab, and saved per case. Key variant significance and knowledge graph databases are updated monthly, so that the most up-to-date information is available during analysis.
Sequencing lab information section reports sequencing run technicalities as indicated during case creation:
Lab
Instrument
Reagents
Kit type
Expected coverage
Protocol
The Emedgene platform is divided into two applications:
Analyze—genomic analysis workbench
Curate—the knowledge management system
Go to the nine-dot app launcher icon located on the top navigation panel and select Curate from the dropdown menu.
Go to the nine-dot app launcher icon located on the Curate navigation panel and select Analyze from the dropdown menu.
MRJD
FORCED_GENOTYPING
SMALL_VARIANT
CNV_READ_DEPTH
SV_SPLIT_END
UNKNOWN
Variants are considered identical if they share the same:
Chromosome
Position
Reference allele (REF)
Alternate allele (ALT)
When applied to Copy Number Variants (CNVs), this approach may merge variants even if they have different lengths.
For DRAGEN versions earlier than 4.2, when ingesting a DRAGEN Manta VCF containing SVs of type INS, replace the following line in the VCF header:
##source=DRAGEN <version>with
##source=MANTA-DRAGEN <version>Example:
Replace
##source=DRAGEN 05.121.645.4.0.3with
Note: Variant types currently annotated and displayed in Emedgene are DEL, DUP and INS.
##source=MANTA-DRAGEN 05.121.645.4.0.3Preview the case report

Welcome to Emedgene, where we unlock genomic insights for hereditary disease and streamline your tertiary analysis workflows.
So you've signed in and can't wait to get started? Here we will guide you through the platform architecture, case creation, and results review. You can dive a bit deeper by following the links and exploring manuals for the platform's applications:
Analyze: Genomic analysis workbench, where you can accession, interpret, curate and report on your cases, while also efficiently managing the lab workflow
Curate: A repository for all of your organizational curated knowledge
Contact Illumina technical support at techsupport@illumina.com.
The platform is operated from the .
From here you can enter:
To enter the flow, click on the namesake button on the . Here:
Select file type.
Upload files.
Create a family tree.
Select a case to review on the . You'll be directed to the that:
Showcases an AI-curated shortlist of variants suggested to be checked first, namely and
Provides numerous customizable to help you explore the total list of genetic variants by yourself
The Cases table navigation panel provides several tools to help you customize your table view and manage cases. It includes the following components:
You can use the Case search tab in the top bar to search for cases by the Case ID or Proband ID.
In order to prevent accidental data loss, deleting cases in Emedgene includes a staging step before permanent case deletion.
Result:
After deletion is confirmed:
All cases marked Trash bin are permanently removed
An activity entry is recorded
Email notifications are sent to users who have opted in
Click on the icon in the top navigation panel to open the Help dropdown menu.
From there, you can access:
Help Center: Find feature guides, step-by-step instructions, and tips to help you get the most out of the platform.
Walkthroughs: View short interactive demos of workflows (in development).
Feature requests: Share your ideas and feedback.
What's new: Stay updated with the latest release notes.
About: View general information such as your organization name and platform version.
To directly import files from your own storage, link it to an organization's storage in Emedgene.
Click on the user initials or profile picture at the rightmost corner of the top navigation panel and select Settings
Select the Management tab and proceed to Storage card that lists currently linked storages.
To add a new storage:
Click Add Storage
Choose a storage type from:
Illumina BioInsight Platform Core
Illumina Basespace (BSSH)
Check the connection to confirm that the storage is successfully linked.
To do this, find the storage in the list and check the cloud icon status:
If it's green, the connection is set correctly
If it's red and strikethrough, something went wrong. Hover over the icon to see details
Click Manage on the right to the storage details.
Click Delete on the right to the storage details.
You can choose one of the following options:
Existing sample: Pick one of the samples already loaded on the platform
Upload new sample: Upload files from your PC and enter sample name
Choose from storage: Choose files from your cloud storage and enter sample name
No sample: Postpone uploading files but proceed with case creation or skip uploading files for family members other than Proband
Note: The fields marked with (*) are mandatory.
Options: Male, Female, Unknown.
Indicates the family relationship of a subject to the Proband automatically inferred from the pedigree. Options: Father, Mother, Sibling, Child, Other.
Expected format: mm/dd/yyyy.
Mark the checkbox if you want to exclude the sample from the AI Shortlist analysis and Inheritance filters while preserving genotype data.
If a sample shares some phenotypes with the Proband, you can copy them by checking this box. Proband's phenotypes will appear in a newly created Related Phenotypes section. To remove any of the proband's phenotypes not observed in a current individual, click the ☒ button next to the HPO term in the Related Phenotypes section.
Phenotypes not shared with a Proband. They can be added one by one (Selection mode) or in batch (Batch mode).
Please follow the steps described below for each phenotype:
Enter an HPO term (e.g., Hypoplasia of the ulna), an HPO ID (e.g., HP:0003022), or a descriptive phenotype name (e.g., Underdeveloped ulna) in the search box;
Select a matching term from a dropdown menu and press Complete after you've added all the terms.
Paste a list of comma-separated HPO terms or HPO IDs in the search box and press Complete.
Select the case type in order to define the proper analysis of your case.
Users can utilize a custom region of interest (ROI) BED file to limit analysis results to variants within the designated regions. A ROI BED determines which genomic regions will be included in the variant analysis.
If no custom ROI BED is selected, the system uses the default ROI BED file based on the case type.
You can select any region of interest, regardless of the case type.
When selecting a Custom BED as you region of interest, you must select a specific BED file that is already configured in your organization.
A coverage BED file is used to calculate and determine quality control (QC) metrics for your case. This file defines the genomic regions that should meet coverage requirements during sequencing.
After selecting a coverage BED file, the available reference sequences for this kit will be displayed.
Specify details such as laboratory name, sequencing machine used, sequencing reagent kit, and expected coverage.
You can limit analysis to a gene list in the platform while creating a case. Choose between:
No limitation of the analysis.
Select one of the previously added gene lists from a dropdown list.
Generate a new virtual panel: add a List title and then add all the gene symbols one by one (Selection mode) or in a batch (Batch mode).
A new gene list can be comprised from a combination of configured gene lists and/or individual genes.
A gene list can by configured to hold up to 10,000 genes.
A new gene list can be created by combining configured gene lists and/or individual genes. Each gene list can be configured to contain up to 10,000 genes.
Note: Please use the up-to-date gene symbols approved by the Hugo Gene Nomenclature Committee. When adding gene symbols in a Batch mode, those genes that do not comply with HGNC standards will be automatically excluded from the gene list. These genes will appear for 3 seconds in a black error box at the bottom of the screen.
For each gene please follow the steps described below: Enter a gene symbol in the search box in the right panel (Candidate Genes) and select a matching symbol from a dropdown menu.
After selecting batch mode, paste a list of comma-separated gene symbols in the search box in the right panel (Candidate Genes).
You can choose between two different modes of a gene list feature:
Selected by default.
AI Shortlist is limited to the selected gene panel, no variants in other genes are considered in the results. If this in silico panel is used for analysis of exome or genome data, the gene restriction may be lifted during manual analysis to "open-up" the entire exome or genome for analysis.
Analysis is performed for variants in all the genes. Variants in the targeted genes get upgraded scores during prioritization by the AI Shortlist algorithm.
Classic joint calling consists of calling variants "simultaneously across all sample BAMs, generating a single call set for the entire cohort." (GATK.broadInstitute.org)
When running from BAM or FastQ samples on Emedgene, we do not apply a classic joint calling but a BAM look-up methodology.
This methodology consists of retrieving coverage information from BAM during the VCF merging process. Thus, if a variant does not exist in a parental sample, the algorithm will check the coverage in that position using data from the BAM file. The position will be considered as "REF" allele if it is covered (depth > 3), and "No coverage" or "N/A" (./. in the VCF FORMAT/GT field), if it is below that threshold or has no coverage.
This process involves the creation of a “genome coverage” file as a separate preliminary step. The coverage file could also be provided via a BED or a gVCF file.
BAM look-up approach is slightly different from classic joint calling used by the joint calling option in DRAGEN and other variant callers, and therefore will not produce identical results.
However, it is important to mention that Emedgene platform supports joint called VCF files, as well.
Remark: If a coverage file (ie. BED, BAM, gVCF) is not provided, then it is not possible to estimate the presence of REF allele in empty positions. As a consequence, "No_coverage" value will be assigned to those variants, which can affect the .
Limitation: It should be noted that the current data pipeline has a limitation stemming from the way it merges variants from different samples into the same case (e.g., in a trio). Since it is based on bcftools, variants are identified by the chromosome number, start position, reference allele, and alternate allele. However, it does not take into account the size of the variant itself. As a result, this may sometimes lead to inaccurate merging of CNV-type variants that differ in size. That limitation is not present when joint calling is used.
To select variants with a particular tag, use the Filter candidates dropdown menu in the top right corner. You can select from Most Likely, Candidate, Incidental, Carrier, Not Reviewed, or any custom tags used in your organization.
For each variant on the Candidates tab, you can explore the gene-related disease, gene symbol, main variant details, and variant tag.
When a variant is found in a gene with no known disease association, the gene-related disease cannot be displayed. Such variants appear under the Gene of Unknown Significance heading.
All the relevant Most Likelies and Candidates fitting a сompound heterozygous mode of inheritance are presented together. This refers to both confirmed and assumed compound heterozygosity (cases with at least one parent and singleton cases, respectively).
If you want to inspect the complete variant information, click on the variant bar to continue to the Evidence page. You can visualize evidence in text or graphical format (Click on the interactive text in the top left corner: Show evidence as text or Show evidence graph to toggle between the two).
The Lab tab shows sample and case-level quality metrics so you can check data reliability before starting interpretation.
Summary dashboard: highlights the key quality indicators, with more details provided in the subsequent sections
Sequencing lab information section: reports sequencing run technicalities
Case quality section: summarizes the data quality of the case
: highlights quality metrics for each sample
: displays the results of the relationship validation for each pair of samples in a family tree.
: highlights regions that may not have been adequately sequenced
The Summary dashboard provides a quick overview of key quality indicators at the case and sample levels.
Case quality — displays the overall case quality status.
Sample quality — reflects the sample quality status.
Evaluation kit — specifies the QC BED kit used to evaluate coverage depth and breadth. If no kit is specified when analysis launches, NCBI RefSeqGene is used as the default reference.
Custom gene coverage — indicates whether coverage of genes in the selected panel meets the expected threshold defined by the QC BED.
— displays relationship-validation results and confirms whether the submitted pedigree matches the genetic data.
The Dashboard tab depicts an overview of the user activity on the Emedgene platform and provides a glance at key performance indicators for an organization.
The Diagnostic Yield card shows the percentage of cases classified as Resolved out of the total number of cases of the same type.
The Status Diagram card shows the total number of cases submitted by the organization and the count of cases for each status.
The Stale Cases card highlights cases stalled at intermediate stages of analysis that haven't been finalized.
The Network Activities panel displays a timeline of user activities within the organization. This log includes activity like creating a case, verifying a filter preset, changing a , generating a report, and more.
The Case details panel provides comprehensive information about a particular case.
The Case details panel is organized into three tabs:
Case info: Displays technical, operational, and clinical information about the case.
Family tree: Shows a graphical pedigree and sample details for each family member.
Labels help you mark cases for specific uses and filter case subsets in the Cases tab. You can , , or case labels in the .
In the case's row of the Cases table, select Edit ( button) in the Label column.
In the Edit label
You can filter cases using most of the fields, as well as by the case outcome category (Resolved / Not resolved), which is the only filter not displayed as a column.
Whenever an organization is created, we automatically allocate bucket folders in AWS S3 cloud storage to it:
Path for upload
Folder intended to store input case files.
Authorized user has view and upload privileges.
Path for download
This folder contains a partially annotated (excluding results of proprietary algorithms) VCF file per case.
Log in to your Illumina private domain via URL in the following format: . This opens the Connected Platform Home
In the left navigation panel: User > API keys
Name the key
While adding a new case, you will build a pedigree and annotate each of the samples with data required for analysis.
After the case has been created, the family tree is available in the panel (righthand panel of the Cases page).
Icon fill color in other pedigree members indicates the presence or absence of the proband's phenotypes in a present sample (regardless of the potential presence of additional unrelated phenotypes).
A preset group is a reusable set of filter presets applied for specific case types, as defined by your laboratory SOPs.
Select a preset group in the Case info screen during case creation or .
The group selected for the case determines which presets appear in the Presets tab of the Filtering panel.
If no group is selected, the system automatically applies the default preset group defined in Lab workflow settings.
Download and install node js platform via
Minimum version required: 16
Upgrade existing installation: nvm install --lts
Download the batch case create script.
Replace my-domain with your Emedgene domain.
Illumina cloud: my-domain.emg.illumina.com
Legacy Emedgene cloud: my-domain.emedgene.com
For DRAGEN versions earlier than 4.2, when ingesting a VCF containing STRs (either a DRAGEN STR VCF or a DRAGEN ExpansionHunter VCF), add the following line to the VCF header:
Additionally, ensure that the contig lines are present in the VCF headers. If they are missing, please use the following ones:
GRCh37
GRCh38
Both GRCh37/hg19 and GRCh38/hg38 are supported. You can run cases with both reference genomes in the same organization.
GRCh38/hg38: Multigenome Graph hg38-alt_masked.cnv.graph.hla.rna-10-r4.0-1.tar.gz.
Pre-built multigenome hash tables for hg38. The hash table builds include DNA, RNA, CNV, and HLA tables. Download .
GRCh37d5: Multigenome Graph hs37d5-cnv.graph.hla.rna-10-r4.0-1.tar.gz
Unlike single-nucleotide variants (SNVs), a multi-nucleotide variant (MNV) represents a single event involving multiple consecutive bases. In Emedgene, small variants are recognized as those comprising an MNV if they are located within a 2-nucleotide distance.
Emedgene recognizes MNV as a distinct variant type and supports ingestion from VCF, annotation, and filtering.
Each MNV is represented and annotated as:
An MNV itself (eg, AG>TC)
Individual SNVs derived from the MNV (eg, A>T and G>C), for compatibility with existing tools and workflows
Both the MNV and its underlying SNVs display the "Suspected MNP" badge in the
Annotations from organization databases appear in various parts of the platform, each showing certain details.
Allele frequency—in "[Organization DB] AF (%)" column
Allele count—in "[Organization DB] AF (#)" column
Summary tab Population summary card
The user can enter a specific case from the by clicking Full details in the corresponding row of the case table.
—displays a Case ID and and includes Case interpretation, Edit case info, and Report preview buttons
—highlights a shortlist of variants, suggested to be reviewed first - Most Likely Candidates and Candidates
You can update the case status either from the individual case page or from the Cases table.
Finalized case status can be applied only from the individual case page to prevent unintended .
Open the case page.
In the page header, select the dropdown () icon next to the current case status.
The Sample quality section in the Lab tab gives you a quick view of the reliability of sequencing or array data used in your case.
The metrics displayed in the Sample quality section and their underlying calculation vary depending on the case type (see below).
NGS sample quality metrics provide an overview of QC validation and coverage results for each sample.
Add phenotypes.
Specify analysis details.
Launch the analysis!
Investigate the evidence on the Variant page and assign appropriate tags to the variants of interest.
When you're ready to finalize the case, indicate the end result of the analysis and variants to be reported in the Case interpretation widget.

Coverage metrics for a target region defined by a QC BED file (or RefSeq coding regions if no kit is provided) included in the Sample quality section:
Average coverage Average depth of coverage for a target region
% Bases with coverage >10x percentage of a target region that is covered at a minimum depth of 10x
% Bases with coverage >20x percentage of a target region that is covered at a minimum depth of 20x
Blue bars represent each of these parameters per sample, while a vertical line represents a general metric across all the samples of the same case type in the account.
The Quality status provides a quick assessment of array data reliability for each sample:
High
Call rate ≥ 0.99 and Log R dev ≤ 0.2
Low If either condition is not met
N/A
If the QC file not available
Use the Quality status to quickly screen whether a sample meets minimal QC thresholds before starting detailed interpretation.
The Autosomal call rate field displays percentage of loci on the array for which a genotype call was successfully made, that only includes autosomes.
A high call rate indicates a high-quality sample and successful genotyping. Low call rates can signify problems with the DNA sample (poor quality or quantity) or issues during the array processing.
Displayed to three decimal places.
Percentage of reads mapped to the reference sequence.
Blue bars represent each of these parameters per sample, while a vertical line represents a general metric across all the samples of the same case type in the account.
To streamline case review, the AI Shortlist pre-selects the list of variants likely to be causative for each case: Most Likely Candidates and Candidates.
Variants that are most promising for solving the case. This list is limited to 10 top-scored variants but may include more if more than one variant is tagged per gene (suggesting compound heterozygosity). We can change the Most Likely Candidates number limit upon request.
Several dozen highly scored variants worth considering.
The ranking of variants by AI Shortlist considers:
SNVs
CNVs
SNV + CNV compound heterozygotes
SVs
mtDNA variants
STRs
The AI Shortlist rates variants based on predicted variant effects, alternative allele frequency, familial segregation pattern, phenotypic match, in silico predictions, and other relevant information from scientific papers and databases.
During the case review, you can untag variants selected by the AI Shortlist or manually tag ones not selected by the AI Shortlist.
The Call rate field displays the percentage of loci on the array for which a genotype call was successfully made.
Call rate is one of the key metrics used to determine array sample quality, alongside log R deviation.
A high call rate indicates a high-quality sample and successful genotyping. Low call rates can signify problems with the DNA sample (poor quality or quantity) or issues during the array processing.
Displayed to three decimal places.
The Log R Deviation (or Log R Ratio standard deviation) quantifies the variability of the the signal intensity for each SNP marker on an array, ie, noise level.
Log R deviation is one of the key metrics used to determine array sample quality, alongside call rate.
Lower values indicate more consistent signal intensities. A high Log R Deviation can indicate a poor-quality sample or potential issues with CNV calling.
Displayed to three decimal places.
Sequencing error rate refers to the frequency at which incorrect base calls are made during sequencing process.
Error rate is calculated as number of low quality variants / total variants.
Blue bars represent each of these parameters per sample, while a vertical line represents a general metric across all the samples of the same case type in the account.




The Add New Case flow does not validate that sample IDs are unique or that input files are uncorrupted. Please ensure sample IDs are unique and that input files are valid before creating the case.
A case won't run if Proband sample files are missing. However, sample files are not mandatory for the rest of the family members (although highly recommended).
When choosing an existing file path, the samples used may be cached from the original run. For a top-up flow please use a new file path.
When you are loading sample files from your PC or choosing them from the storage, and there is more than one file per sample, please ensure that all the necessary files are simultaneously selected in the upload pop-up. You may only select one file type per case (i.e. you may not select both a .vcf and a .bam at the same time).














Select Create new label.
Select Save.
In the case's row of the Cases table, select Edit ( button) in the Label column.
In the Edit label menu, search for the label by name.
Select Save.
In the case's row of the Cases table, select Edit ( button) in the Label column.
In the Edit label menu, select Remove ( button) next to the label name.
Select Save.

During data processing, MNVs are split into consecutive SNVs. The resulting SNVs are annotated with INFO and FORMAT fields that mirror the original record.
SNVs that comprise an MNV display the "Suspected MNP" badge in the Clinical significance tab.
Allele count—in "Allele count" field
Hom/Hemi count—in "Hom/Hemi count" field
The last 10 samples—in "Last 10 samples" field
Population statistics tab Organization DBs
Allele frequency—in "Allele frequency" column
Allele count—in "Allele count" column
Hom/Hemi count—in "# of Homozygotes" column
Allele number—in "Total" column
Visualization tab Population data "[Organization DB]" tracks display variants from organization databases. Left-click a variant in a track to review variant details:
Allele frequency
Allele count
Allele number
Het count
Hom/Hemi count
The last 10 samples
Color-coded "[Organization DB]" badge based on pathogenicity in the Curated DB—"Known variants" column
Summary tab Clinical significance card Color-coded "[Organization DB]" badge based on pathogenicity in the Curated DB
Clinical significance tab Clinical significance card Color-coded "[Organization DB]" badge based on pathogenicity in the Curated DB
##source=ExpansionHunterV4.2##contig=<ID=chr1,length=249250621>
##contig=<ID=chr2,length=243199373>
##contig=<ID=chr3,length=198022430>
##contig=<ID=chr4,length=191154276>
##contig=<ID=chr5,length=180915260>
##contig=<ID=chr6,length=171115067>
##contig=<ID=chr7,length=159138663>
##contig=<ID=chr8,length=146364022>
##contig=<ID=chr9,length=141213431>
##contig=<ID=chr10,length=135534747>
##contig=<ID=chr11,length=135006516>
##contig=<ID=chr12,length=133851895>
##contig=<ID=chr13,length=115169878>
##contig=<ID=chr14,length=107349540>
##contig=<ID=chr15,length=102531392>
##contig=<ID=chr16,length=90354753>
##contig=<ID=chr17,length=81195210>
##contig=<ID=chr18,length=78077248>
##contig=<ID=chr19,length=59128983>
##contig=<ID=chr20,length=63025520>
##contig=<ID=chr21,length=48129895>
##contig=<ID=chr22,length=51304566>
##contig=<ID=chrX,length=155270560>
##contig=<ID=chrY,length=59373566>
##contig=<ID=chrM,length=16571>Edit the downloaded batchCases.csv file. See CSV format requirements for more details.
Execute the batch cases creator as java script using the command below.
Replace my-domain with your Emedgene domain and my-email with your user email.
A prompt for your Emedgene password will appear, enter the password and press Enter.
In case of validation errors in the input CSV, an output CSV called batchCases_results.csv will be created in the same location with detailed error results.
-l will create a log file in the same location.
More information can be found by running
##contig=<ID=chr1,length=248956422>
##contig=<ID=chr2,length=242193529>
##contig=<ID=chr3,length=198295559>
##contig=<ID=chr4,length=190214555>
##contig=<ID=chr5,length=181538259>
##contig=<ID=chr6,length=170805979>
##contig=<ID=chr7,length=159345973>
##contig=<ID=chr8,length=145138636>
##contig=<ID=chr9,length=138394717>
##contig=<ID=chr10,length=133797422>
##contig=<ID=chr11,length=135086622>
##contig=<ID=chr12,length=133275309>
##contig=<ID=chr13,length=114364328>
##contig=<ID=chr14,length=107043718>
##contig=<ID=chr15,length=101991189>
##contig=<ID=chr16,length=90338345>
##contig=<ID=chr17,length=83257441>
##contig=<ID=chr18,length=80373285>
##contig=<ID=chr19,length=58617616>
##contig=<ID=chr20,length=64444167>
##contig=<ID=chr21,length=46709983>
##contig=<ID=chr22,length=50818468>
##contig=<ID=chrX,length=156040895>
##contig=<ID=chrY,length=57227415>
##contig=<ID=chrM,length=16569>curl https://my-domain.emg.illumina.com/v2/js/batchCasesCreator.js --output batchCasesCreator.jsnode batchCasesCreator.js saveTemplateFilenode batchCasesCreator.js create -h https://my-domain.emg.illumina.com -c batchCases.csv -u my-email -lnode batchCasesCreator.js --helpnode batchCasesCreator.js create --helpAzure Blob
Azure Data Lake
AWS S3
Google Cloud
File Transport Protocol (FTP)
Secure File Transport Protocol (SFTP)
Fill in the required credentials
Click Add storage
If data is deleted or moved from the customer's storage, it might adversely affect the case. To learn more about possible consequences, check out this table:
Click on the row of the case you want to view. A pop-up side Case details panel will appear on the right.
To close the panel, click the icon in the top right corner.
To expand the Case details panel, click the (left-pointing arrow) icon on the right edge of the screen.
To collapse it, click the (right-pointing arrow) icon at the top left of the panel.
Status
Resolved or Not resolved
Case details
Type
Label
Quality
Participants
Participants
Go to the Filters menu in the Cases table navigation panel.
Under Field, select the field you want to filter by.
Under List, select a value from the dropdown or manually enter one.
Select Apply to activate the filter.
To add another filter, select Add new under the active filter and repeat steps 1-4.
Go to the Filters menu in the Cases table navigation panel.
To remove a specific filter, click the icon next to it.
In the Cases table navigation panel, click the icon next to the Filters menu.
Identification
Case ID
Sample ID (Proband ID)
Case processing stage
Path for DRAGEN output
This folder contains DRAGEN output files.
Authorized user has view and download privileges.
To get access to your upload, download and DRAGEN output folders, you need to get a key pair consisting of an access key ID and a secret access key. Creating, deactivating, activating and deleting credentials is available for users with Manager and Manage S3 Credentials roles.
You can create and use up to two dynamic access keys at the same time.
When you require technical support, you have the option to generate a new key pair specifically for the troubleshooting process. After the issue has been resolved, you can delete the credentials to ensure security of your system.
The newly generated credentials will only be saved in AWS Identity and Access Management (IAM) and not in our database.
In Settings > Management > S3 Credentials, click on Create Access Key.
You can retrieve the secret access key only when you initially create the key pair. If you lose it, you have to create a new key pair. To immediately copy the secret access key to a secure location, use the Copy to clipboard button.
In Settings > Management > S3 Credentials, click on Deactivate in the corresponding key pair card.
In Settings > Management > S3 Credentials, click on Activate in the corresponding key pair card.
In Settings > Management > S3 Credentials, click on Delete in the corresponding key pair card. Only inactive key pairs can be deleted.
<Region_Cloud>-emg-auto-samples/<org_name>/upload/<Region_Cloud>-emg-downloads/<org_name>/ <Region_Cloud>-emg-auto-results/<org_name>/ Choose one of the following options:
A. Grant access to all workgroups across the domain If your domain includes multiple workgroups and you want the API key to apply universally, select "All current and future Workgroups and roles (Global API Key)"
B. Grant access to specific workgroups Select one or more workgroups from the list. For each selected workgroup, assign the following application roles:
Emedgene Has Access
Illumina Connected Analytics - Has Access
Platform-home Workgroup Admin
Click Generate. Once the API key is generated, copy it to your clipboard or download it as a file.
⚠️ Important: The API key is only accessible while the API Key Generated popup window is open. After closing the window, the key cannot be retrieved. If you didn’t copy or download it, you’ll need to generate a new key.
Log into your Emedgene domain and go to the workgroup where you want to link Platform Core storage
Click on the user avatar and select Settings from the dropdown
Select the Management tab
In the Storage card, click Add Storage
Select Illumina Connected Analytics (not Illumina Connected Analytics V1!) from the Storage type dropdown
Fill the storage credentials:
"Api_key"—the API key before
"Project"—the name of the Project in Platform Core that contains and will contain the data you want to connect
"Path"—the folder within the project where the data is located. This can be used to restrict the user to only be able to access data within the specified folder. Using only “ / “ will allow all folders within your Platform Core project
Click Add Storage
Filled: The individual is affected by all of the proband's phenotypes.
Half-filled: The individual is affected by some of the proband's phenotypes.
Empty: The individual is not affected by any of the proband's phenotypes.
2. Icon color intensity denotes whether sample files have been uploaded for the particular individual.
Full color: The sample has files loaded in the case.
Faded color: No sample files are available.
3. Icon line type indicates whether the sample is considered or excluded during analysis (relevant to samples with uploaded files only)
Solid: The sample is included in the analysis.
Dashed: The sample is ignored by Inheritance filters and the AI Shortlist algorithm, but you still can explore its genotypes.


Pre-built multigenome hash tables for GRCh37d5. The hash table builds include DNA, RNA, CNV, and HLA tables. Download here.
GRCh38/hg38: Multigenome Graph hg38-alt_masked.cnv.graph.hla.rna-9-r3.0.tar.gz.
Pre-built multigenome hash tables for hg38. The hash table builds include DNA, RNA, CNV, and HLA tables. Download here.
GRCh37d5: Multigenome Graph hs37d5-cnv.graph.hla.rna-9-r3.0.tar.gz.
Pre-built multigenome hash tables for GRCh37d5. The hash table builds include DNA, RNA, CNV, and HLA tables. Download here.
GRCh38/hg38: Multigenome Graph hg38-alt_masked.cnv.graph.hla.rna-8-r2.0-1.tar.gz.
Pre-built multigenome hash tables for hg38. The hash table builds include DNA, RNA, CNV, and HLA tables. Download here.
GRCh37d5: Multigenome Graph hs37d5-cnv.graph.hla.rna.tar.gz.
Pre-built multigenome hash tables for GRCh37d5. The hash table builds include DNA, RNA, CNV, and HLA tables. Available on demand.
GRCh38/hg38: GCA_000001405.15_GRCh38_no_alt_analysis_set.fna.gz.
Contains the sequences of the chromosomes, the rCRS mitochondrial sequence, unlocalized scaffolds, and unplaced scaffolds. Download here.
GRCh37/hg19: hs37d5.fa.gz.
Includes data from GRCh37, the rCRS mitochondrial sequence, Human herpesvirus 4 type 1 and the concatenated decoy sequences. Download here.
Genome view tab—provides an interactive overview of genomic structure, ideal for analyzing CNV and ROH/LOH events
Analysis tools tab—provides numerous customizable filters to help you explore the total list of genetic variants in compliance with your organization's standard case review process. You can export shortlisted variants in .xlsx format
Versions tab—documents versions of all the resources used during case analysis
Select the new status you want to apply.
Open the Cases tab.
In the Cases table, locate the relevant case row and select the case status.
From the dropdown menu, select the new status.
*
:
Average coverage
% Bases with coverage >10x
% Bases with coverage >20x
*Available only for whole genome FASTQ cases.
(overall)
For eligible cases, users can also review the results of DRAGEN QC: interactive DRAGEN QC report and DRAGEN QC metric files.
(overall)
*Available only for whole genome FASTQ cases.
Sample quality (overall)
:
Average coverage
% Bases with coverage >10x
% Bases with coverage >20x
Warning: Once trash folder is emptied, this action cannot be undone. Review cases pending deletion before proceeding!
Use these options to customize the Cases table view for your workflow.
Click Fields.
In the Fields menu, use the toggle switch next to each field name to show or hide columns based on your preferred view.
In the Cases table, click the column title you want to hide.
From the dropdown menu, select Hide column.
You can reorder columns in three ways: drag and drop the column, reorder columns via the Fields menu, or move a column using a dropdown menu.
Hover over the column title.
Click the six-dot icon () that appears to the left of the title.
Drag and drop the column.
Click Fields in the Cases table navigation panel.
In the Fields menu, hover over the field name.
Click the six-dot icon () that appears to the left of the title.
Click the column header.
From the dropdown menu, select Move left or Move right.
Hover over the left or right border of the column header cell.
When the resize cursor () appears, click and drag the border to your desired width.
When creating a new case, the first step is to select the sample input type. This determines how your data will be processed and which quality metrics will be available later in the analysis.
You can choose from the following supported formats: FASTQ, Project VCF, and VCF.
Use this option if you want the platform to perform secondary analysis and variant calling.
Accepted file types:
.fastq.gz
.fq.gz
.bam
.cram. Make sure you understand the current limitation for using CRAM files by expanding the section below.
Use when working with a joint VCF file containing multiple samples.
Accepted file types:
.pvcf
.vcf
.pvcf.gz
Use for cases where variants have already been called externally, or for cytogenetic array inputs.
Accepted file types:
.vcf
.vcf.gz
.targeted.json
While creating a new case, you can choose whether to include secondary findings for the proband. This option is available on the Family Tree screen → Create family tree panel → Show Secondary Findings.
Secondary findings are genetic variants that are not related to the primary indication for testing. These variants are automatically assigned the Incidental tag when they meet American College of Medical Genetics and Genomics (ACMG)-defined criteria for reportable secondary findings.
A variant is automatically tagged as a secondary finding if it meets all of the following criteria:
Classification: Previously classified as pathogenic or likely pathogenic in ClinVar or Curate variant databases
Zygosity: Heterozygous or homozygous (only homozygous for the HFE gene)
Allele frequency: Less than 5%
Read depth: 10× or higher
Variant quality: Any value except LOW
Affected gene: Listed in the ACMG SF v3.2 or 3.3 gene list for reporting secondary findings (PMID: 37347242, 40568962)
ACTA2, ACTC1, ACVRL1, APC, APOB, ATP7B, BAG3, BMPR1A, BRCA1, BRCA2, BTD, CACNA1S, CALM1, CALM2, CALM3, CASQ2, COL3A1, DES, DSC2, DSG2, DSP, ENG, FBN1, FLNC, GAA, GLA, HFE, HNF1A, KCNH2, KCNQ1, LDLR, LMNA, MAX, MEN1, MLH1, MSH2, MSH6, MUTYH, MYBPC3, MYH11, MYH7, MYL2, MYL3, NF2, OTC, PALB2, PCSK9, PKP2, PMS2, PRKAG2, PTEN, RB1, RBM20, RET, RPE65, RYR1, RYR2, SCN5A, SDHAF2, SDHB, SDHC, SDHD, SMAD3, SMAD4, STK11, TGFBR1, TGFBR2, TMEM127, TMEM43, TNNC1, TNNI3, TNNT2, TP53, TPM1, TRDN, TSC1, TSC2, TTN, TTR, VHL, WT1.
Includes all v3.2 genes plus newly added genes:
PLN
ABCD1
CYP27A1
This brings the total to 84 reportable genes.
The ethnicities of the proband's mother and father can be specified during the process of UI or API case creation. Please refer to the following list of supported ethnicities.
A "Afghan Jews" "Afghani" "African" "African American" "Afro-Brazilian"
"Alaska Native" "Algerian" "Algerian Jews" "Amish" "Anatolian" "Arab" "Argentinian/Paraguayan" "Armenian" "Ashkenazi Jews" "Asian" "Asian Brazilian" "Australian Native" "Azerbaijan Jews"
B "Bedouin" "Bengali/Northeast Indian" "British/Irish" "Bulgarian Jews"
C "Caribbean Australian"
"Caucasus Jews" "Central African" "Central Asian" "Chilean" "Chinese" "Chinese Dai" "Christian Arab" "Circassian" "Colombia"
D "Druze" "Dutch"
E "East African" "East Asian" "East European" "Egyptian" "Egyptian Jews" "Emirates" "Ethiopia" "Ethiopian / Eritrean" "Ethiopian Jews" "Ethiopian Jews - Beta Israel" "European" "European American"
If you're comfortable with scripting and API usage, you can upload multiple cases at once using those methods. But if you're not a technical expert, don't worry. There is a user-friendly alternative available—importing a CSV file directly through the user interface.
Please follow the steps as described below.
Caution: Please note that refreshing or leaving the page, exiting the Add new case tab, or power failure of your computer before you've completed a batch case upload will result in loss of the case creation progress.
CSV (Comma-Separated Values) is a simple file format used to store data in tabular form. A row represents a sample, and a column represents a data field.
Start by downloading a CSV template with an example line and mandatory and non-mandatory fields from the Add new case page set to Batch mode (see step 2). Fill the file with your data according to CSV format requirements.
Click on the + New case button on the top navigation panel.
Click on the Switch to batch button in the top right corner. You'll be directed to the Select file page of the Batch upload flow. Note: Here you can download a CSV template in the valid format.
Drag and drop a CSV file into the box or upload it from the file explorer. Wait for file upload and validation to finish.
After validation is complete, you will be directed to the Batch validation page. It features validation results details for you to review:
File name,
Number of rows in the file,
Number of cases to be created
Number of errors found,
Click on Create. A progress bar will appear on the right as the cases are created (Cases creation page).
If the cases have been created successfully, the Cases summary page will display the total number of cases that were created.
If there were any errors during the batch case creation process, the Cases summary page will display a table indicating the number of cases that were successfully created and the number of cases that failed.
You will have the option to download a CSV file containing two additional columns: Errors and Case ID. The Errors column will contain error messages for samples where case creation failed, while the Case ID column will contain the Case ID of a successfully created case for the lines where case creation was successful.
The Sex validation column indicates whether the biological sex inferred from genomic data matches the sex information provided during case creation. This helps identify potential sample mix-ups or metadata errors before interpretation begins.
Sex validation results:
Pass
Reported sex matches the estimated sex
Fail A mismatch was detected between reported and estimated sex.
N/A QC file not available; validation could not be performed.
Array sample quality metrics provide an overview of QC validation and call-rate results for each sample.
(overall)
The Case info tab includes the following information:
Case ID—a unique identifier assigned to each case by Emedgene, formatted as EMGXXXXXXXXX
Case type—the type of analysis performed:
The Activity tab offers a timeline of case actions and enables users to leave comments. It supports key functions that enhance case management and review:
Traceability—Maintains a complete, time-stamped history of case actions
Error recovery—Allows users to identify and trace changes, such as variant edits or disease associations, made in error
Real-time collaboration – Enables teams to monitor each other’s updates as they happen, ensuring transparency
Before you proceed to this article, make sure you understand .
In > Management Tab, add or edit the required credentials: CLIENT_ID, CLIENT_SECRET, TENANT_ID, and ACCOUNT_URL.
See the table below to learn where to look for them in your Azure account.
Case status reflects the current stage of case processing, either by the Emedgene platform or your team. Statuses enable case progress tracking and support a consistent, collaborative case review workflow.
Starting in v100.40.0, the Status column on the Cases page includes a for cases with the In progress, Re-Analysis, or Issue reported status.
You can view and the current case status in the Cases table and in the top bar of the individual case page.
The Candidates tab displays all tagged variants, whether tagged by the AI Shortlist or manually by a user.
Variants are automatically tagged as:
Most Likely Candidates and Candidates
Variants prioritized by the AI Shortlist
Secondary findings
Variants that meet ACMG-defined criteria for secondary findings and automatically tagged with an Incidental tag (if enabled)
The Contamination column reports whether a sample shows signs of DNA contamination, helping ensure data reliability before interpretation.
Contamination is detected using calculations, which estimate the proportion of reads that do not match the expected genotype. This estimate is based on the idr_baf score.
idr_baf stands for the interdecile range of the B-allele frequency—calculated as the difference between the 90th and 10th percentiles of the distribution of alt / (ref + alt) ratios across all variant sites.
A larger idr_baf value indicates greater variability in allele balance, which may suggest sample contamination, particularly from another human DNA sample.
No contamination detected: idr_baf < 0.200.
The Ploidy column shows results from the DRAGEN Ploidy Estimator.
The estimator detects aneuploidies and infers sex karyotype in whole genome cases.
Ploidy values come from the DRAGEN *.ploidy_estimation_metrics.csv output file.
All autosomes fall within the expected ploidy range. No large-scale autosomal copy number deviation is detected.
At least one autosome has a median ploidy score below 0.9 or above 1.1.
Hover over the result to identify the affected chromosomes.
Ploidy metrics are not displayed in Sample quality. This occurs when:
The Sex validation column indicates whether the biological sex inferred from genomic data matches the sex information provided during case creation.
This helps identify potential sample mix-ups or metadata errors before interpretation begins.
Sex validation results:
Pass
Reported sex matches the estimated sex
Fail A mismatch was detected between reported and estimated sex.
This page explains the requirements for viewing the DRAGEN QC report for each case type.
Run a FASTQ case in Emedgene.
Result: Because DRAGEN analysis is integrated into Emedgene secondary analysis pipeline, QC reports are automatically generated in the system.
Prerequisite: or
Run DRAGEN analysis externally. This approach is referred to as "Bring your own DRAGEN (BYOD)".
The is generated by the Illumina DRAGEN Bio-IT Platform and covers the entire analysis workflow—from raw sequencing reads to variant calls.
A visual summary that includes interactive plots of key quality metrics.
When available, a DRAGEN report link appears below the sample name in the Sample quality section of the Lab tab.
Clicking the link opens the detailed quality control metrics report in a new browser tab. This integration allows users to assess sample quality directly from the Emedgene interface.
A set of detailed CSV files containing sample-level quality metrics. These files are and support in-depth review and documentation.
F "Fijian Australian" "Filipino" "Filipino Austronesian" "Finnish" "French" "French Canadian"
G "Georgian Jews" "Germans" "Ghanaian / Liberian / Sierra Leonean" "Greece Jews" "Greek Americans" "Greek / Balkan" "Guam/Chamorro"
H "Hawaiian"
I "Iberian" "India - Bene Israel Jews" "India - Cochin Jews" "Indian" "Indigenous Amazonian" "Indigenous peoples in Canada" "Indonesian" "Inuit" "Iranian" "Iranian Persian Jews" "Iraq" "Iraqi Jews" "Irish" "Italian" "Italian Americans" "Italian Jews"
J "Japanese" "Japanese Brazilian" "Jordan"
K "Kenyan" "Korean" "Kurdish" "Kurdish Jews"
L "Latino/Hispanic Americans" "Lebanese Jews" "Levantine" "Libyan" "Libyan Jews"
M "Maasai" "Malayali Indian" "Melanesian" "Mesoamerican and Andean" "Mexican American" "Middle Eastern" "Mongolian / Manchurian" "Mormon" "Moroccan" "Moroccan Jews" "Muslim Arab"
N "Native American" "Nepali" "Nigerian" "North African" "North and West European" "Northern Asian" "Northern Indian"
O "Other Pacific Islander"
P "Pakistani" "Papuan" "Polynesian" "Portuguese in Northern Brazil" "Portuguese in Southern Brazil"
R "Russian Jews" "Russians"
S "Samaritan" "Samoan" "Sardinian" "Saudi" "Scandinavian" "Senegambian / Guinean" "Siberian" "Somali" "South African" "South Asian" "Southern East African / Congolese" "Southern European" "Southern Indian" "Southern Indian / Sri Lankan" "Southern South Asian" "Spaniards" "Spanish Jews" "Sub-Saharan African" "Sudanese" "Swedes" "Syrian Jews" "Syrian-Lebanese"
T "Tajikistan Jews" "Thai / Cambodian / Vietnamese" "Tunisian" "Tunisian Jews" "Turkish" "Turkish / Anatolian" "Turkish Jews"
U "Ukraine" "Ukraine Jews" "Uzbekistan/ Bukharan Jews"
V "Venezuela"
W "West African"
Y "Yemenite" "Yemenite Jews"
Lab workflow settings
Manage presets and preset groups.
Default preset group
Set the preset group applied automatically when no group is selected.
Presets tab
Use predefined combinations of filters that reflect your laboratory’s SOPs.
If the sex was marked as unknown during case creation, the system will display the predicted sex instead of a validation status.
If no errors were detected, a success message will be displayed
If any errors were detected, an error message will be displayed.
You will be given the option to download a file with error details to help you diagnose and correct any issues with the data. Once you've corrected the CSV file, reupload it.
API/batch upload limitations
When using the API or batch upload, note that applying multiple gene lists can inadvertently exceed a combined limit of 10,000 genes across panels. The platform may not provide an explicit error message in such cases. Plan gene-panel combinations carefully.
Combining gene lists at case creation is available via the UI only and cannot be performed through API/batch upload.
API/batch upload cannot add phenotypes for an unaffected parent.
JSON files cannot be uploaded via API/batch upload.













Case status
Overview of case status behavior, history, and related workflows.
Case statuses in a case lifecycle
Review status transitions, control types, and the reference status list.
Case progress indicator (v100.40.0+)
Understand progress indicator states for cases in processing or failure states.
Case status management
Create custom statuses and reorder them for your organization.


Drag and drop the field.






When Emedgene was first released, the term “incidental findings” was adopted in alignment with genomics standard at the time. The 2013 ACMG recommendations defined incidental findings as “the results of a deliberate search for pathogenic or likely pathogenic alterations in genes that are not apparently relevant to a diagnostic indication for which the sequencing test was ordered” (PMID: 23788249).
As the field evolved, the ACMG and broader community began to distinguish between “incidental findings” (unexpected, not actively sought) and “secondary findings” (intentionally analyzed and reportable). This shift was reflected in the updated 2016 ACMG guidance (PMID: 27854360).
To reflect this change, Emedgene introduced the term “secondary findings” into the platform. However, “incidental findings” remains in use throughout the platform for technical consistency.
Warnings:
Secondary findings are limited to the ACMG-defined gene lists. Variants outside these lists will not be tagged automatically.
Only variants with adequate sequencing depth and quality are tagged. Low-quality calls may require manual review.

v100.41+: The sample contains fewer than 10000 variants. Below this threshold, there are too few variants to validate ploidy reliably.
The case is a whole genome Bring your own DRAGEN (BYOD) VCF case.
The case was run with the pipeline older than v32.
Ploidy is available in the Ploidy column only for whole genome FASTQ cases.
Ploidy appears in the DRAGEN QC report and the Ploidy column.
The DRAGEN pipeline generates *.ploidy_estimation_metrics.csv. Its values populate Sample quality.
Ploidy displays N/A in the Ploidy column.
Metrics in the supplied *.metrics.tar.gz archive generate only the DRAGEN QC report. They do not populate Sample quality.
Review ploidy in the DRAGEN QC report for these cases.
Check ploidy early in case review. It can identify potential large-scale chromosomal abnormalities.
Compare the inferred sex karyotype with sex validation. This helps identify possible sample swaps.
A failed result does not confirm an abnormality. Interpret it with other QC metrics and genomic visualizations.
N/A QC file not available; validation could not be performed.
Sex validation is performed by comparing the observed homozygous/heterozygous genotype ratio on the X chromosome with the expected ratios:
<2 for females
>2 for males
Prerequisites:
Only high-quality SNVs from targeted regions—either kit-specific or RefSeq coding regions—are used for sex validation
A minimum of 50 variants is required to generate a reliable result. If this threshold is not met, sex validation cannot be performed, and no result is displayed
Emedgene uses the sex provided in the sample metadata to determine the expected copy number for sex chromosomes CNVs.
If the sample’s sex is marked as unknown, Emedgene defaults to Female for CNV calling.
If the predicted sex does not match the reported sex, we recommend updating it and re-analyzing the case.
If the sex was marked as unknown during case creation, the system will display the predicted sex instead of a validation status.
Prepare the DRAGEN QC data. Use one of these options:
DRAGEN report HTML file (v100.40.0+): Download the DRAGEN report as a .report.html file per sample.
or
DRAGEN metrics TAR file: Download DRAGEN QC metrics files and prepare a TAR archive as described here.
Upload the HTML file or TAR archive together with the sample VCF file.
Run the case.
Result:
If you uploaded an HTML file: The system directly visualizes the uploaded HTML file instead of generating the report from metrics files.
If you uploaded a TAR file: The system generates an interactive HTML report from metrics files.
Array cases start from VCF input files.
Prerequisites:
DRAGEN Array v1.3.0 and later
Emedgene v100.39.0 and later
Run DRAGEN analysis externally. This approach is referred to as "Bring your own DRAGEN (BYOD)".
Prepare the DRAGEN QC data. Use one of these options:
DRAGEN report HTML file (v100.40.0+): Download DRAGEN report as a .report.html file per sample.
or
DRAGEN metrics files: Download the .annotated_cyto.json DRAGEN QC metrics file, the sample VCF file, and the .gt_sample_summary.json file.
Upload the HTML file or metrics files together with the sample VCF file.
Run the case.
Result:
If you uploaded an HTML file: The system directly visualizes the uploaded HTML file instead of generating the report from metrics files.
If you uploaded metrics files: The system generates an interactive HTML report from metrics files.
.vcf.gz.gt_sample_summary.json (DRAGEN Array v1.2+).annotated_cyto.json (v100.39.0+, DRAGEN Array v1.3+)
Tips:
Choose the input type carefully — it cannot be changed after the case is created.
Keep file paths simple (avoid spaces, parentheses, or very long names >255 characters). This helps prevent errors during upload.
Warning:
If files are incomplete or corrupted, the case may still be created but will fail during processing. Double-check your files before uploading.
For large files (BAM/CRAM/FASTQ), browser upload is not recommended. Use Batch Upload, CLI, or cloud-to-cloud transfer instead to avoid incomplete or truncated uploads.
Exome
Custom Panel
Array
Sample type—the format of the sample files used in the case:
FASTQ: *.fastq.gz, *.fq.gz, *.bam, *.cram.
Project VCF: *.pvcf, *.vcf, *.vcf.gz, *.pvcf.gz
VCF: *.vcf, *.vcf.gz, *.targeted.json, *.gt_sample_summary.json
Gene list—defines whether gene list was used during analysis and how it was applied:
All genes—AI Shortlist was neither confined to nor prioritized a specific gene list
Virtual panel (In silico panel)—AI Shortlist was limited to only the genes in the gene list
Boosted gene list—AI Shortlist analyzed variants in all genes, but variants in the gene list were given higher priority
Analysis type:
If field is not present—carrier analysis was not performed
Carrier—carrier analysis was performed for the selected gene list
Human reference—the genome reference used during case analysis
Ordered by—the user who created the case and the case creation date
Signed by—the user who finalized the case
Related cases—the Case IDs of other cases that share one or more samples with the selected case
Due Date—the user-defined deadline for finalizing the case. To enter or edit the Due Date, click the calendar icon
Participants—Users involved in the case, whether in submission, analysis, finalization, or those subscribed to updates. To receive email notifications, click the Subscribe icon. To unsubscribe, hover over your avatar and click the X icon
Patient Information—basic demographic details:
Sex. Specified by the user
Age. Automatically calculated in years based on the provided date of birth
Clinical Information:
Proband phenotypes—HPO terms used to describe clinical findings in the proband
Suspected disease—if provided, includes the suspected condition, penetrance (%), and severity (mild, moderate, severe, or profound)
Maternal and Paternal —ethnic background of the proband’s parents
Parental consanguinity
Additional case information can be added using custom fields, either via the API or by including extra columns in your CSV during batch case creation.
This allows you to extend the case details panel with project-specific data.
To enable this feature or learn more, please contact techsupport@illumina.com.
Training & quality control – Helps identify patterns in variant interpretation and supports consistent application of evidence criteria
Audit compliance – Supports clinical and laboratory documentation standards (e.g., CAP/CLIA) by providing a verifiable action history
Each activity entry includes:
Timestamp (date + time)
User name of the person who performed the action
Action description
Activity logs are kept for at least six years for full traceability.
Case-related
Case created Case status changed Case participants updated Case labels modified Report created Case moved to trash Case data edited, no reanalysis initiated Case data edited and reanalysis launched
Comments
Comments left in the Activity tab
Variant tagging
Viewing activity logs
In the Cases table, the Activity tab within the Case details panel displays only comments and case-related activities. To view the full list of all activities, open the Case details panel directly from the individual case page.
Edits are permanent. Even if a change is undone, the original action remains recorded for traceability
Logs are case-specific. Activity entries do not reflect changes made in other cases or in the Curate database
Time zone awareness. Timestamps follow the system’s configured time zone, which may differ from your local time—especially in international collaborations.
CLIENT_ID
application_id.
Format: ########-####-####-####-############
(letters/numbers)
CLIENT_SECRET
Value of the client_secret tuple (Value, Secret ID).
Format: #####-#######-######-######
(letters/digits/special chars)
TENANT_ID
ID of the tenant.
Format: ########-####-####-####-############
(letters/numbers)
ACCOUNT_NAME
An arbitrary name that the customer must supply to define the ACCOUNT_URL.
Format: string
CONTAINER_NAME
In Microsoft Entra ID, click on App registrations.
Select New registration.
Fill the name of the application & press "register."
You got to the registered app page: (CLIENT_ID / TENANT_ID) From this you can retrieve: Application ID and Tenant ID. Both are marked in the screenshot.
Press "Certificates & secrets"
Press on "New Client secret"
Fill the "Description" and change expires to 12 months. (or according to your organization policy), than press "Add"
8. Get the CLIENT_SECRET from this page.
Give this App registration roles and read access to the relevant Blob.
Go to Azure Storage accounts
Get into the relevant Storage account
Press on "containers"
Press on the relevant container
Press on "Properties"
Copy the ACCOUNT_URL
Errors for bad connections can be found in CloudWatch on particular FRY log stream
Search for: BlobApi, BlobFs, azure.

Each time a case status is updated, the change is logged and recorded in the case activity history.
To review case status updates:
In the relevant case, open the Case details panel.
Select the Activity tab.
Filter logs by selecting Case‑related activities from the dropdown list.
Result: The status history is shown, along with other case-level logs.


Carrier variants
Variants identified by the carrier analysis pipeline (if enabled)
During review in the Candidates tab, additional tags can be applied to a variant alongside the original automatic tag.
A set of the most promising variants based on scores calculated by the AI Shortlist. These variants are initially tagged by the system.
Variant types assessed:
SNVs and indels
CNVs
SVs
mtDNA variants
STRs
Secondary findings are variants that are automatically assigned the Incidental tag when they meet the criteria for secondary findings as defined by the American College of Medical Genetics and Genomics (ACMG).
Tagging is applied only when the Secondary findings checkbox is selected during case creation.
A variant is automatically tagged as an incidental (secondary) finding if it meets all of the following criteria:
Classification: Previously classified as pathogenic or likely pathogenic in ClinVar or Curate variant databases
Zygosity: Heterozygous or homozygous (only homozygous for the HFE gene)
Allele frequency: Less than 5%
Read depth: 10× or higher
Variant quality: Any value but LOW
Affected gene: Listed in the ACMG SF v3.2 or 3.3 gene list for reporting secondary findings (PMID: 37347242, 40568962)
ACTA2, ACTC1, ACVRL1, APC, APOB, ATP7B, BAG3, BMPR1A, BRCA1, BRCA2, BTD, CACNA1S, CALM1, CALM2, CALM3, CASQ2, COL3A1, DES, DSC2, DSG2, DSP, ENG, FBN1, FLNC, GAA, GLA, HFE, HNF1A, KCNH2, KCNQ1, LDLR, LMNA, MAX, MEN1, MLH1, MSH2, MSH6, MUTYH, MYBPC3, MYH11, MYH7, MYL2, MYL3, NF2, OTC, PALB2, PCSK9, PKP2, PMS2, PRKAG2, PTEN, RB1, RBM20, RET, RPE65, RYR1, RYR2, SCN5A, SDHAF2, SDHB, SDHC, SDHD, SMAD3, SMAD4, STK11, TGFBR1, TGFBR2, TMEM127, TMEM43, TNNC1, TNNI3, TNNT2, TP53, TPM1, TRDN, TSC1, TSC2, TTN, TTR, VHL, WT1.
Includes all v3.2 genes plus newly added genes:
PLN
ABCD1
CYP27A1
This brings the total to 84 reportable genes.
Variants identified by the Carrier analysis pipeline. Carrier variants are automatically tagged only if you've selected the Carrier Analysis checkbox while creating a case. Analysis requirements and a list of targeted regions are specified by the organization's manager. This Carrier analysis flow is implemented by request.
Variants that were manually selected to be reported.
When Emedgene was first released, the term “incidental findings” was adopted in alignment with genomics standard at the time. The 2013 ACMG recommendations defined incidental findings as “the results of a deliberate search for pathogenic or likely pathogenic alterations in genes that are not apparently relevant to a diagnostic indication for which the sequencing test was ordered” (PMID: 23788249).
As the field evolved, the ACMG and broader community began to distinguish between “incidental findings” (unexpected, not actively sought) and “secondary findings” (intentionally analyzed and reportable). This shift was reflected in the updated 2016 ACMG guidance (PMID: 27854360).
To reflect this change, Emedgene introduced the term “secondary findings” into the platform. However, “incidental findings” remains in use throughout the platform for technical consistency.
idr_baf < 0.241.Contamination suspected: 0.241 ≤ idr_baf < 0.300.
Contamination confirmed: idr_baf ≥ 0.300.
No data is available:
v100.41+: The sample contains fewer than 10000 variants. Below this threshold, there are too few variants to validate contamination reliably.
idr_baf = 0.000.
The case is an older case.
Hover over the value to display a tooltip showing the HET ratio (proportion of sites that are heterozygous) and the HET count (number of heterozygote calls in sampled sites).
Tips:
Always review contamination results before starting interpretation to rule out technical issues that could explain unexpected variant calls.
Cross-check contamination results with other QC metrics (e.g., depth, ploidy, sex validation) for a more complete picture of sample quality.
Warnings:
Panels:
v100.41+: Contamination is not performed for samples with fewer than 10000 variants and shows N/A.
Requirements by case type
Check DRAGEN QC requirements by case type.
Download DRAGEN QC metrics files
Download sample-level DRAGEN QC metrics files from the Lab tab.

Cases table lists key details of all genomic sequencing cases submitted by the organization.
You can customize the table by hiding, showing, rearranging fields, or adjusting column widths, except for Case ID, which is fixed as the first column and always visible.
Case ID
A unique case identifier (EMGXXXXXXXXX).
This field is fixed and cannot be hidden or repositioned in the table.
Please provide this code to Tech Support when reporting any issues.
Proband ID
Identifier of the proband.
For , this is the Sample Name; for , it is the BioSample Name of the test subject.
Go to the google cloud Console.
Navigate to IAM & Admin - In the left sidebar, go to IAM & Admin > Service Accounts.
Create a New Service Account: Click on the "Create Service Account" button at the top.
Fill in the Service Account Details:
Service account name: Give your service account a name.
Service account ID: This will be automatically generated based on the name.
Description: Optionally, provide a description for the service account.
Click "Create and Continue".
example:
Assign Roles to the Service Account:
In the Grant this service account access to project step, you’ll assign the necessary roles.
Grant these role:
"storage object viewer" (read-only access)
Add the above 3 values into the appropriate fields:
Client_credentials_base64: pasting the output of 8.
Bucket: the bucket name.
Path: for default, fill with / else, put your path in the bucket. Seperate directories with /
Download and install the Google Cloud SDK from the Google Cloud SDK Install page.
Select Your Platform (Windows, macOS, or Linux), download and run.
Initialize and Authenticate with Google Cloud: In the Cloud SDK Shell/terminal, run:
gcloud init
This will open a browser window to authenticate your Google account. Follow the instructions to log in and select your project.
notice:
origin: if using Illumina cloud:
https://host_name.emg.illumina.com
else, Emedgene cloud:
https://host_name.emedgene.com
Apply CORS Configuration to Your Bucket: run the next command.
gcloud storage buckets update gs://your-bucket-name --cors-file=cors.json
Verify the CORS Configuration:
gcloud storage buckets describe gs://your-bucket-name
Log in to Emedgene and navigate to Settings in the upper right-hand corner of the page.
Click on the Management tab and then on Add Storage.
Choose Illumina BaseSpace storage type.
Fill Client Key, Client Secret and App Token as provided from BaseSpace (a description on how to get this information is provided below) and click Add storage to complete the setup.
Install BaseSpace CLI (Command Line Interface)
Follow the instructions on the if needed. Be aware of the Basespace Regional Instance you are working on (us, euc1, aps2, euw2)
On BSSH, login to the workgroup you want to connect as the storage.
Once the BaseSpace CLI is installed, run the authentication command in the terminal.
The command will direct you to a link which requires to login.
After the authentication was completed successfully, find the access token in the config file.
The result should look like -
Populate the App_token with the accessToken value, and Server with the apiServer URL from the BSSH config file.
Client_key will be displayed in subsequent menus, so a descriptive name such as the workgroup name can be used.
Client_secret is unused when the App_token is available and can be set to "x".
Go to the BaseSpace and login. Be aware of the Basespace Regional Instance you are working on (us, , , )
Go to My Apps and click Create a new Application.
Fill details for the application and click on create an application.
Fill details and press save.
You will need to fill all the fields that it requested, please add “NA” to them.
Go to My Apps and click on your new app. Then go to the credentials tab.
You will find the Client ID (Client Key), Client Secret and App Token to enter to Emedgene platform.
Log in into the desired Emedgene organization.
Go to Settings
Go to Management tab
Click on Add Storage
Add the information from your “Credentials” of the App previously created in BSSH.
Note: The fields marked with (*) are mandatory.
Options: Male, Female, Unknown.
When a sample is user-assigned "Unknown" sex, the system assumes "Female". This affects CNV interpretation on sex chromosomes in case the genetic sex is actually male:
Chromosome X: CN = 2 is considered reference (REF) for a female genome, so CNVs with two copies are hidden by default. This may cause chromosome X duplications to be missed.
The default fixed value for Proband is Test Subject.
Expected format: mm/dd/yyyy.
Options: Affected, Healthy.
The default value for Proband is Affected, but you may change it to Healthy.
To add all relevant phenotypes for the Proband, use one of the following methods:
, or
Automatically infer disease-associated phenotypes (see below).
Please follow the steps described below for each phenotype:
Enter an HPO term (e.g., Hypoplasia of the ulna), an HPO ID (e.g., HP:0003022), or a descriptive phenotype name (e.g., Underdeveloped ulna) in the search box.
Select a matching term from a dropdown menu and press Complete after you've added all the terms and additional proband information below.
Paste a list of comma-separated HPO terms or HPO IDs in the search box and press Complete.
Enter the disease name in the search box, select a matching term from a dropdown menu and press Complete. All the associated phenotypes will be automatically added to the Proband Phenotypes.
Selecting a disease only fetches its associated phenotypes for convenience—it does not affect downstream analysis. You can edit this list to match the proband’s case. Only the phenotypes you keep or add influence analysis, not the disease selection itself.
To remove any phenotype described for the disease but not observed in the proband, click the button next to the HPO term in the Proband Phenotypes list.
Enter the suspected disease penetrance as a percentage.
Select the appropriate category to indicate the severity of disease symptoms observed in the proband: Mild, Moderate, Severe, Profound.
Mark the checkbox if applicable.
Paternal and Maternal. Enter the name in the search box and select a matching term from a dropdown menu.
A region of interest (ROI) BED file determines which genomic regions are included in variant analysis. It functions as a preprocessing filter, determining which variants proceed to annotation and interpretation.
If no custom ROI BED kit is applied to a case, the system applies a default ROI BED file based on the case type. All default ROI BED files are available for download (see Default ROI kit details).
A BED file covering a wide range of genomic regions. It contains:
"RefSeq ALL" transcripts and "GENCODE" full gene regions, with 5 Kbp upstream and 5 Kbp downstream
Within this range, all “Clinical Regions” are included
All dosage regions (HI/TS sig level 1, 2, or 3)
Moreover, liftover versions of both reference regions are included for the current and previous range versions.
Liftover is done using CrossMap (v0.5.2), chain hg19ToHg38.over.chain.gz
NCBI RefSeq regions are based on release 105 (hg19) and release 110 (hg38)
GENCODE regions are based on release V19 (hg19) and release V41 (hg38)
All microRNA genes are based on the HGNC miRNA definition from December 2022
Download files used in v100.39.0+
Download files used up to v38.0
This BED file includes regions relevant to disease interpretation, specifically:
“RefSeq Curated” and “GENCODE” regions with 50 bp flanking regions on each side of all exons, including coding exons and UTRs, for protein-coding genes
OMIM disease-related RNA genes (flanking 50 bp)
All ClinVar pathogenic variant regions (flanking 50 bp)
Promoter regions (EPDnew human version 006, flanking 50 bp)
For consistency, the GRCh38 version includes the lifted-over regions from GRCh37 (using CrossMap for liftover).
Download files used in v100.39.0+
Download files used up to v38.0
Each variant has a main_effect and main_gene chosen based on the most prioritized transcript for this variant. This selection influences how variants are displayed, interpreted and classified across the platform.
From 100.39 case pipeline and up, Emedgene introduces improvements to Curate transcript prioritization and updates the RNA gene prioritization logic.
Emedgene uses VEP and EFF for transcript annotations and supports organization-defined canonical and preferred transcripts from Curate.
VEP transcripts are prioritized over EFF transcripts.
If the case is a Virtual Panel, prioritize transcripts from genes in the case gene list (not applied for Boosted Genes panel types).
Prioritize transcripts defined in Curate variants
Curate variant-level preferred transcripts now receive high priority.
Requires the new organization setting, enabled by Illumina Bioinformatics support.
Prioritize RNA genes associated with disease
(See Appendix 1: Updated RNA gene list)
This rule does not apply to upstream or downstream RNA variants.
RNA gene prioritization has been refined in 100.39 case pipeline
De-prioritize readthrough biotype transcripts.
Prioritize intronic based on variant impact:
HIGH → MODERATE → LOW → MODIFIER
Prioritize intronic > UTR > upstream effects
(See Appendix 2 for MODIFIER effect prioritization)
Prioritize organization canonical transcripts
Defined in Curate
Always applied; no additional settings required
Prioritize canonical transcripts based on APPRIS.
Prioritize transcripts from genes in the case gene list.
Prioritize genes without a “ — ” in their symbol.
From v100.39, Emedgene has changed how RNA genes are prioritized relative to protein-coding genes.
Prior to v100.39, if a variant overlapped an RNA gene from the prioritized list, the RNA transcript was often chosen as the main_gene, even when a protein-coding gene had a more impactful variant.
Starting with v100.39 case pipeline
Protein-coding genes with stronger effects now take priority over RNA genes.
RNA genes are still considered, but no longer override coding transcripts with higher significance.
This results in more more clinically meaningful main_gene selection.
Here is a list of ordered rules for transcript prioritization:
VEP transcripts are prioritized over EFF transcripts.
If the case is a virtual panel, prioritize transcripts from genes in the case gene list (but not for Boosted Genes type panels).
Prioritize RNA genes associated with disease (See appendix 1 for prioritized list RNA genes). Importantly this does not apply to upstream and downstream RNA variants.
Variant effect
For each variant that is mapped to the reference genome, Emedgene uses Ensembl’s Variant Effect Predictor (VEP) and the RefSeq (NCBI) library of transcripts to calculate variant effect. VEP uses a set of consequence terms defined by the Sequence Ontology (SO), including immediately recognizable terms like “missense_variant” and “frame_shift_variant” as well as some more esoteric ones like “non_coding_transcript_exon_variant”.
The full list of terms, along with detailed descriptions and severity impact categories can be found in the link below.
Importantly, each variant has a "main_effect" and "main_gene" chosen based on the most prioritized transcript for this variant. Transcript prioritization depends on many different parameters and on different Emedgene pipeline versions as described here.
Variant severity
Variant severity, also known as variant impact, is a subjective assessment of the severity of a variant consequence.
Severity is usually categorized as modifier, low, moderate or high:
Modifier severity is used for non-coding variants or variants affecting non-coding genes, where predictions are difficult or there is no evidence of impact. Inter-genic and non-coding variants are classic examples.
Low severity is used for variants that are assumed to be mostly harmless or unlikely to change protein function. This includes synonymous variants.
Moderate severity is used for non-disruptive variants that might change protein effectiveness, such as missense variants and in-frame insertions/deletions.
High severity is used for variants that are assumed to have a disruptive impact on abundance protein, such as by causing protein truncation, loss of open reading-frame, and/or triggering nonsense mediated decay.
Most of the time, variant effect and variant severity on Emedgene are consistent with VEP. However, genomics is a field defined by exceptions. There are key factors, outlined below, the Emedgene genetic team believes are critical to account for when assigning severity.
For small variants (SNV):
Splice prediction: Small variants will be upgraded to HIGH severity if its is high or MODERATE if its splicing prediction is moderate.
Conservation: Synonymous variants and splice region variants that are will be upgraded to MODERATE.
Non-coding RNA disease genes: The severity of a small variant will be upgraded to MODERATE if the variant is within a known to be associated with disease.
For CNV/SV:
VEP annotates CNVs with overlapping genomic features and designates them with the following effects: transcript amplification (DUP), feature elongation (DUP, INS), feature truncation (DEL), and transcript ablation (DEL). However, the severity assigned by VEP for CNVs does not reflect the complexity of CNV effects on protein function and in our experience is not suitable for genome analysis and filtering.
On Emedgene, variants are annotated in regards to its overlap with three different types of regions: ‘coding regions’, ‘clinical regions’, and ‘full gene’ region (see for a more detailed description about the BED files used in the system).
The region annotation is then used to assess severity for CNV and SV as follow:
Table 1: CNV/SV severity table. For each category of CNV/SV, the types of regions that overlap a given variant required to trigger the severity classification are shown.
For STR variants:
Emedgene is using an internal annotation for STR variants. More details can be provided by request to techsupport@illumina.com.
Known limitations
List of RNA genes known to be associated with disease is updated overtime as part of pipeline update.
Emedgene does not provide VEP annotation for non-coding regulatory data.
If you have an Enterprise account and you would like Emedgene-managed DRAGEN solution to save the DRAGEN output files in your own bucket, reach out to .
Emedgene directly from your AWS S3 bucket. In order to do it, you should enable for the Emedgene application URLs.
This guide provides a step-by-step process for creating a new case via the user interface. Detailed instructions for each step are available in the corresponding pages of the .
Click on the Add New Case button on the top navigation panel.
At the page, choose the file type for your case analysis (FASTQ, gVCF, or VCF).
Click Next to proceed.
The Case quality section shows the results of and the . Case quality is based on validation of the proband sample.
Use case quality to judge, at a glance, whether a case has enough high-quality, well-annotated variant data to support interpretation.
Case quality is based on five checks of the proband sample:
Chromosome validation
gnomAD validation
The overall sample quality indicator summarizes sequencing reliability for each sample. Sample quality is evaluated using different metrics, depending on the sample file type.
The overall FASTQ sample quality reflects the confidence level that the sequencing data can support accurate variant calling.
For samples processed from FASTQ files, quality is determined based on the lowest result among these metrics:
When creating NGS cases that start from VCF, you can create a browsable from the DRAGEN metrics files. Due to security restrictions, CSV files are not directly ingested, but they can be included when packaged in a TAR file.
Navigate to the local directory that contains the metrics files for a specific sample.
Define the sample name as a variable:
samplename="NA12878"
For family cases, check that no contamination is flagged before relying on inheritance-based filters.
v100.40 and earlier: Contamination estimates may be less reliable because fewer variants are available. Cross-check with other QC metrics when interpreting these results.
Do not use in isolation: A "Likely" or "Yes" result should not immediately be considered diagnostic — review case setup, sequencing quality, and sample handling first.
Case statuses in a case lifecycle
A visual overview of case status transitions and a reference description of each status.
How to update a case status
Update case status.
Case status management
Create custom case statuses and reorder them according to your workflow.
Case progress indicator (v100.40.0+)
Understand progress indicator states for cases in processing or failure states.
Issue reported status - error codes
Use this reference to identify and resolve errors that occur when the Emedgene pipeline processes a case.
De-prioritize biotype readthrough transcripts.
Prioritize based on impact in the following order: HIGH > MODERATE > LOW > MODIFIER.
Prioritize introns over UTR over upstream (Appendix 2: MODIFIER effects prioritization).
Prioritize organization canonical transcripts (Defined in Curate. Always applied, no settings needed).
Prioritize canonical transcripts (Based on Appris).
Prioritize transcripts from genes in the case gene list.
Prioritize gene without "-" in their Name.
Warning: If Curate preferred transcripts are enabled, transcript selection may differ from previous versions. This may change the displayed main gene/effect for some variants.
Tip: To ensure consistent results across teams, confirm whether Curate transcript prioritization is enabled for your organization.
Variant tag updated — this log entry includes a link to the relevant variant page for immediate review
Evidence notes
Evidence notes updated — this log entry includes a link to the relevant variant page for immediate review
Evidence pathogenicity
Variant pathogenicity updated — this log entry includes a link to the relevant variant page for immediate review
Evidence graph
Evidence graph updated — this log entry includes a link to the relevant variant page for immediate review
ACMG pathogenicity
ACMG evidence updated (logs any changes made via the ACMG classification wizard) — this log entry includes a link to the relevant variant page for immediate review
Transcript changes
Reference transcript updated
Report secondary findings—specifies whether secondary findings analysis was requested
Sharing consent
Clinical note—free-text notes provided at the time of launching the analysis
The mapping/alignment stage (which produces the CRAM file)
The variant calling stage (Emedgene secondary pipeline)
A mismatch in reference genome assembly files prevents the system from decompressing the CRAM file, leading to case analysis failure.
Best practices
Confirm reference compatibility with your organization settings before launching a run
If you receive CRAM files from an external lab, verify the specific reference genome file used to generate them
If the reference is unknown or incompatible, convert CRAM → BAM and upload the BAM file instead
ClinGen dosage regions, December 2022
Promoters from EPDnew human version V6
mtDNA CRS
RNA disease genes based on OMIM and HGNC (Dec 2022): ATXN8OS, TERC, IL12A-AS1, FAAHP1, NUTM2B-AS1, GAS8-AS1, RNU12, MIR204, IGHG2, SLC7A2-IT1, MIR99A, RMRP, XIST, MEG3, DIRC3, MIR17HG, GNAS-AS1, LRTOMT, LINC00299, DUX4L1, MIR137, MIR140, MIR605, SNORD118, RNU4ATAC, HELLPAR, IGHG1, IGHM, MIR19B1, RNU7-1, LINC00237, MIR2861, MIR4718, IGHV3-21, IGHV4-34, IGKC, KCNQ1OT1, MIR184, MIR96, H19, HYMAI, PCDHA9, UGT1A1, AFG3L2P1, DISC2, SNORA31, TRU-TCA1-1, PCDHGA4, TRAC, ECEL1P3, MIAT
ClinVar variants (ClinVar Dec 2022) with any pathogenic or likely pathogenic significance, and some drug responses associated with pathogenicity
50K STR regions based on the DRAGEN 4.0 Specification file
Known STR regions (DRAGEN 4.0 specification file)
All microRNA genes (flanking 50 bp, based on HGNC)
Full mtDNA region
Research Genome
None
Whole Genome
Exome
Custom Panel
An arbitrary name that the customer must supply to define the ACCOUNT_URL.
Format: string
ACCOUNT_URL
The account_url of the Azure account.
Format: https://account_name.blob.core.windows.net/container_name








Gain (DUP)
Intragenic (coding regions but not entire gene region)
Coding Regions / Clinical Regions not intragenic
Full gene and not in Clinical Regions
No overlap with any BED
Insertion (INS)
Coding regions
Clinical Regions and not in Coding regions
Full gene and not in Clinical regions
None
Deletion (DEL)
Coding regions
Clinical Regions and not in Coding regions
Full gene and not in Clinical Regions
No overlap with any BED












Create the Service Account:
After assigning the roles, click "Done".
Generate and Download a Key:
Find your newly created service account, click the three dots on the right, and select "Manage Keys".
Click Add Key > Create New Key and choose the JSON format.
Download the key and store it securely, as it is used for authentication in your code or applications.
Encode the key in base 64:
use python function: put this function and your json (here named json_file.json) in the same directory and run.\
save the output printed.
cors.json) on your machine with the CORS rules.
Example\ it should look like:



# Linux
$ wget "https://launch.basespace.illumina.com/CLI/latest/amd64-linux/bs" -O $HOME/bin/bs
# Mac
$ wget "https://launch.basespace.illumina.com/CLI/latest/amd64-osx/bs" -O $HOME/bin/bs
# or
$ brew tap basespace/basespace && brew install bs-cli
# Windows
$ wget "https://launch.basespace.illumina.com/CLI/latest/amd64-windows/bs.exe" -O bs.exe$ bs auth$ cat .basespace/default.cfgapiServer = https://api.basespace.illumina.com
accessToken = import json
import base64
def encode_json_to_base64(json_file):
# Read JSON data from file
with open(json_file, 'r') as file:
json_data = json.load(file)
# Convert the JSON data to a string
json_str = json.dumps(json_data)
# Encode the string to bytes, then to Base64
json_bytes = json_str.encode('utf-8')
base64_bytes = base64.b64encode(json_bytes)
# Convert Base64 bytes back to a string
base64_str = base64_bytes.decode('utf-8')
# Print the Base64-encoded string
print(base64_str)
encode_json_to_base64('json_file.json')[
{
"origin": ["https://<host_name>.emg.illumina.com"],
"method": ["GET"],
"responseHeader": ["emgauthorization"],
"maxAgeSeconds": 3600
}
]To include these variants in the analysis, enable the Include Reference Homozygosity and No Coverage Calls toggle in Workbench & Pipeline Settings.
Warning: Select valid HPO phenotypic abnormality terms
When adding proband phenotypes, ensure that all selected HPO terms originate from the “Phenotypic abnormality (HP:0000118)” branch of the HPO ontology. Terms outside this branch are not supported for case analysis.





Reanalysis will fail (will be fixed)
FASTQ
CRAM (Output)
Reanalysis will fail
FASTQ
VCFs
Reanalysis will fail
FASTQ
CSV, etc
Reanalysis will fail
VCF
BAM/CRAM (visualizations)
Visualization will fail
VCF
VCF (input)
Reanalysis will fail
VCF
CSV, etc
Reanalysis will fail (will be fixed)
This feature is only related to saving Dragen output files in your own bucket when using Dragen through Emedgene (without ICA).
If you are looking to:
Import data from AWS S3 to Emedgene go to Manage data storages
Integrating any data storage to Emedgene go to Manage data storages
Download any data from Emedgene go to Manage S3 credentials
Bring Your Own Bucket, also known as BYOK, enables you to control your DRAGEN file outputs.
Emedgene-managed DRAGEN solution saves the DRAGEN output files in a detected AWS S3 bucket that you have access to using your S3 credentials.
However, if you have an Enterprise account and you would like Emedgene-managed DRAGEN solution to save the DRAGEN output files in your own bucket, reach out to techsupport@illumina.com and follow this steps:
Emedgene requires access to the root folder, which means a dedicated bucket might be appropriated.
Bucket policy should allow Emedgene user access to the bucket.
Example bucket policy:
Emedgene visualizes data in IGV directly from your AWS S3 bucket. In order to do it, you should enable CORS for the Emedgene application URLs.
Example CORS policy:
We will require to run a case and validate the managed DRAGEN pipeline finish successfully and all features are available in the platform.
If a customer enables an AWS S3 Lifecycle policy in order to archive or change the S3 tiers for different files, they might create an adverse effect on the platform.
FASTQ
FASTQ/BAM/CRAM (input)
Reanalysis will fail (will be fixed)
FASTQ
CRAM (Output)
FASTQ
FASTQ/BAM/CRAM (input)
{
Coming Soon
}{
Coming Soon
}The page is divided into two panels: Create family tree (left) and Add patient information (right).
Use the visual tool to build the pedigree.
Add Clinical Notes (optional) in free text.
Select suspected Inheritance mode(s) (for record only; not used in the analysis).
Decide whether to include Secondary findings in the proband for the AI Shortlist (checkbox).
For each family member:
Add a sample (use a unique file path unless reusing samples).
Fill in a sample name (for VCF input, this must match the header in the file).
Complete the required subject details: for a proband and for non-proband samples.
Click Next to proceed to the Case info screen.
Here you define how the analysis will run:
Case type: Choose Array, Custom Panel, Exome, Whole Genome, or Other.
For Exome cases, variants outside exons ±50 bp are automatically filtered.
Carrier Analysis: Optional checkbox. Requires a targeted gene list.
Select an enrichment kit (if applicable) or "No kit".
If provided, kit details (Lab, Machine, Reagents, Expected coverage) will be used to compare coverage depth and breadth.
If no kit is provided, RefSeq coding regions will be used as reference.
options:
All genes
Phenotype-based genes
Existing gene list
: Select the Preset group appropriate for this case type.
If none is selected, the default Preset group is applied automatically (marked as default).
Consent: Confirm subject consent for extended sharing.
Additional case info (optional):
Indication for testing (free text).
Labels (choose from predefined organization labels; these cannot be changed later).
At the Summary stage, confirm case type, gene list, and other selections.
After the case is created:
The Case ID is displayed.
You may add participants so colleagues receive notifications on status changes or updates.
Caution: Please note that refreshing or leaving the page, exiting the Add new case tab, or power failure of your computer before you've completed adding a new case will result in loss of the case creation progress.
The Add New Case flow does not validate that sample IDs are unique or that input files are uncorrupted. Please ensure sample IDs are unique and that input files are valid before creating the case.
If a QC metrics file (metrics.tar.gz) is uploaded from BSSH, it will not be processed.
Caution: Clicking Next here will finalize case creation. After delivery, only the proband’s phenotypes can be edited without reanalysis.
ClinVar validation
AI Shortlist validation
mtDNA reference validation
Each check returns one of three results:
Passed
Failed
Skipped — the check does not apply to this case and does not affect case quality.
Passed: Every checked chromosome has at least one high-quality variant.
Failed: At least one checked chromosome has no high-quality variants.
Skipped: The check does not apply to this case.
Passed: Every checked chromosome has at least one gnomAD-annotated variant.
Failed: At least one checked chromosome has no gnomAD-annotated variants.
Skipped: The check does not apply to this case.
Passed: Every checked chromosome has at least one ClinVar-annotated variant.
Failed: At least one checked chromosome has no ClinVar-annotated variants.
Skipped: The check does not apply to this case.
Passed: The AI Shortlist tagged at least one variant.
Failed: The AI Shortlist tagged no variants.
Skipped: The gene list has fewer genes than the gene list threshold.
Passed: Mitochondrial DNA variants use the rCRS reference.
Failed: Mitochondrial DNA variants use a different reference.
Skipped: The case has no mitochondrial DNA data.
The overall case quality status appears:
In the Summary dashboard on the Lab tab.
Next to the Case quality section title on the Lab tab.
In the Quality column of the Cases table.
The overall case quality status depends on whether any individual check failed.
Pass: Every check either passed or was skipped.
Pass with exceptions (v100.41+) or Fail (v100.40 and earlier): At least one check failed.
Evaluate sex validation
The system reviews the sex validation result for the sample:
If sex validation failed, the system evaluates the overall FASTQ sample quality as low, and no further metrics are considered.
If sex validation passed or is unavailable, the system proceeds to evaluate coverage metrics.
Evaluate coverage metrics
The system reviews the results for average coverage and the percentage of target bases with coverage ≥20× and evaluates them as low, moderate, or high using the thresholds in Table 1.
Assign overall quality
The system uses the lower result to assign the overall FASTQ sample quality.
Table 1. Overall FASTQ sample quality thresholds
Low
Failed
≤30×
<70%
The overall VCF sample quality reflects the confidence level that the sequencing data can support accurate variant interpretation.
For samples processed from VCF files, quality is determined based on:
Evaluate sex validation
The system reviews the sex validation result for the sample:
If sex validation failed, the system evaluates the overall VCF sample quality as low, and no further metrics are considered.
If sex validation passed or is unavailable, the system proceeds to evaluate error rate.
Evaluate error rate
The system reviews the results for error rate and evaluates them as low, moderate, or high using the thresholds in Table 2.
Assign overall quality
The system uses the error rate result to assign the overall VCF sample quality.
Table 2. Overall VCF sample quality thresholds
Low
Failed
>45%
Moderate
Passed or N/A
Combine the find and tar commands to package the files into a *.metrics.tar.gz file.
Use this command to find files matching the required patterns:
find . \( -name "*.csv" -o -name "*.tsv" -o -name "*.counts" -o -name "*.counts.gz" -o -name "*.counts.gc-corrected" -o -name "*.counts.gc-corrected.gz" -o -name "*.ploidy.vcf" -o -name "*.correlation.txt.gz" -o -name "*.correlation.txt" -o -name "*.repeats.vcf" -o -name "*.ploidy.vcf.gz" -o -name "*.repeats.vcf.gz" -o -name "*.annotated_cyto.json" \) | xargs tar -czf "${samplename}.metrics.tar.gz"Upload the metrics.tar.gz file to the storage location used for case creation.
Add metrics.tar.gz to the case creation API JSON payload using the corresponding storage ID.
If the extension is not included in the filename, such as for files from BaseSpace, set "sample_type": "dragen-metrics" in the JSON payload.
Phenotypes
Phenotypes submitted for the proband.
Status
Current case status.
You can update the status directly from the table.
Type
Case type: whole genome, exome, custom panel, or array.
Label
Custom case labels. Click the pencil icon to add a new label, select an existing one, or remove a label from the case.
Quality
Overall case quality:
Pass
Pass with exceptions (v100.41+) or Fail (≤v100.40)
Not available
Hover over the icon for a brief summary, or view detailed results in the Lab tab. Sortable (Pass > Pass with exceptions > Fail > Not available).
Creation date
Date the analysis was initiated. Sortable.
Due date
Customizable due date.
Click the calendar icon to set a date. To change it, click the existing date and select a new one. Remove the date by clicking the cross icon. Sortable.
Participants
Users subscribed to case updates.
To receive email alerts for case updates, click the Subscribe icon. To unsubscribe, hover over your avatar and click the button.
Lab directors and other authorized roles can assign cases directly to analysts, making workload management easier.
User groups
User groups defined in Settings; each group appears as its own column.
Bring Your Own Key (BYOK) is a security feature that allows organizations to use their own encryption keys to protect their data. This ensures that they maintain control over their encryption keys and, consequently, their data.
Illumina integrates with leading Key Management Services (KMS), including Azure Key Vault and AWS KMS, so organizations can maintain full control over their encryption keys. These integrations combine Illumina’s Bring Your Own Key (BYOK) feature with your preferred KMS provider to deliver robust key management and enhanced data security.
Azure Key Vault is a cloud service that provides a secure way to store and manage sensitive information like API keys, passwords, and certificates. It offers robust features for key management, including key generation, storage, and lifecycle management.
AWS Key Management Service (KMS) allows you to create and control encryption keys used to encrypt your data across a wide range of AWS services and applications. It provides centralized management of encryption keys and integrates seamlessly with other AWS services.
Losing the encryption key means that all data encrypted with that key will be inaccessible. This can lead to permanent loss of access to crucial information.
It is crucial to securely store and manage your keys to prevent such risks.
The API server encrypts the organization's information before storing it in the database and decrypts it when needed (e.g., during pipeline execution). The key vault is managed by the organization.
To configure encryption in Emedgene, you need the following information from Azure Key Vault:
Application tokens:
Client Id
Tenant Id
Client Secret
The key information:
Key URL
Navigate to App registrations
Click Register to create a new application and and fill in the required details
After registration, copy and save the Application (Client) ID and Directory (Tenant) ID
In the left menu, select Certificates & Secrets
Click New client secret. Copy and save the Value (Client Secret) immediately, as it is shown only once.
Click New Key (Create key vault)
Specify the key vault name, region (for example, East US), and pricing tier
Click Next to go to Access Policies
Navigate to the newly created Key vault
In the left menu, select Keys, and then select the key
Select the current version
Description is coming soon.
The API server will encrypt the client's information before storing it in a database and decrypt that information when needed (e.g., running the pipeline). The key vault is managed by the client, and Emedgene will only be provided with access to encrypt/decrypt functions in that key vault. This guarantees that clients control access to the information.
Illustration of data flow when creating a case in Emedgene platform:
Illustration of data flow when reading a case data from emedgene platform:
A preliminary step to this solution is having a key vault owned by the client, and a key that Emedgene is given access to.
The client will create an access policy in the key vault of type “Application” and provide the matching key and secret to Emedgene. The access policy must contain permissions to perform encrypt and decrypt actions.
In order for Emedgene to integrate with the key, depending on the key vault provider, the client needs to provide the following information:
Client Id
Client Secret
Tenant Id
Key vault name
Since some of our platform search capabilities run directly on the DB, we can’t directly search any data that is encrypted. To overcome this, we will implement a hashing search functionality as follows.
The case data will still be fully encrypted in the DB as it is today
Specific fields we want to make “searchable” - as defined by the customer, we will save their hash value alongside the encrypted data.
Hashing will be done using SHA-256, and will include a secure random generated salt of 32 characters, which will be added to the value.
The salt is unique and will not be used anywhere else in the platform.
Illustration of data flow when searching in Emedgene platform:
Illustration of data flow when creating a case with searchable field in Emedgene platform:
Use the case progress indicator to determine where a case is in the analysis workflow and whether you need to take action.
The indicator appears in the Status column on the Cases page for cases with In progress, Re-Analysis, or Issue reported status.
This section explains what each progress indicator state means for cases with In progress or Re-Analysis status.
Indicator: Grey striped progress bar.
Meaning: The FASTQ‑based case is running DRAGEN secondary analysis.
Indicator: Empty segmented progress bar.
Meaning: The system has accepted the VCF (either a VCF case or the VCF output from integrated DRAGEN secondary analysis), but processing has not started yet.
Blue segments fill from left to right as the case moves through the five stages of the Emedgene tertiary analysis pipeline.
A filled blue segment indicates a completed stage
A half-filled blue segment indicates the stage in progress
Hover over the indicator to view:
The current tertiary analysis stage
Elapsed time since tertiary analysis started
Indicator: Segment 1 is half-filled.
Meaning: The system standardizes and prepares the uploaded data for downstream processing.
Indicator: Segment 1 is filled and segment 2 is half-filled.
Meaning: Variants are enriched with genomic, clinical, population, transcript, phenotype, and knowledge base annotations.
Indicator: Segments 1-2 are filled and segment 3 is half-filled.
Meaning: The annotated data is indexed to prepare the case for review within the platform.
Indicator: Segments 1-3 are filled and segment 4 is half-filled.
Meaning: The system performs quality validation.
Indicator: Segments 1-4 are filled and segment 5 is half-filled.
Meaning:
The system performs AI-driven phenotype matching.
The AI Shortlist performs variant prioritization, resulting in the list of for the case.
After all five tertiary analysis stages complete, the case status changes to Delivered.
In case of processing failure, Issue reported status may be supplemented with the progress indicator.
Indicator: Red striped progress bar.
Meaning: Failure during DRAGEN secondary analysis.
Indicator: No progress indicator is shown.
Meaning: Input VCF file issues.
Indicator: Progress bar shows completed tertiary analysis stages in blue and the failed stage in red.
Meaning: Processing stopped during the stage corresponding to the segment filled in red.
Hover over the indicator to confirm the failed stage. Select the question-mark icon () to view error details.
In Emedgene, case status indicates the current stage of a case: from data upload through analysis, review, and results finalization.
Statuses are assigned either automatically by the system or by authorized users, depending on the workflow stage and user permissions. Different statuses require different for assignment and reassignment.
System-controlled: Assigned automatically by the system; cannot be reassigned by users.
User-controlled:
Reanalysis will fail
FASTQ
VCFs
Reanalysis will fail
FASTQ
CSV, etc
Reanalysis will fail
VCF
BAM/CRAM (visualizations)
Visualization will fail
VCF
VCF (input)
Reanalysis will fail
VCF
CSV, etc
Reanalysis will fail
(will be fixed)
Moderate
Passed or N/A
30−45×
70−80%
High
Passed or N/A
>45×
>80%
25−45%
High
Passed or N/A
<25%
{
"test_data":
{
"consanguinity": false,
"inheritance_modes":
[],
"sequence_info":
{},
"type": "Whole Genome",
"notes": "",
"samples":
[
{
"bam_location": "",
"fastq": "NA12878-PCRF450-1",
"status": "uploaded",
"directoryPath": "",
"sampleFiles":
[
{
"filename": "NA12878-PCRF450-1.metrics.tar.gz",
"sample_type": "dragen-metrics",
"path": "/analysis_output/demo_data_germline_v4_3_6_v2-DRAGEN_Germline_Whole_Genome_4-3-6-v2-75b081e8-a8aa-433e-862b-a20d2d65e492/NA12878-PCRF450-1/NA12878-PCRF450-1.metrics.tar.gz",
"size": 0,
"storage_id": 420,
"status": "uploaded",
"vcf_column_name": "NA12878-PCRF450-1",
"vcf_column_names":
[
"NA12878-PCRF450-1"
],
"loadingSample": false
},
{
"filename": "NA12878-PCRF450-1.hard-filtered.vcf.gz",
"sample_type": "vcf",
"path": "/analysis_output/demo_data_germline_v4_3_6_v2-DRAGEN_Germline_Whole_Genome_4-3-6-v2-75b081e8-a8aa-433e-862b-a20d2d65e492/NA12878-PCRF450-1/NA12878-PCRF450-1.hard-filtered.vcf.gz",
"size": 0,
"storage_id": 420,
"status": "uploaded",
"vcf_column_name": "NA12878-PCRF450-1",
"vcf_column_names":
[
"NA12878-PCRF450-1"
],
"loadingSample": false
}
],
"storage_id": 420,
"sampleType": "vcf"
}
],
"sample_type": "vcf",
"patients":
{
"proband":
{
"fastq_sample": "NA12878-PCRF450-1",
"gender": "Male",
"healthy": false,
"relationship": "Test Subject",
"notes": "",
"phenotypes":
[
{
"id": "phenotypes/EMG_PHENOTYPE_0001324",
"name": "Muscle weakness"
}
],
"detailed_ethnicity":
{
"maternal":
[],
"paternal":
[]
},
"zygosity": "",
"quality": "",
"dead": false,
"ignore": false,
"id": "proband"
},
"other":
[]
},
"diseases":
[],
"disease_penetrance": 100,
"disease_severity": "",
"boostGenes": false,
"selected_preset_set": "",
"incidental_findings": null,
"labels":
[],
"gene_list":
{
"type": "all",
"id": 1,
"visible": false
}
},
"should_upload": false,
"sharing_level": 0
}You may combine multiple gene lists into one, or add specific genes to an existing list during case creation. The merged list behaves like any other list in the platform.
Always ensure sample IDs are unique to prevent case failure.
If using joint gVCF input, place the proband first for accurate insufficient region calculation.
The UI does not allow reusing the same gVCF file for multiple samples.









Select Add access policy, and set Key permissions:
Key Management Operations
Cryptographic Operations: Decrypt, Encrypt, Unwrap Key, Wrap Key
Set Secret permissions:
Secret Permission: Get
Select Principal: select the application you created earlier
Finish with Review + create
Copy the Key Identifier (Key URL):
When the user enters a string to search, we will hash that value using all the salt values, and search those hash values.
https://<key-vault-name>.vault.azure.net/keys/<key-name>/<key-version>Client->Emedgene API: Add New Test Request
note right of Emedgene API: Process Request
Emedgene API->Key Vault: PHI
note right of Key Vault: Encrypt
Key Vault->Emedgene API: Encrypted PHI
Emedgene API->Emedgene DB: Store Encrypted PHIClient->Emedgene API: Get Test Request
emedgene DB->Emedgene API: Encrypted PHI
Emedgene API->Key Vault: Encrypted PHI
note right of Key Vault: Decrypt
Key Vault->Emedgene API: Decrypted PHI
Emedgene API->Client: Decrypted PHIClient->Emedgene API: Add New Test Request
note right of Emedgene API: Process Request
Emedgene API->Key Vault: PHI
note right of Key Vault: Encrypt
Key Vault->Emedgene API: Encrypted PHI
Emedgene API-> Emedgene DB: Get Salt
Emedgene API-> Emedgene API: Hash Value using Salt
Emedgene API->Emedgene DB: Store Encrypted PHI + Hashed valueClient->Emedgene API: Search string
Emedgene API->AWS Secrets: Get Salt
Emedgene API-> Emedgene API: Hash string using Salt
Emedgene API->Emedgene DB: Search hashed string
Emedgene DB->Emedgene API: Search results
Emedgene API->Client: Search resultsCase status
Overview of case status behavior, history, and related workflows.
Case statuses in a case lifecycle
Review status transitions, control types, and the reference status list.
How to update a case status
Change the current case status from the case page or the Cases table.
Case status management
Create custom statuses and reorder them for your organization.











System-assigned, user-reassignable: Assigned automatically by the system but can be reassigned by authorized users.
Out-of-the-box: Default options provided by the platform.
Custom: User-configured to align with specific workflows.
Each status represents a distinct stage in the case lifecycle. Figure 1 shows the possible transitions between statuses and the control type for each assignment, indicated by solid and dashed arrows. Table 1 provides an overview of case statuses.
Table 1. Case status reference
"Uploading"
Data upload in progress.
System-controlled
Out-of-the-box
"In progress"
Analysis currently running.
System-controlled
Out-of-the-box
"Delivered"
Analysis completed; case ready for review.
System-assigned, user-reassignable
Out-of-the-box
Custom status
Indicates a custom case processing stage between "Delivered" and "Finalized".
User-controlled
Custom
"Finalized"
The analysis and review of the case by the analyst group have been completed.
Typically assignment and reassignment of the "Finalized" status is configured to be restricted to organization managers and/or lab directors.
User-controlled
Out-of-the-box
"Trash bin"
Case marked for deletion; access restricted.
Typically assignment and reassignment of the "Trash bin" status is configured to be restricted to organization managers and/or lab directors.
User-controlled
Out-of-the-box
"Pending sequencing"
Case created; awaiting sequencing data.
System-controlled.
Exception: user-reassignable to "Trash bin"
Out-of-the-box
"Issue reported"
The case failed to run.
Please check the integrity of the uploaded files and ensure that the variant caller used is on Emedgene list of accepted .
System-controlled.
Exception: user-reassignable to "Trash bin"
Out-of-the-box
"Re-Analysis"
The system is re-running the AI Shortlist algorithm.
System-controlled
Out-of-the-box
Case status
Overview of case status behavior, history, and related workflows.
Case progress indicator (v100.40.0+)
Understand progress indicator states for cases in processing or failure states.
How to update a case status
Change the current case status from the case page or the Cases table.
Case status management
Create custom statuses and reorder them for your organization.
Emedgene provides the tightest integration with DRAGEN for germline variation analysis, providing accuracy, comprehensiveness, and efficiency, spanning variant calling through interpretation and report generation.
4.5
100.40.0+
*to modify the case pipeline version refer to .
The Emedgene platform supports a variety of variant callers and applies specific quality parameters for each. The quality assessment is an essential step in the Emedgene pipeline because variants with low quality will not be considered by the AI components.
If the variant caller is not supported or not recognized, a default quality function will be applied. The default parameters are built on GT (genotype), depth (DP) and allele bias (AB). These fields are mandatory, and their absence will induce “Low quality” for all variants.
The following variant callers are currently supported on the Emedgene pipeline, providing a header with the variant caller command line should be present within the VCF headers.
Additional callers can be supported on demand under license.
The following are the general format requirements for a CSV file used to create multiple cases:
The file must have a .csv extension.
The file must contain a [Data] header.
The row after [Data] header must include the field names identifying the data in each column. The column names are case-sensitive.
The row after the column name header and each subsequent row represents a sample.
Each column represents a data field.
It is essential that there are no empty rows between the [Data] header and the last sample row.
Number of cases per file can't be greater than 50.
Must be present in the sample table at all times.
Case Type;
Family Id;
Phenotypes OR Phenotypes Id.
If these fields are left empty, it will result in the creation of an empty sample.
BioSample Name;
Files Names;
Storage Provider Id;
This field is mandatory if Files Names is empty:
Sample Type.
This field is required if the auto option is used for Files Names (only relevant for BSSH):
Default Project.
The sample table may include these supported optional columns.
Assignee ID (v100.40.0+)
Boost Genes
Clinical Notes
The sample table may contain custom columns to suit your specific needs and include any relevant information that is important for your workflow.
Each custom field must be assigned a unique name without spaces. Data from custom columns is saved per case under the Additional information section of .
(highlighted in red), (highlighted in orange), and fields should be filled in according to the following rules.
For BSSH, it is necessary to use the actual names (numbers):
instead of aliases
The batch upload process allows you to provide a human-readable path in your batch CSV for BSSH files.
When a batch CSV includes a human-readable path, the system performs the following validations for paths in BSSH storage:
Single File in the Path:
If the provided path contains exactly one file or dataset, the batch upload proceeds successfully.
Two Files in the Path:
If the path contains two files with the same name (for example, two pairs of fastqs in a dataset) , the system will:
Multiple QCPassed Datasets: If two datasets in the same path are marked as QCPassed, the batch upload will fail with a descriptive error indicating the conflict.
Excessive Files in the Path: If more than two files are found for the provided path, the batch upload will fail, instructing the user to provide a more specific or valid path.
Enables customers to use intuitive, human-readable paths in their workflows.
Automatically handles dataset selection based on quality control status.
1.2
37.0+
Cyto
CNVReadDepth
5.12, 5.20
SmallVariant
N/A
SmallVariant
1.38
CNVReadDepth
Clair3
v37.0+
SmallVariant
N/A
SmallVariant
ClinSV
N/A
SVSplitEnd
N/A
CNVReadDepth
CNVReporter
0.01
CNVReadDepth
1.0
CNVReadDepth
N/A
CNVReadDepth
cuteSV
2.02
v37.0+
SVSplitEnd
Multi-Sample Viewer:1.0.0.71
Unknown
1.0.0
SmallVariant
N/A
SVSplitEnd
0.1
CNVReadDepth
ExomeDepthAM
0.1
Private fork of ExomeDepth
CNVReadDepth
N/A
SmallVariant
3, 3.4, 3.5, 2014, 4, 4.1
SmallVariant
GATK
N/A
SmallVariant
Scramble
Running: scramble2vcf.pl
SmallVariant
1.4
SmallVariant
4.x, 5.x and not: 5.12, 5.20
SmallVariant
CNV
5.16
CNVReadDepth
2.2.0
SVSplitEnd
N/A
SmallVariant
2.X
SmallVariant
2.1.1
SVSplitEnd
2.2.4
SmallVariant
2.2.4
SVSplitEnd
2.X
SVSplitEnd
5.2.9
SmallVariant
201808, 201911, 202010
SmallVariant
201808.03
SmallVariant
2.0.6, 2.0.7, 2.5
SVSplitEnd
0.0.2
SmallVariant
2.0.1
CNVReadDepth
Spectre
v37.0+
CNVReadDepth
2.4.5
SmallVariant
N/A
SmallVariant
N/A
SVSplitEnd
SNV, CNV, STR, SV (del/dup/ins), Targeted, MRJD, Ploidy TruPath: SNV, SV, MRJD
100.39.0+
SNV, CNV, STR, SV (del/dup/ins), Targeted, MRJD, Ploidy
36.0+
SNV, CNV, STR, SV (del/dup/ins), Targeted, MRJD, JSON PGx*
All
SNV, CNV, STR, SV (del/dup/ins), SMN, JSON PGx*
4.2
All
SNV, CNV, STR, SV (del/dup/ins), SMN
4.0
All
SNV, CNV, STR, SV (del/dup/ins)
3.10
All
SNV, CNV, STR, SV (del/dup/ins)
3.6-3.9
All
SNV
1.4
100.40.0+
Cyto
1.3
100.39.0+
AED CNV
N/A
Cyto
Affymetrix Extensible Data. converted to VCF
Date Of BirthDue Date
Execute now
Gender. See an important note
Gene List Id
Intersect Bed Id
Kit Id
Label Id
Opt In
Relation
Selected Preset. See an important note
Visualization Files
Custom
Free text
24-02-2022
Sample_Type
Custom
Free text
Amniotic Fluid
BioSample Name
Conditionally mandatory. An empty sample will be created if the field is left blank.
Free text
NA24385
Boost Genes
Optional.
Indicates whether the Boost genes mode will be used. TRUE means that variants in the targeted genes will receive upgraded scores during prioritization by the AI Shortlist algorithm.
Default value is FALSE.
Only considered for proband.
1. TRUE
2. FALSE
TRUE
Case Type
Mandatory. Only considered for proband.
1. Whole Genome
2. Exome
3. Custom Panel
4. Array
5. Custom case type
Whole Genome
Clinical Notes
Optional
Free text
A 14-year-old boy with a visual acuity of 20/200 in both eyes in whom hearing loss was first noted at 5 years of age on routine screening; audiometry revealed sensorineural hearing loss.
Date Of Birth
Optional
YYYY-MM-DD
2013-01-22
Default Project
Conditionally mandatory.
Must be filled in if the auto option is used for Files Names (only relevant for BSSH).
Free text
GIAB
Due Date
Optional
YYYY-MM-DD
2023-05-03
Execute now
Optional.
Default value is TRUE. Use FALSE if you don't want to run the case upon uploading the file.
Only considered for proband.
1. TRUE
2. FALSE
FALSE
Family Id
Mandatory
Free text
RM8392
Files Names
Conditionally mandatory.
An empty sample will be created if the field is left blank.
The existing option automatically locates FASTQ files based on the BioSample Name.
Note: If data files for an existing case were sourced from the customer's external bucket and later removed, attempting to create a case from those files will result in an error.
Learn about the current limitation for CRAM file input.
With the auto option, BSSH users can automatically locate FASTQ files based on the BioSample Name and Default Project provided.
When using BSSH without the auto option, ensure that your file path is .
1. Semicolon-separated list of paths to .fastq, .fastq.gz, .vcf, .vcf.gz, .bam, .cram, .gt_sample_summary.json, .annotated_cyto.json files without spaces
2. existing
3. auto (BSSH)
/GIAB_cases/1/NA24385.dragen.hard-filtered.gvcf.gz;/QA_cases/Other/NA24385.dragen.cnv.vcf.gz;/QA_cases/Other/NA24385.dragen.repeats.vcf;
Gender
Optional.
Default value is U. See an .
1. F
2. M
3. U
M
Gene List Id
Optional. Must be the id of a previously defined Gene List. Only considered for proband.
Integer
12345
Kit Id
Optional.
ID of a Coverage BED. Must be the id of a previously defined kit. Only considered for proband.
Integer
23456
Intersect Bed Id
Optional. ID of a Region of interest BED. Must be the id of a previously defined kit. Only considered for proband.
Integer
78957
Label Id
Optional. Must be the id of a previously defined Case Label. Only considered for proband.
Integer
34567
Opt In
Optional.
Indicates whether the case subject consented to the extended sharing of data with your network(s).
Default value is TRUE.
1. TRUE
2. FALSE
FALSE
Phenotypes
Mandatory for proband sample if Phenotypes Id is empty.
List must be under 100.
It is possible to include non-HPO terms if Phenotypes Id is empty.
Semicolon-separated list of HPO phenotype terms
Unaffected is used for non-affected family members.
Abnormal pupillary function;Orthotopic os odontoideum;
Phenotypes Id
Mandatory for proband sample if Phenotypes is empty.
List must be under 100.
Semicolon-separated list of HPO phenotype IDs
HP:0007686;HP:0025375;
Relation
Optional.
Default value is proband.
Values proband, father, mother can be only used once per Family ID.
One sample with Relation proband is required per Family ID.
1. proband
2. mother
3. father
4. sibling
mother
Sample Type
Conditionally mandatory.
Required if Files Names is empty.
Only considered for proband.
1. FASTQ
2. VCF
FASTQ
Selected Preset
Optional. Must be the name of a previously defined preset group. The specified preset group appears in the Presets tab for the case.
If set to Default, the default preset group is used.
If left empty, no preset is applied.
See an .
1. Free text
2. Default
Exome trio
Storage Provider Id
Conditionally mandatory.
Required if Files Names is not empty.
Must be from the configured storage provider ID list.
Integer
208
Visualization Files
Optional
Semicolon-separated list of paths to sequence alignment data files of extension .bam, .cram, .tn.bw, .baf.bw, .roh.bed, .lrr.bedgraph, .baf.bedgraph
/giab_project/NA24385.bam
Select the dataset marked as QCPassed.
Fail the batch upload if both datasets are marked as QCPassed, as this indicates conflicting data.
More Than Two Files in the Path:
If the path contains more than two files or datasets, the system fails the batch upload, as the path is considered ambiguous or invalid.
Institution
Custom
Free text
GenoMed Solutions
Assignee ID (v100.40.0+)
Optional. Users subscribed to case updates.
Appears as Participants in the Case info tab and the Cases table.
Comma-separated list of user IDs
/projects/3824821/appresults/2319318/files/119675608/projects/ABC_DEF_2022-12-22_DEv395/appresults/ABC-GM58342-def/files/ABC-GM58342-def.hard-filtered.vcf.gzWhen a sample is user-assigned "Unknown" sex, the system assumes "Female". This affects CNV interpretation on sex chromosomes in case the genetic sex is actually male:
Chromosome X: CN = 2 is considered reference (REF) for a female genome, so CNVs with two copies are hidden by default. This may cause chromosome X duplications to be missed.
Chromosome Y: CN = 0 is considered reference (REF) for a female genome, so CNVs with zero copies are hidden by default. This may cause chromosome Y deletions to be missed.
To include these variants in the analysis, enable the Include Reference Homozygosity and No Coverage Calls toggle in Workbench & Pipeline Settings.
Despite its name, the Selected Preset field specifies a preset group used in the case, not an individual preset.
Sample_Received_Date
3,10,14
Short tandem repeats (STRs) are genomic regions composed of repeated short DNA sequences. When STRs expand beyond the normal range, they can cause mutations known as repeat expansions, which may alter gene expression or function.
Expansion of these sequences in certain genes is the underlying mechanism for a group of inherited conditions called repeat expansion disorders. These disorders primarily affect the nervous and muscular systems, leading to diseases such as Fragile X syndrome, amyotrophic lateral sclerosis and Huntington’s disease.
The platform shows repeat counts per allele and highlights values that fall in normal, intermediate or pathogenic ranges (where known).
STR variants outside the genomic regions specified in the DRAGEN v4.2 expanded variant catalog are excluded before tertiary analysis. This applies regardless of whether DRAGEN (internal or external) or another pipeline was used for the secondary analysis and irrespective of the DRAGEN version.
Variant calling reliability and availability of meaningful annotations for STR loci can vary. To reduce noise and ensure high-confidence calls,
Validated genotype–phenotype associations and established pathogenicity. Loci must have well-documented links to disease and specific phenotypes. This avoids reporting variants with unclear implications.
Reliable calling performance. STR calling is technically challenging. Some loci are prone to false positives or inaccurate sizing. The subset includes loci that DRAGEN-STR can call consistently and accurately. This avoids reporting low-quality variants.
Expansion thresholds. Loci with defined pathogenic thresholds (e.g., number of repeats linked to disease) are prioritized. This avoids reporting variants that lack clear interpretation guidelines.
STRs that are not tagged by AI Shortlist are still included in the analysis. This means they will not be automatically prioritized or tagged as Most Likely or Candidate by the AI. However, these loci remain available for manual review and tagging.
Table 1. STR loci considered by AI Shortlist
AR
X:67545316-67545385
Spinal and bulbar muscular atrophy (SBMA)
XR
GCA
ATN1
12:6936716-6936773
Dentatorubral-pallidoluysian atrophy (DRPLA)
AD
CAG
ATN1
12:6923960-6923977
TG
ATN1
12:6925117-6925140
GT
ATN1
12:6930800-6930827
AATA
ATXN1
6:16327633-16327723
Spinocerebellar ataxia 1 (SCA1)
AD
TGC
ATXN10
22:45795354-45795424
Spinocerebellar ataxia 10 (SCA10)
AD
ATTCT
ATXN2
12:111598949-111599018
Spinocerebellar ataxia 2 (SCA2)
AD
GCT
ATXN2
12:111461472-111461481
AAAAT
ATXN2
12:111464247-111464282
GT
ATXN2
12:111483194-111483237
AC
ATXN2
12:111487453-111487472
TTTA
ATXN2
12:111568991-111569015
TTTTG
ATXN2
12:111577256-111577279
TTAT
ATXN2
12:111588188-111588219
AAAT
ATXN2
12:111591364-111591375
TAAAAA
ATXN2
12:111591756-111591773
TTG
ATXN2
12:111596205-111596244
AC
ATXN2
12:111596373-111596388
CA
ATXN2
12:111600935-111600958
AAAT
ATXN2
12:111604087-111604106
AT
ATXN3
14:92071009-92071041
Spinocerebellar ataxia 3 (SCA3)
AD
GCT
ATXN3
14:92059829-92059848
AAAG
ATXN3
14:92065889-92065900
AAT
ATXN3
14:92078361-92078376
TTAT
ATXN7
3:63912684-63912714
Spinocerebellar ataxia 7 (SCA7)
AD
GCA
ATXN7
3:63912714-63912725
GCC
ATXN8OS
13:70139383-70139428
CTG
ATXN8OS
13:70139353-70139382
Spinocerebellar ataxia 8 (SCA8)
AD
CTA
ATXN8OS
13:70115027-70115046
GT
ATXN8OS
13:70126607-70126620
TA
ATXN8OS
13:70129375-70129402
TG
C9ORF72
9:27573528-27573546
Frontotemporal dementia and/or amyotrophic lateral sclerosis 1 (FTDALS1)
AD
GGCCCC
CACNA1A
19:13207858-13207897
Spinocerebellar ataxia 6 (SCA6)
AD
CTG
CNBP
3:129172576-129172656
CAGA
DMPK
19:45770204-45770264
Myotonic dystrophy 1 (DM1)
AD
CAG
FMR1
X:147912050-147912110
Fragile X syndrome (FXS)
XD
CGG
FXN
9:69037286-69037304
GAA
FXN
9:69037261-69037285
Friedreich ataxia (FRDA)
AR
A
HTT
4:3074876-3074933
Huntington disease (HD)
AD
CAG
HTT
4:3074939-3074965
CCG
JPH3
16:87604287-87604329
Huntington disease-like 2 (HDL2)
AD
CTG
NOP56
20:2652733-2652757
Spinocerebellar ataxia 36 (SCA36)
AD
GGCCTG
NOP56
20:2652757-2652774
CGCCTG
PPP2R2B
5:146878727-146878757
Spinocerebellar ataxia 12 (SCA12)
AD
GCT
TBP
6:170561906-170562017
Spinocerebellar ataxia 17 (SCA17)
AD
GCA