For the complete documentation index, see llms.txt. This page is also available as Markdown.

OMIM

Overview

OMIM is a comprehensive, authoritative compendium of human genes and genetic phenotypes that is freely available and updated daily.

Publications

Amberger JS, Bocchini CA, Scott AF, Hamosh A. OMIM.org: leveraging knowledge across phenotype-gene relationships. Nucleic Acids Res. 2019 Jan 8;47(D1):D1038-D1043. doi:10.1093/nar/gky1151. PMID: 30445645.

Amberger JS, Bocchini CA, Schiettecatte FJM, Scott AF, Hamosh A. OMIM.org: Online Mendelian Inheritance in Man (OMIM®), an online catalog of human genes and genetic disorders. Nucleic Acids Res. 2015 Jan;43(Database issue):D789-98. PMID: 25428349.

Parse OMIM data

Illumina Connected Annotations uses gene symbols as the gene identifiers internally. To generate the OMIM database, we first map the MIM numbers, which are the primary identifiers used by OMIM, to gene symbols supported by Illumina Connected Annotations. Please note that there can be multiple MIM numbers mapped to one gene symbol. Only MIM numbers successfully mapped to an Illumina Connected Annotations gene symbol are further processed. The OMIM API is used to fetch all the information associated with a gene MIM number, except the gene symbols.

mim2gene.txt

This mim2gene.txt (http://omim.org/static/omim/data/mim2gene.txt) file provides the mapping between MIM numbers and gene symbols. An example of this file is given below:

# MIM Number    MIM Entry Type (see FAQ 1.3 at https://omim.org/help/faq)   Entrez Gene ID (NCBI)   Approved Gene Symbol (HGNC) Ensembl Gene ID (Ensembl)
100050  predominantly phenotypes
100070  phenotype   100329167
100100  phenotype
100200  predominantly phenotypes
100300  phenotype
100500  moved/removed
100600  phenotype
100640  gene    216 ALDH1A1 ENSG00000165092
100650  gene/phenotype  217 ALDH2   ENSG00000111275
100660  gene    218 ALDH3A1 ENSG00000108602
100670  gene    219 ALDH1B1 ENSG00000137124
100675  predominantly phenotypes
100678  gene    39  ACAT2   ENSG00000120437

The information in the "Entrez Gene ID (NCBI)", "Approved Gene Symbol (HGNC)" and "Ensembl Gene ID (Ensembl)" columns are used to find the proper gene symbol supported by Illumina Connected Annotations, which may or may not be the same as the gene symbol listed here.

OMIM API

Illumina Connected Annotations retrieves the OMIM annotations from the OMIM API JSON responses. The "entry" handler is used to fetch all the annotations associated with a given OMIM gene. A sample JSON response from the API is provided there.

Content from the OMIM API JSON response is reorganized as shown in the Illumina Connected Annotations JSON Output

Mappings between the Illumina Connected Annotations JSON output and OMIM JSON API are listed in the table below:

Illumina Connected Annotations JSON key chain
OMIM API JSON key chain

omim:mimNumber

omim:entryList:entry:mimNumber

omim:geneName

omim:entryList:entry:geneMap:geneName

omim:description

omim:entryList:entry:textSectionList:textSection:textSectionContent

omim:phenotypes:mimNumber

omim:entryList:entry:geneMap:phenotypeMapList:phenotypeMap:mimNumber

omim:phenotypes:phenotype

omim:entryList:entry:geneMap:phenotypeMapList:phenotypeMap:phenotype

omim:phenotypes:description

omim:entryList:entry:textSectionList:textSection:textSectionContent

omim:phenotypes:mapping

omim:entryList:entry:geneMap:phenotypeMapList:phenotypeMap:phenotypeMappingKey (see mapping below)

omim:phenotypes:inheritances

omim:entryList:entry:geneMap:phenotypeMapList:phenotypeMap:phenotypeInheritance

omim:phenotypes:comments

omim:entryList:entry:geneMap:phenotypeMapList:phenotypeMap:phenotype (see mapping below)

Mapping key to content

1 to disorder was positioned by mapping of the wild type gene 2 to disease phenotype itself was mapped 3 to molecular basis of the disorder is known 4 to disorder is a chromosome deletion or duplication syndrome

Phenotype character to comment

? to unconfirmed or possibly spurious mapping [/] to nondiseases {/} to contribute to susceptibility to multifactorial disorders or to susceptibility to infection

There are different types of link in the OMIM description section. For example, in above JSON response, we have the description of MIM entry 100640:

The ALDH1A1 gene encodes a liver cytosolic isoform of acetaldehyde dehydrogenase ({EC 1.2.1.3}), an enzyme involved in the major pathway of alcohol metabolism after alcohol dehydrogenase (ADH, see {103700}). See also liver mitochondrial ALDH2 ({100650}), variation in which has been implicated in different responses to alcohol ingestion.\n\nALDH1 is associated with a low Km for NAD, a high Km for acetaldehyde, and is strongly inactivated by disulfiram. ALDH2 is associated with a high Km for NAD, and low Km for acetaldehyde, and is insensitive to inhibition by disulfiram ({4:Hsu et al., 1985}).

As the descriptions will be shown as plain text, we remove the curry brackets surrounding links and try to make the text still readable with minimal modifications. Briefly:

  • Links referring to another MIM entry (e.g. {100650}) will be removed. Any word(s) specifically associated with the removed link will also be removed. For example, "(ADH, see {103700})" will become "(ADH)" after the process.

  • Links referring to a literature reference will be processed to remove the internal index and curry brackets. For example, "{4:Hsu et al., 1985}" becomes "Hsu et al., 1985".

  • All the other links will simple have their curry brackets removed. For example, "{EC 1.2.1.3}" becomes "EC 1.2.1.3".

  • If the content within a pair of parentheses becomes empty after being processed, the parentheses need to be removed as well and its surrounding white spaces should be properly processed. For example, "ALDH2 ({100650})," will become "ALDH2,".

Here is a list of examples about how the description section supposed to be processed:

Original text
Processed text

({516030}, {516040}, and {516050})

(e.g., D1, {168461}; D2, {123833}; D3, {123834})

(e.g., D1; D2; D3)

(desmocollins; see DSC2, {125645})

(desmocollins; see DSC2)

(e.g., see {102700}, {300755})

(ADH, see {103700}). See also liver mitochondrial ALDH2 ({100650})

(ADH). See also liver mitochondrial ALDH2

(see, e.g., CACNA1A; {601011})

(see, e.g., CACNA1A)

(e.g., GSTA1; {138359}), mu (e.g., {138350})

(e.g., GSTA1), mu

(NFKB; see {164011})

(NFKB)

(see ISGF3G, {147574})

(see ISGF3G)

(DCK; {EC 2.7.1.74}; {125450})

(DCK; EC 2.7.1.74)

JSON output

Field
Type
Notes

mimNumber

int

OMIM ID for gene

geneName

string

gene name

description

string

phenotypes

object array

see Phenotype entry below

Phenotype

Field
Type
Notes

mimNumber

int

phenotype

string

description

string

mapping

string

see possible values below

inheritance

string array

see possible values below

comments

string array

see possible values below

Mapping

  1. disorder was positioned by mapping of the wild type gene

  2. disease phenotype itself was mapped

  3. molecular basis of the disorder is known

  4. disorder is a chromosome deletion or duplication syndrome

Inheritance

  • autosomal recessive

  • autosomal dominant

Comments

  • contributes to the susceptibility to multifactorial disorders

  • variations that lead to apparently abnormal laboratory test values

  • unconfirmed mapping

Building the supplementary files

There are 2 ways of building your own OMIM supplementary files using SAUtils.

The first way is to use SAUtils command's subcommands downloadOMIM and omim.

The second way is to use SAUtils command's subcommands AutoDownloadGenerate. To use AutoDownloadGenerate, read more in SAUtils section.

Using subcommands downloadOMIM and omim

The first step in builing the OMIM .nga files is to use the SAUtils command's subcommand downloadOMIM to download the necessary data. In order to download the data the user must possess an API key obtained from OMIM. This key has to be set as the environment variable OmimApiKey.

Once the download has succeeded, the nga files can be produced using the SAUtils command's subcommand omim.

Last updated

Was this helpful?