ClinVar-GKM¶
ClinVar-GKM provides a standardized, machine-readable representation of ClinVar release data using the GKM (Genomic Knowledge Model) schema set — VRS, Cat-VRS, and VA-Spec — from the GA4GH Genomic Knowledge Standards workstream. It is developed and maintained by the ClinGen driver project.
New to GKM datasets? Start with the GKM Starter Kit
The GA4GH GKM Starter Kit is a practical entry point for working with GKM data. It brings together the reference libraries (vrs-python, cat-vrs-python, va-spec-python), the GKM Toolkit for loading and exploring published bundles, and user stories from projects putting GKM to work. clinvar-gkm is a GKM data producer — those tools work directly on its bundles.
Why ClinVar-GKM¶
ClinVar is one of the most widely used public archives of human genetic variation and its relationship to disease. However, the native ClinVar XML format presents challenges for programmatic consumption — inconsistent structures, deeply nested records, and representations that do not align with emerging genomic data standards.
ClinVar-GKM addresses these challenges by transforming every ClinVar release into a consistent, semantically rich format built on GA4GH specifications:
- Normalized variant identifiers — Every variant receives a computable VRS identifier, enabling unambiguous cross-system matching
- Categorical variant representations — Variants are represented as Cat-VRS categorical variants with defining allele constraints and expressions
- Structured classification statements — Every submitted (SCV), aggregate (VCV), and condition-level (RCV) classification is represented as a VA-Spec statement with explicit propositions, evidence, and provenance
- Semantic consistency — Classifications, propositions, conditions, genes, and cross-references use standardized structures with typed references
What It Covers¶
The pipeline processes the entirety of each ClinVar XML release — every variation, submitted classification, and aggregate record is represented. While 100% of variant, SCV, VCV, and RCV records from the corresponding ClinVar XML release are included, some data types within ClinVar are not yet part of the v1 release.
| ClinVar Data Type | v1 Status |
|---|---|
| Variation Aggregate Classifications (VCVs) | Included |
| Variation-Condition Aggregate Classifications (RCVs) | Included |
| Germline and Somatic Submissions (SCVs) | Included |
| Variations (incl. Genes, HGVS, SPDI and VCF extensions) | Included |
| Submitters | Included |
| Aggregate & Case-level observations (example) | Planned |
| Functional data submissions (example) | Planned |
Items marked Planned are not currently included in the v1 release. If this data is requested by the community, it will be added as demand indicates.
Functional Data Submissions
Functional data is submitted to ClinVar as an SCV and may or may not be associated with a germline or somatic classification record. Functional data SCVs are excluded from the v1 release regardless of whether they are linked to a classification SCV. They will be included in a future release when functional data support is added.
Feedback and feature requests are welcome via the GitHub issue tracker.
New releases are produced within a day or two after ClinVar's XML releases and are intended to be synchronized with ClinVar's release dates.
Who It's For¶
- Variant scientists seeking a clear, consistent representation of ClinVar classifications with explicit propositions, conditions, and evidence
- Platform engineers building systems that consume ClinVar data and need a structured, well-documented format for integration
- GA4GH implementers using ClinVar-GKM as a real-world validation of VRS, Cat-VRS, and VA-Spec schemas
Getting Started¶
Download the Latest Release¶
The latest ClinVar-GKM release is available as a single compressed JSON file:
# Download the latest monthly release
curl -O https://pub-f0ad0e0dac0345408dcc95bda20beb42.r2.dev/datasets/clinvar-gkm_00-latest.json.gz
# Decompress
gunzip clinvar-gkm_00-latest.json.gz
What's in the File¶
The release file is a single JSON object with bundle sections at the root level. Each section is a keyed collection of objects — the key is the object's unique identifier, and the value is the object itself.
{
"sequenceReference": { "SQ.abc123": { ... } },
"location": { "ga4gh:SL.xyz789": { ... } },
"allele": { "ga4gh:VA.def456": { ... } },
"gene": { "ncbigene:3077": { ... } },
"variation": { "clinvar:10": { ... } },
"condition": { "clinvar.trait:9580": { ... } },
"conditionSet": { "clinvar.traitset:1234": { ... } },
"therapy": { "clinvar.therapy:{sha256}": { ... } },
"therapyGroup": { "clinvar.therapygroup:{sha256}": { ... } },
"submitter": { "clinvar.submitter:500139": { ... } },
"varcond-proposition": { "SCV001234567-PATH": { ... } },
"vartumor-proposition": { "SCV002345678-ONCO": { ... } },
"vartherapy-proposition": { "SCV003456789-TR": { ... } },
"varcustom-proposition": { "SCV004567890-RF": { ... } },
"scv": { "clinvar.submission:SCV001234567.1": { ... } },
"vcv": { "VCV000012582.63-G-PATH-CP": { ... } },
"rcv": { "RCV000012345.8-G-PATH-CP": { ... } }
}
Objects reference each other using #/ JSON pointer strings. For example, an allele references its location as "#/location/ga4gh:SL.xyz789", and an SCV statement references its proposition as "#/varcond-proposition/SCV001234567-PATH". Propositions are delivered in four datatype-specific sections (varcond-proposition, vartumor-proposition, vartherapy-proposition, varcustom-proposition) — the pointer names the section the proposition lives in.
See Output Format for the full structure and reference patterns.
Quick Example¶
To find the classification statements for a specific variant, start with the variation ID. ClinVar variation 10 (the HFE p.His63Asp variant) has the key clinvar:10 in the variation section:
{
"variation": {
"clinvar:10": {
"id": "clinvar:10",
"type": "CategoricalVariant",
"name": "NM_000410.4(HFE):c.187C>G (p.His63Asp)",
"members": ["#/allele/ga4gh:VA.ELQCnIBGqaTl0AEE0Az18XZ2cgIHAQIY"],
"constraints": [ ... ],
"extensions": [ ... ],
"mappings": [ ... ]
}
}
}
The SCV statements for this variant reference it via #/variation/clinvar:10 in their propositions. To find them, look for entries in the proposition sections (varcond-proposition, vartumor-proposition, vartherapy-proposition, varcustom-proposition) where subject (or subject, for custom) matches, then find the corresponding scv entries that reference those propositions.
Key Concepts¶
Statements are the core unit of ClinVar-GKM. Each statement represents a classification — either submitted (SCV), aggregated per variation (VCV), or aggregated per condition (RCV). Statements carry:
- A classification — the clinical significance label (e.g., Pathogenic, Likely benign)
- A proposition — what the classification asserts (variant X causes condition Y)
- Direction and strength — whether the evidence supports, disputes, or is neutral toward the proposition
- Evidence lines — links to the contributing submissions or lower-level aggregations
- Extensions — provenance metadata including submitted conditions, review status, and submission details
Propositions define the relationship being classified — a variant's causal role for a condition, its oncogenic potential, or its clinical impact. Each proposition has a type (e.g., VariantPathogenicityProposition), a predicate (e.g., isCausalFor), a subject variant, and an object condition.
Conditions represent the diseases or phenotypes that classifications are made against. Single conditions reference #/condition/clinvar.trait:{id}, while multi-condition sets reference #/conditionSet/clinvar.traitSet:{id}.
Data Access¶
ClinVar-GKM is distributed as a monthly full bundle plus weekly deltas. Each is a gzip-compressed JSON file with typed Parquet (one file per section). The files are freely available for download from Cloudflare R2 object storage with no authentication required and no egress fees.
Release Schedule¶
- Weekly deltas are published for every ClinVar release under
deltas/<yyyy-mmdd>/— each carries only the records added or updated since the prior release, plus amanifest.jsonlisting per-section adds, updates, and deletes - Monthly full bundles are published once a month under
datasets/, aligned to ClinVar's own monthly VCV releases — eachClinVarVCVRelease_YYYY-MMtriggers ourclinvar-gkm_YYYY-MMfull, built from the most recent weekly release before ClinVar's monthly cut - At the start of each year, the previous year's monthly full bundles move to
archives/
The stable filenames clinvar-gkm_00-latest.json.gz (monthly full) and clinvar-gkm-delta_00-latest.json.gz (weekly delta) always point to the most recent full and delta respectively. To reconstruct the current state, take the latest monthly full and replay the weekly deltas published since it.
Downloads¶
Download the most recent monthly full bundle:
Download the most recent weekly delta and its manifest:
curl -O https://pub-f0ad0e0dac0345408dcc95bda20beb42.r2.dev/deltas/00-latest/clinvar-gkm-delta_00-latest.json.gz
curl -O https://pub-f0ad0e0dac0345408dcc95bda20beb42.r2.dev/deltas/00-latest/manifest.json
See Downloads for the full directory layout, the manifest shape, and the consumer replay model.
Directory Structure¶
datasets/
clinvar-gkm_00-latest.json.gz latest monthly full bundle
clinvar-gkm_yyyy-mm.json.gz monthly full bundles (current year)
deltas/00-latest/
clinvar-gkm-delta_00-latest.json.gz latest weekly delta bundle
manifest.json latest delta manifest
deltas/yyyy-mmdd/
clinvar-gkm-delta_yyyy-mmdd.json.gz weekly delta bundle (added + updated records)
manifest.json per-release change manifest
archives/{yyyy}/
clinvar-gkm_yyyy-mm.json.gz monthly full bundles from prior years
Release Notes¶
Pipeline changes that affect the structure or content of the output are documented in the release_notes/ directory. These notes cover additions, bug fixes, or schema changes specific to the ClinVar-GKM pipeline — they do not replicate ClinVar's own release notes.
How It Works¶
The pipeline runs on Google BigQuery using SQL stored procedures, with an external VRS Python processing step. The two most expensive stages — variation identity and VRS processing — run incrementally, recomputing only the variations that changed since the prior release and carrying the rest forward. A per-release change log drives the weekly delta published to R2.
See the Pipeline Overview for the full workflow.
Next Steps¶
- Output Format — detailed guide to the bundle format
- Variations — how ClinVar variations are represented
- SCV Statements — submitted classification statements
- Data Model — class hierarchy and schema reference
- Pipeline Overview — how the data is produced from ClinVar XML
- Examples — annotated JSON examples
License¶
This project is licensed under CC0 1.0 Universal (public domain dedication). The output data carries the same terms as the source ClinVar data.