One consistent, open source of truth — usable far beyond the XML engineers who can parse it today.
A spreadsheet or VCF of variants you care about — your panel, your cohort.
Join to ClinVar-GKM by computable variant id — no manual normalization.
Classification, condition, review status & provenance — ready to read.
Most people would never download ClinVar and build a report. With typed Parquet (today) — and friendly no-code tools the standard makes possible — they won't have to.
Filter by gene, significance, and review quality, then follow the submitter-provided citations to the PubMed evidence — the kind of discovery that's impractical against raw XML.
Reference Python libs parse records into validated models (vrs-python · cat-vrs-python · va-spec-python).
A weekly delta ships for every ClinVar release — added, updated & deleted records with a manifest — typically within a day or two. Stay synced by applying deltas, not re-downloading and re-parsing the world.
ClinVar's own VCF sometimes carries too much or too little for a given use. Regenerate full or partial VCFs from the standardized data — scoped to your gene panel, quality bar, or significance.
Because each release is captured consistently, you can study how classifications evolve — reclassification rates, emerging consensus, submitter trends — with friendly tools built on one stable shape.
A semantically sound, uniform representation — the same meaning, every record — is far better AI input than inconsistent XML. Stable identifiers and explicit propositions mean models and agents can trust what they read.
Aggregate and individual case evidence that sits deep in submissions — rich signal that almost no one extracts today.
Functional evidence submitted alongside classifications — hard to access and reuse in its current form.
Some of ClinVar's most valuable content is effectively locked by the effort it takes to extract. Standardizing it in GKM puts it within reach of everyone — not just teams who can invest in bespoke parsing.
Where variants and classifications already live in your workflow.
Validate, compare against current ClinVar, and prepare submissions — on a standard model.
Integrate submission & status tracking without bespoke glue code.
A shared, standard representation makes it realistic to build an enhanced submission reviewer and submit/track tooling that fits existing systems — lowering the barrier to contributing high-quality data back.
Query your own data and the literature — no parser, no normalization.
Delete the XML pipeline; sync with weekly deltas and validated models.
Consistent, semantic inputs for apps, agents, and knowledge graphs.
Tailored VCFs and trustworthy, quality-ranked classifications.
Review and submit in your workflow; see your footprint clearly.
Shared tooling, temporal insight, and data that was out of reach.
Start here: clingen-data-model.github.io/clinvar-gkm · Downloads · GKM Starter Kit · Roadmap — upvote what you need
| → Space | Next slide |
| ← | Previous slide |
| Home / End | First / last slide |
| S or N | Toggle speaker notes |
| T | Toggle light / dark theme |
| F | Toggle fullscreen |
| ? | Show / hide this help |