A ClinVar-GKM opportunity brief · ClinGen

What ClinVar-GKM unlocks

From a weekly XML archive to a platform — for analysis, developer tooling, discovery, AI, and giving everyone access to data that was hard to reach.
GA4GH Connect · Oct 2, 2026 Visual brief Companion to the Primer
What's realistic today vs. tomorrow: Available now Now within reach On the roadmap
Use ← / → to navigate · press ? for shortcuts
The shift

From parsing pipelines to a queryable platform

Today, to use ClinVar fully

Download the XML release
Write & maintain parsers
Normalize variant identity yourself
Track every schema change
…repeat, every release
→

With ClinVar-GKM

One standardized bundle (+ typed Parquet)
Stable VRS identity, built in
Query, build, and analyze directly
Weekly deltas keep you current
Shared tooling, not bespoke code
Democratize Available now

Free the data

One consistent, open source of truth — usable far beyond the XML engineers who can parse it today.

ClinVar-GKM
one standardized, CC0 bundle
Researchers
cohort & gene-level analysis
Clinicians
consistent classifications
Developers
drop-in, no XML parser
Data analysts
query Parquet directly
Submitters
see their footprint clearly
AI / tools
clean, semantic inputs
Analysis for everyone Parquet today No-code: within reach

Analyze your own data — without a parser

Bring

Your variant list

A spreadsheet or VCF of variants you care about — your panel, your cohort.

→
Match

On VRS identity

Join to ClinVar-GKM by computable variant id — no manual normalization.

→
Get

Answers

Classification, condition, review status & provenance — ready to read.

Most people would never download ClinVar and build a report. With typed Parquet (today) — and friendly no-code tools the standard makes possible — they won't have to.

Discovery Available now

Ask better questions

Example: pathogenic BRCA1 variants from trusted submitters — and the literature behind them
gene = BRCA1 classification = Pathogenic / Likely pathogenic review status ≥ 1★ (criteria provided) has citations
▼
1 query
typed Parquet join — no XML parsing
variants ↓
filtered to ≥1★ quality submitters
PubMed →
citations linked from each submission

Filter by gene, significance, and review quality, then follow the submitter-provided citations to the PubMed evidence — the kind of discovery that's impractical against raw XML.

For developers Available now

Delete your XML pipeline

Maintain today

XML fetch & unpack
Bespoke parser (deeply nested)
Variant-identity normalization
Patch for every schema change
Load & reconcile
→

With ClinVar-GKM

Fetch bundle or weekly delta
Load typed Parquet / JSON
— that's it —

Reference Python libs parse records into validated models (vrs-python · cat-vrs-python · va-spec-python).

Stability & maturity Available now

A schema that won't surprise you

ClinVar XML  — shape can change when you least expect it
Each ● is an XML structure change your parser must absorb — on ClinVar's schedule, not yours.
GKM  — versioned, community-governed, maturing fast
1.0 1.1 1.x
Changes arrive as semantic versions through the GA4GH standards process — predictable, documented, and being hardened now as production adopters normalize the model.
Realtime Available now

Always current — no re-parsing

FULL
monthly
Δ
wk 1
Δ
wk 2
Δ
wk 3
Δ
wk 4
FULL
monthly
Δ
wk 1
Δ
wk 2
replay deltas onto the latest full → always up to date →

A weekly delta ships for every ClinVar release — added, updated & deleted records with a manifest — typically within a day or two. Stay synced by applying deltas, not re-downloading and re-parsing the world.

Tailored outputs Now within reach

Rebuild VCFs — your way

ClinVar-GKM
standardized bundle
→
filter: gene in panel
review status ≥ 2★
significance P / LP
→
Your VCF
full or partial — exactly what you need

ClinVar's own VCF sometimes carries too much or too little for a given use. Regenerate full or partial VCFs from the standardized data — scoped to your gene panel, quality bar, or significance.

Temporal analysis On the roadmap

See how knowledge changes over time

Illustrative — variant reclassifications per release
VUS → Pathogenic / Likely pathogenic resolutions trending up as evidence accrues
0 250 500 750 2022 2023 2024 2025 2026 rising

Because each release is captured consistently, you can study how classifications evolve — reclassification rates, emerging consensus, submitter trends — with friendly tools built on one stable shape.

Roadmap · chart is illustrative, not live data
AI-ready Now within reach

Reliable ground truth for AI

Explicit, consistent semantics
variant — isCausalFor → condition
classification: Pathogenic
quality: reviewed by expert panel
id: ga4gh:VA.…
→
Dependable inputs for
RAG & grounding for LLMs
Agents & decision-support tools
Knowledge graphs
Model training & evaluation

A semantically sound, uniform representation — the same meaning, every record — is far better AI input than inconsistent XML. Stable identifiers and explicit propositions mean models and agents can trust what they read.

Reach the unreachable On the roadmap

Unlock the data few can use today

▲

Case-level observations

Aggregate and individual case evidence that sits deep in submissions — rich signal that almost no one extracts today.

Planned for GKM
▲

Functional data (incl. MaveDB)

Functional evidence submitted alongside classifications — hard to access and reuse in its current form.

Planned for GKM

Some of ClinVar's most valuable content is effectively locked by the effort it takes to extract. Standardizing it in GKM puts it within reach of everyone — not just teams who can invest in bespoke parsing.

Close the loop Now within reach

Better submissions, in your workflow

Your systems

Apps & LIMS

Where variants and classifications already live in your workflow.

⇄
Enabled by GKM

Review & submit tooling

Validate, compare against current ClinVar, and prepare submissions — on a standard model.

⇄
ClinVar

Submit & track

Integrate submission & status tracking without bespoke glue code.

A shared, standard representation makes it realistic to build an enhanced submission reviewer and submit/track tooling that fits existing systems — lowering the barrier to contributing high-quality data back.

The opportunity

Built to be built on

Analysts & researchers

Query your own data and the literature — no parser, no normalization.

Developers

Delete the XML pipeline; sync with weekly deltas and validated models.

Tool & AI builders

Consistent, semantic inputs for apps, agents, and knowledge graphs.

Clinical teams

Tailored VCFs and trustworthy, quality-ranked classifications.

Submitters

Review and submit in your workflow; see your footprint clearly.

The community

Shared tooling, temporal insight, and data that was out of reach.

Start here: clingen-data-model.github.io/clinvar-gkm · Downloads · GKM Starter Kit · Roadmap — upvote what you need

Thank you — what would you build on this? Press ? for shortcuts.
1 / 14