The full, authoritative snapshot — every variation, submission, aggregate, and provenance detail — is released weekly as XML, alongside convenient VCF / CSV reports and targeted Entrez E-utilities queries. ClinVar-GKM adds a consistent, standards-native organization of that same data.
The only representation carrying the complete database snapshot, refreshed weekly. Everything else is derived from it.
ClinVar's VCF, tab-delimited summaries, and E-utilities queries are convenient, targeted views into the full archive.
A GA4GH-standard reorganization of the same complete data — adding computable variant identity and an explicit statement model for consistent, accurate analysis.
ga4gh:VA.… digest everywhere, so you can join across systems without heuristics.
explicit assertions
Statements & propositions — a machine-readable "who said variant X relates to
condition Y, how, and on what evidence."
traceable provenance
Provenance — every value traceable to its submission and to the exact transformation that produced it.
In the weekly VCV XML, a single submitted record (a ClinicalAssertion) lives inside every variant's <VariationArchive>. The format is comprehensive and deeply structured — ClinVar-GKM reorganizes the same information into a flat, consistent shape.
A submitted record lives within each variant's archive — complete detail, across millions of variants, every week.
The variant and condition appear both in aggregate and per submission — rich detail to align across records.
SimpleAllele | Haplotype | Genotype at each level; three classification types; optional lists throughout.
ClinVar-GKM reorganizes this same complete information into a consistent, standardized shape — easier to access and analyze, with stable identity and full traceability.
A consistent, semantically rich representation of the entire ClinVar release — variations, submitted (SCV), variant-aggregate (VCV), and condition-aggregate (RCV) classifications.
A single bundle of ~20 cross-linked sections, plus typed Parquet (one file per section).
Every variation, SCV, VCV & RCV from the matching ClinVar XML release.
Published within a day or two of each ClinVar release; synced to ClinVar's cadence.
Public-domain dedication — same terms as the source ClinVar data. No auth, no egress fees.
Built and maintained by the ClinGen driver project — and a real-world validation of the GA4GH VRS, Cat-VRS, and VA-Spec standards against real-world ClinVar data at full scale.
GKM (Genomic Knowledge Model) is the schema set; GKS (Genomic Knowledge Standards) is the GA4GH workstream that curates it. ClinVar-GKM is built on the newest releases across all four — all freshly published.
ga4gh:VA everywhere.
ClinVar's submitted and aggregate records all map onto one shape — each is a Statement carrying a classification, a proposition, evidence, and provenance.
Submitted classification — what one lab, expert panel, or group asserted. The source for every aggregate.
Variant-level aggregate — all SCVs for one variant.
Condition-level aggregate — SCVs for one variant + condition.
VCV and RCV both roll up the same SCVs — in parallel, not a chain. An RCV may be associated with a VCV, but does not drive the VCV's classification.
Aggregate statements reuse the same proposition type and meaning as their SCVs — only how the aggregate classification is derived differs. A deliberate, community-aligned choice.
A GKM Bundle compacts the data: every reused component — alleles, conditions, submitters —
is stored once with an id in a section dictionary, and everything else references it by a
#/section/id JSON pointer, like a foreign key.
{ scv, vcv, rcv, evidenceLine,
varcond- / vartumor- / vartherapy- / varcustom-proposition,
condition, conditionSet, therapy, therapyGroup, submitter,
variation, allele, location, sequenceReference,
gene, copyNumberCount, copyNumberChange }
GKM Starter Kit · Toolkit (in active development) — reads a bundle's JSON Schema to validate, load, and work with the data in Python, so consumers start from a validated model, not raw JSON.
VRS · variant identity
ga4gh:VA digest.
iCNV CountCopyNumberCount — an absolute copy number over a location (VRS).
iCNV ChangeCopyNumberChange — a relative copy-number gain/loss over a location (VRS).
SequenceLocationSequenceLocation — a position/range on a reference sequence.
SequenceReferenceSequenceReference — the reference sequence (refget accession + molecule type).
Cat-VRS · categorical variation
VA-Spec · annotations
GKM-Core · shared building blocks
MappableConcept and ConceptSet thread through everything — classifications, conditions, genes, therapies. Every component here is exercised by real consumers, and is part of the datatype scope they help mature.
Follow the pointers: a statement's proposition names the variation, which resolves through Cat-VRS → VRS down to the reference sequence.
ClinVar-GKM doesn't invent shapes where the community already defined them. It applies the agreed Cat-VRS recipes and VA-Spec community profiles, then layers its own statement conventions.
How a variation's VRS representation is constrained.
The classification frameworks ClinVar submissions use.
Documented mappings, consistently applied every release.
ClinVar carries assertion types and data values the base specs don't cover. ClinVar-GKM meets them within the standards — via implementer-defined proposition types and typed extensions.
6 standard GA4GH types used where they exist:
8 ClinVar custom types for submissions the specs don't model:
Typed Extension name/value pairs carry the values ClinVar consumers depend on — without breaking schema conformance:
Hosted on Cloudflare R2 — free, no authentication, no egress fees. JSON bundle and typed
Parquet, with stable 00-latest URLs.
Complete bundle (+ Parquet), aligned to ClinVar's monthly VCV release.
Added + updated records for every ClinVar release, plus a change manifest.
One file per section — query with DuckDB / pandas, no JSON parsing.
The transformation isn't a black box. The whole pipeline — ClinVar XML in, GKM bundle out — is documented step by step and verified every release.
Extract & normalize variants (incremental).
vrs-python assigns digest ids.
Cat-VRS, conditions, SCV/VCV/RCV statements.
Change-log → deltas → R2.
BigQuery stored procedures + an external VRS step — all documented, with both a Pipeline track and a consumer Output-Reference track.
Releases are oracle-gated — reconstruction checks must match (0-diff) before a release ships. Release notes record what changed.
Every statement links to its contributing submissions and carries provenance extensions — back to the submitter.
Real consumers are the forcing function: as they build on ClinVar-GKM, they solidify — and normalize — the maturity of the GKM model and the scope of datatypes applied.
Linked Data Hub — consuming ClinVar-GKM as standardized input.
Bringing standardized ClinVar into the clinical/EHR setting.
Additional knowledge bases and platform teams are evaluating it now.
Each adopter exercises — and validates — the GKM components in real workflows:
Their feedback drives what matures first and which datatypes expand.
| → Space PgDn | Next slide |
| ← PgUp | Previous slide |
| Home / End | First / last slide |
| S or N | Toggle speaker notes |
| T | Toggle light / dark theme |
| F | Toggle fullscreen |
| ? | Show / hide this help |