A GA4GH / GKM Primer · ClinGen

ClinVar-GKM

ClinVar — standardized. Every ClinVar release rebuilt as a computable GKM bundle: VRS · Cat-VRS · VA-Spec, on GKM-Core. Built on the newest GKM releases across all four.
GA4GH Connect · Oct 2, 2026 ~8-minute primer New to VA-Spec? Start here CC0 / open data
Full documentation — every topic links to its page: clingen-data-model.github.io/clinvar-gkm
Use ← / → to navigate · press ? for all shortcuts
The starting point

ClinVar's complete truth lives in the weekly XML

The full, authoritative snapshot — every variation, submission, aggregate, and provenance detail — is released weekly as XML, alongside convenient VCF / CSV reports and targeted Entrez E-utilities queries. ClinVar-GKM adds a consistent, standards-native organization of that same data.

XML = source of truth

The only representation carrying the complete database snapshot, refreshed weekly. Everything else is derived from it.

Focused products for focused needs

ClinVar's VCF, tab-delimited summaries, and E-utilities queries are convenient, targeted views into the full archive.

Where ClinVar-GKM helps

A GA4GH-standard reorganization of the same complete data — adding computable variant identity and an explicit statement model for consistent, accurate analysis.

What consumers really want: a stable variant id VRS identity — the same variant resolves to the same ga4gh:VA.… digest everywhere, so you can join across systems without heuristics. explicit assertions Statements & propositions — a machine-readable "who said variant X relates to condition Y, how, and on what evidence." traceable provenance Provenance — every value traceable to its submission and to the exact transformation that produced it.
The source of truth, up close

The complete record: one SCV, five levels in

In the weekly VCV XML, a single submitted record (a ClinicalAssertion) lives inside every variant's <VariationArchive>. The format is comprehensive and deeply structured — ClinVar-GKM reorganizes the same information into a flat, consistent shape.

<ClinVarVariationRelease> weekly snapshot root <VariationArchive Accession="VCV000012345"> one per variant · millions <ClassifiedRecord> <SimpleAllele> Location · HGVSlist · GeneList variant — aggregate form <RCVList> <RCVAccession …/> <Classifications> Germline · Somatic · Oncogenicity 3 types <ClinicalAssertionList> <ClinicalAssertion> ◀ one SCV — 5 levels deep <ClinVarAccession Accession="SCV000045678" SubmitterName="…" OrgID="…"/> <Classification> ReviewStatus · Description per-SCV <SimpleAllele> GeneList · Location · HGVSlist variant — submitted form <TraitSet><Trait> Name · XRef submitted condition <ObservedInList> Sample · Method · ObservedData <Comment/> <Citation/> <XRefList/> <AttributeSet/> </ClinicalAssertion> × N submitters </ClinicalAssertionList> <TraitMappingList/> · <DeletedSCVList/> … </ClassifiedRecord> … </VariationArchive>

Comprehensive & deeply nested

A submitted record lives within each variant's archive — complete detail, across millions of variants, every week.

Represented at multiple levels

The variant and condition appear both in aggregate and per submission — rich detail to align across records.

Flexible by design

SimpleAllele | Haplotype | Genotype at each level; three classification types; optional lists throughout.

ClinVar-GKM reorganizes this same complete information into a consistent, standardized shape — easier to access and analyze, with stable identity and full traceability.

Structure per ClinVar's VCV schema (ClinVar_VCV.xsd), document order; element set abbreviated, ids illustrative. · Pipeline
What it is

Every ClinVar release, rebuilt as one GKM bundle

A consistent, semantically rich representation of the entire ClinVar release — variations, submitted (SCV), variant-aggregate (VCV), and condition-aggregate (RCV) classifications.

1 JSON

A single bundle of ~20 cross-linked sections, plus typed Parquet (one file per section).

100%

Every variation, SCV, VCV & RCV from the matching ClinVar XML release.

~1–2 d

Published within a day or two of each ClinVar release; synced to ClinVar's cadence.

CC0

Public-domain dedication — same terms as the source ClinVar data. No auth, no egress fees.

Built and maintained by the ClinGen driver project — and a real-world validation of the GA4GH VRS, Cat-VRS, and VA-Spec standards against real-world ClinVar data at full scale.

The standards

GKM = three GA4GH specs on a shared core

GKM (Genomic Knowledge Model) is the schema set; GKS (Genomic Knowledge Standards) is the GA4GH workstream that curates it. ClinVar-GKM is built on the newest releases across all four — all freshly published.

VA-Spec 1.1.0 Variant annotations — Statements, Propositions, Evidence Lines, provenance. Every classification is a VA-Spec Statement.
Cat-VRS 1.1.1 Categorical variation — canonical alleles & copy-number categories with defining constraints.
VRS 2.1.1 Variant identity — a computable, digest-based id. Same variant → same ga4gh:VA everywhere.
GKM-Core 1.3.0 Shared building blocks — MappableConcept, ConceptSet, Agent, Extension — used across all three.
All four freshly released · each ships a Python library (vrs-python · cat-vrs-python · va-spec-python) that parses these bundles into validated models. · GA4GH Standards
The core idea

Classifications become VA-Spec Statements

ClinVar's submitted and aggregate records all map onto one shape — each is a Statement carrying a classification, a proposition, evidence, and provenance.

SCV

Submitted classification — what one lab, expert panel, or group asserted. The source for every aggregate.

aggregate by variant →

VCV

Variant-level aggregate — all SCVs for one variant.

by variant + condition →

RCV

Condition-level aggregate — SCVs for one variant + condition.

Two separate aggregations

VCV and RCV both roll up the same SCVs — in parallel, not a chain. An RCV may be associated with a VCV, but does not drive the VCV's classification.

Same proposition & semantics

Aggregate statements reuse the same proposition type and meaning as their SCVs — only how the aggregate classification is derived differs. A deliberate, community-aligned choice.

The output

The GKM Bundle — one JSON object, cross-linked by pointers

A GKM Bundle compacts the data: every reused component — alleles, conditions, submitters — is stored once with an id in a section dictionary, and everything else references it by a #/section/id JSON pointer, like a foreign key.

scv
ClinvarScvStatement · va-spec Statement
clinvar.submission:SCV001571657.2
proposition → #/varcond-proposition/
contributions[].agent → #/submitter/
hasEvidenceLines[] → #/evidenceLine/
proposition 1..1
varcond-proposition
VariantPathogenicityProposition
SCV001571657-PATH
subject → #/variation/
object → #/condition|conditionSet/

{ scv, vcv, rcv, evidenceLine,
  varcond- / vartumor- / vartherapy- / varcustom-proposition,
  condition, conditionSet, therapy, therapyGroup, submitter,
  variation, allele, location, sequenceReference,
  gene, copyNumberCount, copyNumberChange }

  • Compact — each reused object is stored once under an id; others point to it, never copy it.
  • Typed & self-describing — one object type per section.
  • Traversable — follow pointers from a statement to the sequence reference.

GKM Starter Kit · Toolkit (in active development) — reads a bundle's JSON Schema to validate, load, and work with the data in Python, so consumers start from a validated model, not raw JSON.

What's in play

The GKM classes ClinVar-GKM uses (hover any chip)

VRS · variant identity

iAlleleVRS Allele — a specific sequence change at a Location, globally identified by a ga4gh:VA digest. iCNV CountCopyNumberCount — an absolute copy number over a location (VRS). iCNV ChangeCopyNumberChange — a relative copy-number gain/loss over a location (VRS). SequenceLocationSequenceLocation — a position/range on a reference sequence. SequenceReferenceSequenceReference — the reference sequence (refget accession + molecule type).

Cat-VRS · categorical variation

iCategoricalVariantCategoricalVariant — Cat-VRS grouping tying a ClinVar variation to its VRS representation via constraints. Recipes: CanonicalAllele, CategoricalCnvCount, CategoricalCnvChange.

VA-Spec · annotations

iStatementStatement — a complete assertion: classification + proposition + evidence + provenance. SCV / VCV / RCV. iPropositionProposition — the claim being assessed: subject (variant), predicate, object (condition/therapy). iEvidenceLineEvidenceLine — links a proposition to evidence items with direction + strength.

GKM-Core · shared building blocks

iMappableConceptMappableConcept — one concept (conceptType + name + codings/mappings). Used for classifications, conditions, genes, therapies. iConceptSetConceptSet — a group of concepts with a membershipOperator (AND/OR). Multi-condition objects; combination therapies. AgentAgent — a submitter / contributor on a statement's contributions. ExtensionExtension — name/value pair carrying ClinVar-specific metadata outside the core specs.

MappableConcept and ConceptSet thread through everything — classifications, conditions, genes, therapies. Every component here is exercised by real consumers, and is part of the datatype scope they help mature.

Worked example

One variant, end to end — clinvar:10 (HFE p.His63Asp)

Follow the pointers: a statement's proposition names the variation, which resolves through Cat-VRS → VRS down to the reference sequence.

scv
ClinvarScvStatement · Pathogenic
clinvar.submission:SCV00… (illustrative)
classification: Pathogenic direction / strengthDirection = supports / disputes / neutral; Strength = definitive / likely / strong / potential.
quality: criteria provided review statusquality (va-spec — a newer attribute) — the SCV's review status / submission level (e.g. "criteria provided", "expert panel"). It maps to ClinVar's star rating and drives aggregation: the highest review status wins when SCVs roll up into a VCV / RCV.
proposition → #/varcond-proposition/
proposition
varcond-proposition
VariantPathogenicityProposition
subject → #/variation/clinvar:10
predicate: isCausalFor
object → #/condition/ (hemochromatosis)
variation
CategoricalVariant · Cat-VRS (CanonicalAllele)
clinvar:10
name: NM_000410.4(HFE):c.187C>G (p.His63Asp)
members[] → #/allele/
members
allele
Allele · VRS
ga4gh:VA.ELQCnIBGqaTl0AEE…
location → #/location/ → #/sequenceReference/
That ga4gh:VA digest is the same wherever this variant appears — the join key across systems.
Standards in practice

Community profiles & recipes — in play

ClinVar-GKM doesn't invent shapes where the community already defined them. It applies the agreed Cat-VRS recipes and VA-Spec community profiles, then layers its own statement conventions.

Cat-VRS recipes

  • CanonicalAllele
  • CategoricalCnvCount
  • CategoricalCnvChange

How a variation's VRS representation is constrained.

VA-Spec community profiles

  • ACMG/AMP 2015 — germline pathogenicity
  • AMP/ASCO/CAP 2017 — somatic clinical significance
  • CCV / VICC 2022 — oncogenicity

The classification frameworks ClinVar submissions use.

ClinVar-GKM profiles

  • 14 statement types
  • classification → direction + strength
  • review status → star ranks & aggregation

Documented mappings, consistently applied every release.

Implementer profiles

Where ClinVar needs more: custom propositions + extensions

ClinVar carries assertion types and data values the base specs don't cover. ClinVar-GKM meets them within the standards — via implementer-defined proposition types and typed extensions.

Proposition types

6 standard GA4GH types used where they exist:

PathogenicityOncogenicityClinicalSignificance TherapeuticResponseDiagnosticPrognostic

8 ClinVar custom types for submissions the specs don't model:

RiskFactorProtectiveDrugResponse AffectsAssociationConfersSensitivity OtherNotProvided

ClinVar-specific extensions

Typed Extension name/value pairs carry the values ClinVar consumers depend on — without breaking schema conformance:

  • Submitted conditions & mappings
  • Review status & submission details
  • HGVS / SPDI / VCF expressions
  • Conflicting-classification explanations
Standard where possible · custom where ClinVar requires it · always schema-valid. · Propositions · SCV Extensions
Get the data

Downloads: monthly full + weekly deltas

Hosted on Cloudflare R2 — free, no authentication, no egress fees. JSON bundle and typed Parquet, with stable 00-latest URLs.

Monthly full

Complete bundle (+ Parquet), aligned to ClinVar's monthly VCV release.

Weekly delta

Added + updated records for every ClinVar release, plus a change manifest.

Typed Parquet

One file per section — query with DuckDB / pandas, no JSON parsing.

# latest monthly full bundle
curl -O https://….r2.dev/datasets/clinvar-gkm_00-latest.json.gz
# latest weekly delta + manifest (added / updated / deleted)
curl -O https://….r2.dev/deltas/00-latest/clinvar-gkm-delta_00-latest.json.gz
Replay model: latest full+ deltas since= current state manifest-drivenmanifest.json lists per-section Added / Updated / Deleted keys, and names the checkpoint full to replay onto.
Transparency

You can see exactly how every value was produced

The transformation isn't a black box. The whole pipeline — ClinVar XML in, GKM bundle out — is documented step by step and verified every release.

1 · Identity

Extract & normalize variants (incremental).

→

2 · VRS

vrs-python assigns digest ids.

→

3 · Model

Cat-VRS, conditions, SCV/VCV/RCV statements.

→

4 · Publish

Change-log → deltas → R2.

Open method

BigQuery stored procedures + an external VRS step — all documented, with both a Pipeline track and a consumer Output-Reference track.

Verified releases

Releases are oracle-gated — reconstruction checks must match (0-diff) before a release ships. Release notes record what changed.

Traceable records

Every statement links to its contributing submissions and carries provenance extensions — back to the submitter.

Momentum

Groups are adopting this for production

Real consumers are the forcing function: as they build on ClinVar-GKM, they solidify — and normalize — the maturity of the GKM model and the scope of datatypes applied.

ClinGen LDH

Linked Data Hub — consuming ClinVar-GKM as standardized input.

Epic

Bringing standardized ClinVar into the clinical/EHR setting.

…and more early adopters

Additional knowledge bases and platform teams are evaluating it now.

Consumers demonstrate component utility

Each adopter exercises — and validates — the GKM components in real workflows:

MappableConceptConceptSetCategoricalVariant AlleleCNV CountCNV Change StatementPropositionEvidenceLine

Their feedback drives what matures first and which datatypes expand.

Start here

Explore it, build on it, tell us what's missing

Docs & data

Feedback loop

  • GitHub issues — data questions, bugs, format requests
  • Roadmap — driven by labeled GitHub Discussions (upvote to prioritize)
  • Tell us which components & datatypes you need next
Thank you — questions welcome. Press ? for navigation shortcuts.
1 / 15