Downloads¶
ClinVar-GKM releases are hosted on Cloudflare R2 object storage. All downloads are free with no authentication required and no egress fees.
Distribution follows a full + delta model:
- The complete monthly full bundle — a gzip-compressed JSON file (plus typed Parquet, one file per section) — is published once a month. It contains every variation, statement, proposition, condition, and supporting reference record for that release.
- A weekly delta is published for every ClinVar release. Each delta carries only the records that were added or updated since the prior release, in the same section structure as the full bundle, alongside a
manifest.jsonthat lists per-section adds, updates, and deletes.
A consumer that wants the current state takes the latest monthly full and replays the weekly deltas published since it. See Weekly Deltas for the replay model.
Latest Release¶
Download the most recent full bundle and the most recent weekly delta using the stable URLs below:
| Product | Download | Description |
|---|---|---|
| Monthly full (JSON) | clinvar-gkm_00-latest.json.gz | Latest monthly full bundle |
| Weekly delta (JSON) | clinvar-gkm-delta_00-latest.json.gz | Latest weekly delta (added + updated records) |
| Delta manifest | manifest.json | Per-section adds, updates, and deletes for the latest delta |
| Parquet (full) | See download instructions | Typed Parquet files (one per bundle section) at datasets/parquet/00-latest/, always the newest monthly full |
Download with curl¶
# Latest monthly full bundle (JSON)
curl -O https://pub-f0ad0e0dac0345408dcc95bda20beb42.r2.dev/datasets/clinvar-gkm_00-latest.json.gz
# Latest weekly delta + its manifest
curl -O https://pub-f0ad0e0dac0345408dcc95bda20beb42.r2.dev/deltas/00-latest/clinvar-gkm-delta_00-latest.json.gz
curl -O https://pub-f0ad0e0dac0345408dcc95bda20beb42.r2.dev/deltas/00-latest/manifest.json
# Decompress
gunzip clinvar-gkm_00-latest.json.gz
# Download a single Parquet section from the latest monthly full (e.g., SCV statements)
curl -O https://pub-f0ad0e0dac0345408dcc95bda20beb42.r2.dev/datasets/parquet/00-latest/scv.parquet
# Download all Parquet files (latest monthly full)
for section in sequenceReference location allele copyNumberCount copyNumberChange \
gene variation condition conditionSet therapy therapyGroup submitter \
varcond-proposition vartumor-proposition vartherapy-proposition varcustom-proposition \
evidenceLine vcv_evidenceLine rcv_evidenceLine \
scv vcv rcv; do
curl -O "https://pub-f0ad0e0dac0345408dcc95bda20beb42.r2.dev/datasets/parquet/00-latest/${section}.parquet"
done
Download with Python¶
Prerequisites
These snippets use Python 3 (the download examples need only the standard library). New to Python? Install it from python.org/downloads or your OS package manager.
import urllib.request
BASE = "https://pub-f0ad0e0dac0345408dcc95bda20beb42.r2.dev"
# Download latest monthly full bundle
urllib.request.urlretrieve(
f"{BASE}/datasets/clinvar-gkm_00-latest.json.gz",
"clinvar-gkm_00-latest.json.gz"
)
# Download a specific monthly full bundle
urllib.request.urlretrieve(
f"{BASE}/datasets/clinvar-gkm_2026-06.json.gz",
"clinvar-gkm_2026-06.json.gz"
)
# Download the latest weekly delta + manifest
urllib.request.urlretrieve(
f"{BASE}/deltas/00-latest/clinvar-gkm-delta_00-latest.json.gz",
"clinvar-gkm-delta_00-latest.json.gz"
)
urllib.request.urlretrieve(
f"{BASE}/deltas/00-latest/manifest.json",
"manifest.json"
)
# Download a specific weekly delta (release 2026-07-06 -> dir 2026-0706)
urllib.request.urlretrieve(
f"{BASE}/deltas/2026-0706/clinvar-gkm-delta_2026-0706.json.gz",
"clinvar-gkm-delta_2026-0706.json.gz"
)
# Download an archived full bundle from a prior year
urllib.request.urlretrieve(
f"{BASE}/archives/2025/clinvar-gkm_2025-03.json.gz",
"clinvar-gkm_2025-03.json.gz"
)
# Download a Parquet section from the latest monthly full
urllib.request.urlretrieve(
f"{BASE}/datasets/parquet/00-latest/scv.parquet",
"scv.parquet"
)
# Download a Parquet section from a specific monthly full (checkpoint-addressable)
urllib.request.urlretrieve(
f"{BASE}/datasets/parquet/2026-06/scv.parquet",
"scv.parquet"
)
Validate a downloaded bundle
Every JSON bundle conforms to the ClinVar-GKM bundle schema (JSON Schema Draft 2020-12) — the same schema covers the monthly full, weekly deltas, and sub-bundle extracts. The GKM Toolkit validates a bundle against it in one call.
Weekly Deltas¶
A weekly delta is published for every ClinVar release under deltas/<YYYY-MMDD>/. Each release directory contains three artifact kinds:
deltas/2026-0706/
clinvar-gkm-delta_2026-0706.json.gz added + updated records (same section structure as the full bundle)
manifest.json per-section adds, updates, and deletes for this release
parquet/<section>.parquet typed Parquet for the changed records only
The most recent delta is mirrored at deltas/00-latest/ under stable filenames (clinvar-gkm-delta_00-latest.json.gz, manifest.json, parquet/<section>.parquet).
Delta Bundle¶
The delta bundle has the same shape as the monthly full — a single JSON object with bundle sections at the root, each a keyed collection of objects. The difference is content: a delta contains only the records added or updated since its baseline release. Sections with no additions or updates are absent from the delta bundle. Deleted records are not present in the bundle — they are listed only in the manifest.
manifest.json¶
The manifest describes exactly what changed and which full bundle the delta chain roots at:
{
"release": "2026-07-06",
"baseline_release": "2026-06-29",
"compare_release": "2026-07-06",
"pipeline_version": "clinvar-gkm vX.Y.Z @ 2026-07-06T00:00:00Z",
"checkpoint_full": { "path": "datasets/clinvar-gkm_2026-06.json.gz", "release": "2026-06" },
"sections": {
"allele": { "added": 812, "updated": 34, "deleted": ["ga4gh:VA.oldDigest1"] },
"scv": { "added": 1203, "updated": 517, "deleted": ["clinvar.submission:SCV000000001.1"] },
"vcv": { "added": 44, "updated": 96, "deleted": [] }
},
"counts": { "A": 2063, "U": 647, "D": 2 }
}
| Field | Description |
|---|---|
release |
The ClinVar release date this delta represents |
baseline_release |
The prior release this delta was diffed against — null only on the very first release |
compare_release |
The release the changes are computed to (equals release) |
pipeline_version |
The pipeline build stamp that produced the delta |
checkpoint_full |
The monthly full bundle this delta chain replays onto — {path, release}; null before the first monthly full is published |
sections |
Per-section change summary — added and updated counts plus a deleted list of primary keys |
counts |
Roll-up totals across all sections — A (added), U (updated), D (deleted) |
Deletes live only in the manifest. For each section, deleted is the list of keys that must be removed; the delta bundle itself carries only the added and updated records.
Consumer Replay Model¶
To reconstruct the current state, bootstrap from the monthly full that the latest delta's manifest names in checkpoint_full, then replay the contiguous weekly deltas published after it. For each delta, apply the manifest's deletes first, then upsert every record present in the delta bundle — section by section.
checkpoint_full tells you which monthly full to start from ({path, release}, where release is the YYYY-MM of the full). Replay only the deltas from later months — the deltas within the checkpoint's own month are already reflected in the monthly full. Verify chain integrity as you replay: each delta's baseline_release must equal the previous delta's compare_release. A mismatch means a weekly release is missing from the chain — re-bootstrap from the monthly full rather than applying a partial chain.
import gzip
import json
import urllib.request
BASE = "https://pub-f0ad0e0dac0345408dcc95bda20beb42.r2.dev"
def fetch_json(url):
with urllib.request.urlopen(url) as r:
return json.load(r)
def fetch_json_gz(url):
with urllib.request.urlopen(url) as r:
return json.loads(gzip.decompress(r.read()))
# 1. The latest delta's manifest names the monthly full to bootstrap from.
latest = fetch_json(f"{BASE}/deltas/00-latest/manifest.json")
checkpoint = latest["checkpoint_full"] # {"path": "datasets/clinvar-gkm_2026-06.json.gz",
# "release": "2026-06"} (None on initial rollout,
# before any monthly full exists)
# 2. Load that monthly full as the baseline state: {section: {key: record}}.
state = fetch_json_gz(f"{BASE}/{checkpoint['path']}")
checkpoint_month = checkpoint["release"] # "2026-06"
# 3. From the index, take the dated weekly deltas published AFTER the checkpoint month,
# oldest -> newest. (A delta dir is named YYYY-MMDD packed as "2026-0706", so the first
# 7 chars are its YYYY-MM.) Skip the "latest" mirror and any delta in the checkpoint
# month or earlier — those are already folded into the monthly full.
index = fetch_json(f"{BASE}/index.json")
deltas = sorted(
(d for d in index["deltas"]
if d["release"] != "latest" and d["release"][:7] > checkpoint_month),
key=lambda d: d["release"],
)
# 4. Replay each delta onto the baseline, verifying the chain as we go.
prev_compare = None
for d in deltas:
manifest = fetch_json(f"{BASE}/{d['manifest']}")
# Chain check: after the first, each delta must build on the previous compare_release.
# The first delta builds on the checkpoint month's final release. A mismatch => a
# missing week; re-bootstrap from the monthly full.
if prev_compare is not None and manifest["baseline_release"] != prev_compare:
raise SystemExit(
f"broken chain before {manifest['release']}: re-bootstrap from a monthly full"
)
dirname = d["path"].strip("/").split("/")[-1] # "2026-0706"
delta = fetch_json_gz(f"{BASE}/{d['path']}clinvar-gkm-delta_{dirname}.json.gz")
# 4a. Apply deletes (manifest only), then 4b. upsert added + updated records.
for section, info in manifest["sections"].items():
target = state.setdefault(section, {})
for pk in info["deleted"]:
target.pop(pk, None)
for section, records in delta.items():
state.setdefault(section, {}).update(records)
prev_compare = manifest["compare_release"]
# `state` now reflects the most recent weekly release.
Browse All Releases¶
The file browser below shows all available releases — monthly full bundles, weekly deltas, and archives. It is populated from the release index and updated automatically with each weekly delta and monthly full upload.
Loading release index...
Directory Structure¶
datasets/
clinvar-gkm_00-latest.json.gz latest monthly full bundle (stable URL)
clinvar-gkm_YYYY-MM.json.gz monthly full bundles (current year)
datasets/parquet/00-latest/
{section}.parquet typed Parquet for the latest monthly full (stable URL)
datasets/parquet/YYYY-MM/
{section}.parquet typed Parquet for a specific monthly full (current year)
deltas/00-latest/
clinvar-gkm-delta_00-latest.json.gz latest weekly delta bundle (stable URL)
manifest.json latest delta manifest
parquet/{section}.parquet typed Parquet for the latest delta
deltas/YYYY-MMDD/
clinvar-gkm-delta_YYYY-MMDD.json.gz weekly delta bundle (added + updated records)
manifest.json per-release change manifest
parquet/{section}.parquet typed Parquet for the changed records
archives/{YYYY}/
clinvar-gkm_YYYY-MM.json.gz monthly full bundles from prior years
parquet/YYYY-MM/{section}.parquet typed Parquet month sets from prior years
index.json release index (datasets, archives, deltas)
Parquet Files¶
Typed Parquet files are produced for each monthly full and organized by month, mirroring the JSON bundle lifecycle. Each monthly full lands in a dated directory datasets/parquet/YYYY-MM/, and datasets/parquet/00-latest/ always points at the newest monthly full (stable URL). At year boundaries, prior-year month sets move to archives/{YYYY}/parquet/YYYY-MM/ and are retained indefinitely. Because each month set is preserved, a delta chain's checkpoint_full monthly Parquet stays available for reconstruction after later months publish. The weekly delta Parquet (changed records only) lives separately under deltas/<YYYY-MMDD>/parquet/.
Each Parquet file contains one bundle section with typed, query-friendly columns extracted from the JSON objects. Every section includes an id column (the object identifier) and a data column (the full JSON object as a string), plus additional typed columns for key fields — enabling efficient filtering and aggregation without parsing JSON.
Available Parquet files (20 sections):
| File | Description |
|---|---|
sequenceReference.parquet |
NCBI RefSeq sequence references |
location.parquet |
VRS SequenceLocation records |
allele.parquet |
VRS Allele records |
copyNumberCount.parquet |
VRS CopyNumberCount records |
copyNumberChange.parquet |
VRS CopyNumberChange records |
gene.parquet |
Gene records (MappableConcept with NCBI Gene / HGNC codings) |
variation.parquet |
CategoricalVariant records (Cat-VRS) |
condition.parquet |
Condition records (traits) |
conditionSet.parquet |
ConditionSet records (trait sets) |
therapy.parquet |
Therapy records (drug therapies, content-addressed) |
therapyGroup.parquet |
TherapyGroup records (combination therapies) |
submitter.parquet |
Submitter organization records |
varcond-proposition.parquet |
Variant×condition propositions (Pathogenicity, ClinicalSignificance, Diagnostic, Prognostic) |
vartumor-proposition.parquet |
Variant×tumorType propositions (Oncogenicity) |
vartherapy-proposition.parquet |
Variant×therapy propositions (TherapeuticResponse) |
varcustom-proposition.parquet |
Custom variant×condition propositions |
evidenceLine.parquet |
SCV evidence line records |
vcv_evidenceLine.parquet |
VCV evidence line records |
rcv_evidenceLine.parquet |
RCV evidence line records |
scv.parquet |
SCV statement records |
vcv.parquet |
VCV aggregate statement records |
rcv.parquet |
RCV aggregate statement records |
Working with Parquet Files¶
Download the Parquet files you need, then query them locally. The R2 hosting has rate limits and is designed for file downloads, not as a remote query endpoint for tools like DuckDB.
Prerequisites
The examples below use the DuckDB CLI and/or Python with pandas + pyarrow (pip install pandas pyarrow). If you don't already have these tools, install them from the linked pages first — each subsection also shows its own one-line install command.
Download¶
Download individual sections or all files at once:
# 00-latest = the newest monthly full; swap for a dated month (e.g. .../parquet/2026-06) to pin a release.
BASE="https://pub-f0ad0e0dac0345408dcc95bda20beb42.r2.dev/datasets/parquet/00-latest"
mkdir -p clinvar-gkm-parquet && cd clinvar-gkm-parquet
# Download specific sections
curl -O "${BASE}/scv.parquet"
curl -O "${BASE}/varcond-proposition.parquet"
curl -O "${BASE}/condition.parquet"
# Or download all 20 sections
for section in sequenceReference location allele copyNumberCount copyNumberChange \
gene variation condition conditionSet therapy therapyGroup submitter \
varcond-proposition vartumor-proposition vartherapy-proposition varcustom-proposition \
evidenceLine vcv_evidenceLine rcv_evidenceLine \
scv vcv rcv; do
curl -O "${BASE}/${section}.parquet"
done
DuckDB¶
DuckDB is the fastest way to explore Parquet files — it queries them directly with no data loading step.
# Query SCV statements
duckdb -c "
SELECT id, classification.name AS classification, direction, strength.name AS strength, quality.name AS quality
FROM 'scv.parquet'
WHERE classification.name = 'Pathogenic'
LIMIT 10;
"
# Count classifications across all SCVs
duckdb -c "
SELECT classification.name AS classification, direction, COUNT(*) as n
FROM 'scv.parquet'
GROUP BY classification.name, direction
ORDER BY n DESC;
"
# Join SCVs with propositions to find pathogenic variants for a specific condition
duckdb -c "
SELECT s.id, s.classification.name AS classification, p.predicate, p.object_condition_id
FROM 'scv.parquet' s
JOIN 'varcond-proposition.parquet' p ON s.proposition_id = p.id
WHERE s.classification.name = 'Pathogenic'
AND p.object_condition_id LIKE '%clinvar.trait:9580%'
LIMIT 10;
"
DuckDB also works from Python:
import duckdb
df = duckdb.sql("""
SELECT id, classification.name AS classification, direction, strength.name AS strength, quality.name AS quality
FROM 'scv.parquet'
WHERE classification.name = 'Pathogenic'
LIMIT 100
""").df()
print(df)
pandas / pyarrow¶
import pandas as pd
# Load a section into a DataFrame
scv = pd.read_parquet("scv.parquet")
# Filter pathogenic SCVs — classification is a MappableConcept struct (dict), so read .name
pathogenic = scv[scv["classification"].map(lambda c: c and c.get("name")) == "Pathogenic"]
print(f"{len(pathogenic)} pathogenic SCVs")
# Access the full JSON when you need nested fields
import json
record = json.loads(pathogenic.iloc[0]["data"])
print(record["proposition"])
Column Reference¶
Statement sections (scv, vcv, rcv) share a common set of typed columns:
| Column | Type | Description |
|---|---|---|
id |
string | Statement identifier |
type |
string | Statement type |
proposition_id |
string | FK to the matching proposition Parquet — one of varcond-proposition, vartumor-proposition, vartherapy-proposition, varcustom-proposition, per the proposition's datatype |
classification |
struct (MappableConcept) | Use classification.name (or classification.primaryCoding.code) — e.g., "Pathogenic" |
strength |
struct (MappableConcept) | Use strength.name — e.g., "definitive", "likely" |
direction |
string | Evidence direction ("supports", "disputes", "neutral") |
quality |
struct (MappableConcept) | Use quality.name — submission level, e.g., "criteria provided" |
has_evidence_lines |
list\<string> | FK references to evidence line Parquet (evidenceLine for SCV, vcv_evidenceLine for VCV, rcv_evidenceLine for RCV) |
extensions |
string | JSON array of extensions |
data |
string | Full JSON object |
SCV statements include additional columns: description, contributions, reported_in, specified_by.
The four proposition sections are typed per datatype: varcond-proposition (subject_variant_id, predicate, object_condition_id, type, gene_context_name, mode_of_inheritance, penetrance), vartumor-proposition (object_tumor_type_id, …), vartherapy-proposition (object_therapy, condition_qualifier_id, …), and varcustom-proposition (custom_proposition_type, subject_id, object_id, qualifiers) — enabling JOINs across statements, variants, and conditions without parsing JSON.
Every section includes id and data at minimum. Run DESCRIBE in DuckDB to see the full schema for any section.
Example Queries¶
These examples demonstrate cross-section JOINs using typed columns. Most analytical queries can be answered without parsing JSON — the varcond proposition's gene_context_name column carries the gene symbol directly, and the condition's primaryCoding struct carries the MedGen code. (These queries join scv.parquet to varcond-proposition.parquet, the variant×condition group that holds pathogenicity and clinical-significance propositions; the other three proposition groups — vartumor, vartherapy, varcustom — have their own typed columns.)
All SCVs for a gene — detailed view:
-- SCVs for BRCA1: classification, review status, condition
SELECT
s.id AS scv_id,
s.classification.name AS classification,
s.direction,
s.strength.name AS strength,
s.quality.name AS review_status,
p.gene_context_name AS gene,
c.name AS condition_name,
c.primaryCoding.code AS condition_code
FROM 'scv.parquet' s
JOIN 'varcond-proposition.parquet' p ON s.proposition_id = p.id
LEFT JOIN 'condition.parquet' c ON p.object_condition_id = c.id
WHERE p.gene_context_name = 'BRCA1'
ORDER BY s.classification.name;
Classification summary for a gene:
-- Count SCVs by classification and review status for BRCA2
SELECT
p.gene_context_name AS gene,
s.classification.name AS classification,
s.quality.name AS review_status,
s.direction,
COUNT(*) AS scv_count
FROM 'scv.parquet' s
JOIN 'varcond-proposition.parquet' p ON s.proposition_id = p.id
WHERE p.gene_context_name = 'BRCA2'
GROUP BY ALL
ORDER BY scv_count DESC;
Restrict to submissions with criteria provided:
-- Only expert panel and criteria-provided SCVs for a gene
SELECT
s.id AS scv_id,
s.classification.name AS classification,
s.quality.name AS review_status,
c.name AS condition_name
FROM 'scv.parquet' s
JOIN 'varcond-proposition.parquet' p ON s.proposition_id = p.id
LEFT JOIN 'condition.parquet' c ON p.object_condition_id = c.id
WHERE p.gene_context_name = 'TP53'
AND s.quality.name IN ('criteria provided', 'reviewed by expert panel')
ORDER BY s.classification.name;
Cross-gene comparison — classification breakdown for multiple genes:
-- Compare pathogenicity classification distributions across genes
SELECT
p.gene_context_name AS gene,
s.classification.name AS classification,
COUNT(*) AS n
FROM 'scv.parquet' s
JOIN 'varcond-proposition.parquet' p ON s.proposition_id = p.id
WHERE p.gene_context_name IN ('BRCA1', 'BRCA2', 'TP53', 'MLH1')
AND s.quality.name = 'criteria provided'
GROUP BY gene, s.classification.name
ORDER BY gene, n DESC;
Accessing fields not in typed columns:
Some fields — like submitter names, HGVS expressions, and assertion methods — are only available in the data column (full JSON string). Use DuckDB's json_extract_string to access them:
-- SCVs with submitter name and assertion method (from JSON)
SELECT
s.id AS scv_id,
s.classification.name AS classification,
p.gene_context_name AS gene,
json_extract_string(s.data, '$.contributions[0].agent.name') AS submitter,
json_extract_string(s.data, '$.specifiedBy.name') AS method
FROM 'scv.parquet' s
JOIN 'varcond-proposition.parquet' p ON s.proposition_id = p.id
WHERE p.gene_context_name = 'BRCA1'
AND s.classification.name = 'Pathogenic'
AND s.quality.name = 'reviewed by expert panel'
LIMIT 20;
Summary:
| Approach | Best for | Tradeoff |
|---|---|---|
| Typed columns only | Filtering, counting, grouping, JOINs on classification, gene, condition, review status | Fast; covers most analytical questions |
Typed columns + json_extract_string |
Ad-hoc queries needing submitter names, HGVS, methods | Slightly slower; syntax is verbose |
Parse data column in application code |
Bulk processing needing many nested fields | Full flexibility; requires application-side JSON parsing |
Applying Deltas (keeping a Parquet set current)¶
The weekly delta ships typed Parquet too, so you can maintain a current Parquet set without re-downloading the full monthly bundle. The model mirrors the JSON replay: start from the checkpoint monthly full, then apply each weekly delta in order.
Each delta is a per-section keyed upsert plus a delete list:
deltas/<YYYY-MMDD>/parquet/<section>.parquetholds the added and updated rows for that section (keyed byid). Sections with no adds/updates have no file.deltas/<YYYY-MMDD>/manifest.jsonholds the deletes —sections.<section>.deletedis the list ofidvalues removed this release. Deletes are not in the Parquet.
Every section Parquet — full and delta alike — exposes an id column, and the manifest's deleted keys are those same id values, so applying a delta is: drop from the full every row whose id appears in the delta or in the delete list, then append the delta rows. Because the full and delta Parquet for a section are produced by the identical schema, their columns line up exactly (UNION ALL BY NAME).
Download the pieces¶
BASE="https://pub-f0ad0e0dac0345408dcc95bda20beb42.r2.dev"
# 1. The checkpoint monthly full named by the delta's manifest (e.g. 2026-06).
curl -s "${BASE}/deltas/00-latest/manifest.json" -o manifest.json
CKPT=$(python3 -c "import json;print(json.load(open('manifest.json'))['checkpoint_full']['release'])")
# 2. The checkpoint full Parquet set -> full/, and the delta set -> delta/.
mkdir -p full delta
for section in sequenceReference location allele copyNumberCount copyNumberChange \
gene variation condition conditionSet therapy therapyGroup submitter \
varcond-proposition vartumor-proposition vartherapy-proposition varcustom-proposition \
evidenceLine vcv_evidenceLine rcv_evidenceLine scv vcv rcv; do
curl -sf "${BASE}/datasets/parquet/${CKPT}/${section}.parquet" -o "full/${section}.parquet" || true
curl -sf "${BASE}/deltas/00-latest/parquet/${section}.parquet" -o "delta/${section}.parquet" || true
done
Apply one delta¶
import duckdb, json, os, shutil
con = duckdb.connect()
manifest = json.load(open("manifest.json"))
# Start from a copy of the checkpoint full set; overwrite only the changed sections.
shutil.copytree("full", "updated", dirs_exist_ok=True)
for section, info in manifest["sections"].items():
full = f"full/{section}.parquet" # checkpoint (baseline) section
delta = f"delta/{section}.parquet" # added + updated rows (may be absent)
out = f"updated/{section}.parquet"
deleted = info["deleted"] # ids removed this release
has_delta = os.path.exists(delta)
if not os.path.exists(full):
# Section new since the checkpoint: the delta rows ARE the section.
if has_delta:
shutil.copyfile(delta, out)
continue
if not has_delta:
# Deletes only — filter the baseline, nothing to append.
con.execute("""
COPY (SELECT * FROM read_parquet(?) WHERE id NOT IN (SELECT unnest(?)))
TO ? (FORMAT PARQUET)
""", [full, deleted, out])
else:
# Drop updated + deleted keys from the baseline, then append the delta rows.
con.execute("""
COPY (
SELECT * FROM read_parquet(?) -- baseline full
WHERE id NOT IN (SELECT id FROM read_parquet(?)) -- drop updated keys
AND id NOT IN (SELECT unnest(?)) -- drop deleted keys
UNION ALL BY NAME
SELECT * FROM read_parquet(?) -- append added + updated
) TO ? (FORMAT PARQUET)
""", [full, delta, deleted, delta, out])
# `updated/<section>.parquet` now reflects the release the manifest names.
Chaining multiple weeks¶
To advance across several weekly deltas, apply them oldest→newest, feeding each step's output back in as the next step's full/. Use the index.json deltas list (skip latest, keep only releases after the checkpoint month) and verify baseline_release == prior compare_release at each step — a mismatch means a missing week, so re-bootstrap from the monthly full rather than applying a partial chain. This is the same contiguity rule as the JSON replay; only the per-section apply differs (a keyed Parquet upsert instead of a dict merge).
A single-file convenience alternative: because the pipeline also refreshes datasets/parquet/00-latest/ on every monthly full, you can skip replay entirely and pull the current full set directly — replay is for reconstructing a specific historical release or minimizing download size between monthly fulls.
Release Cadence¶
A weekly delta is published for every ClinVar release, typically within 1-2 days of each ClinVar XML release. Each delta lands under deltas/<YYYY-MMDD>/ and is mirrored at deltas/00-latest/.
A monthly full bundle is published once a month, aligned to ClinVar's own monthly VCV releases. When ClinVar posts ClinVarVCVRelease_YYYY-MM.xml.gz at its XML index (early in month YYYY-MM), our clinvar-gkm_YYYY-MM full is built from the most recent weekly release before ClinVar's monthly cut datetime — so e.g. the 2026-07 full comes from the 2026-06-27 release, the last one before ClinVar's _2026-07 cut. That upload writes datasets/clinvar-gkm_YYYY-MM.json.gz and its Parquet month set datasets/parquet/YYYY-MM/, and refreshes the datasets/clinvar-gkm_00-latest.json.gz and datasets/parquet/00-latest/ pointers. Each delta manifest's checkpoint_full records which monthly full its chain replays onto — and because each Parquet month set is retained, that checkpoint's Parquet stays reconstructable after later months publish.
At year boundaries, the prior year's monthly full bundles and Parquet month sets are moved to archives/{YYYY}/ (archives/{YYYY}/parquet/YYYY-MM/ for Parquet). All monthly archives are retained indefinitely.
There is no weekly full bundle — weekly changes are distributed as deltas only. Consumers that need the full weekly state reconstruct it by replaying deltas onto the latest monthly full, as shown in Consumer Replay Model.
Feedback¶
This project is in active development and we welcome community feedback. If you encounter data quality issues, have questions about the output format, or want to suggest improvements:
- Open an issue on GitHub
- Include the release date and specific records involved