Export & Distribute¶
The final pipeline step exports the gkm_dict_* tables from BigQuery, assembles them into a single keyed JSON bundle, exports Parquet files directly from BigQuery, and uploads to Cloudflare R2 for public distribution.
The gkm_dict_* tables are the published product — they are built directly by the statement procedures (steps 5–7) and the change-log step (step 8). The retired gkm_json_proc is not part of this path; its inlined JSON-render tables were never published. See Retired: gkm_json_proc.
Distribution follows a full + delta model:
- The complete monthly full bundle (JSON + Parquet) is published once a month via
release-gkm.sh/upload-gkm-to-r2.sh. - A weekly delta — added and updated records plus a change manifest — is published for every release via
release-gkm-delta.sh/upload-gkm-delta-to-r2.sh.
Workflow¶
The export and distribution process uses four steps, executed in sequence. The release-gkm.sh wrapper runs all four automatically, or each step can be run individually.
Step 1: Export Dictionaries to GCS¶
export-gkm-dicts.sh exports all dictionary and statement tables from BigQuery to Google Cloud Storage in two formats:
- NDJSON — sharded, gzip-compressed files for JSON bundle assembly (
gkm-dicts/) - Parquet — Snappy-compressed files exported natively via
bq extract(gkm-dicts-parquet/)
# Export both NDJSON and Parquet
./src/scripts/export-gkm-dicts.sh clinvar_2026_06_14_v2_5_0 clinvar-gkm gkm-dicts
# Export Parquet only (skip NDJSON)
./src/scripts/export-gkm-dicts.sh clinvar_2026_06_14_v2_5_0 clinvar-gkm gkm-dicts --parquet-only
The script exports the following 19 tables:
| Table | NDJSON Output | Parquet Output |
|---|---|---|
gkm_dict_sequence_reference |
sequenceReference-*.ndjson.gz |
sequenceReference.parquet |
gkm_dict_location |
location-*.ndjson.gz |
location.parquet |
gkm_dict_allele |
allele-*.ndjson.gz |
allele.parquet |
gkm_dict_copy_number_count |
copyNumberCount-*.ndjson.gz |
copyNumberCount.parquet |
gkm_dict_copy_number_change |
copyNumberChange-*.ndjson.gz |
copyNumberChange.parquet |
gkm_dict_gene |
gene-*.ndjson.gz |
gene.parquet |
gkm_dict_variation |
variation-*.ndjson.gz |
variation.parquet |
gkm_dict_condition |
condition-*.ndjson.gz |
condition.parquet |
gkm_dict_condition_set |
conditionSet-*.ndjson.gz |
conditionSet.parquet |
gkm_dict_submitter |
submitter-*.ndjson.gz |
submitter.parquet |
gkm_dict_proposition |
proposition-*.ndjson.gz |
proposition.parquet |
gkm_dict_evidence_line |
evidenceLine-*.ndjson.gz |
evidenceLine.parquet |
gkm_dict_vcv_proposition |
vcv_proposition-*.ndjson.gz |
vcv_proposition.parquet |
gkm_dict_vcv_evidence_line |
vcv_evidenceLine-*.ndjson.gz |
vcv_evidenceLine.parquet |
gkm_dict_rcv_proposition |
rcv_proposition-*.ndjson.gz |
rcv_proposition.parquet |
gkm_dict_rcv_evidence_line |
rcv_evidenceLine-*.ndjson.gz |
rcv_evidenceLine.parquet |
gkm_dict_scv |
scv-*.ndjson.gz |
scv.parquet |
gkm_dict_vcv |
vcv-*.ndjson.gz |
vcv.parquet |
gkm_dict_rcv |
rcv-*.ndjson.gz |
rcv.parquet |
BigQuery EXTRACT shards large NDJSON tables across multiple files automatically. Parquet files are exported as single files per table; BigQuery may auto-shard very large tables with numeric suffixes.
Step 2: Assemble Bundle¶
assemble-gkm-dicts.py reads all NDJSON shard files and assembles them into a single keyed JSON bundle file.
<source> is a local path or gs:// URI containing the NDJSON shards. <date> is the ClinVar release date (YYYY-MM-DD); the output path is derived as /tmp/clinvar-gkm-{date}.json.gz. By default, source files are deleted after assembly; use --keep-source to retain them.
The script assembles 20 bundle sections in a fixed order: sequenceReference, location, allele, copyNumberCount, copyNumberChange, gene, variation, condition, conditionSet, therapy, therapyGroup, submitter, varcond-proposition, vartumor-proposition, vartherapy-proposition, varcustom-proposition, evidenceLine, scv, vcv, rcv. The therapy / therapyGroup sections hold content-addressed (deduplicated) drug therapies referenced by VariantTherapeuticResponseProposition.object via #/therapy/ and #/therapyGroup/. Each section is a keyed object where the key is the record's unique identifier. Propositions from SCV, VCV, and RCV are delivered in four datatype-homogeneous sections keyed by their (subject, object) signature — varcond-proposition (variant×condition), vartumor-proposition (variant×tumorType), vartherapy-proposition (variant×therapy), varcustom-proposition (custom variant×condition); evidence line shards are merged into a single evidenceLine section.
Install orjson for best performance:
Step 3: Download and Merge Parquet from GCS¶
release-gkm.sh downloads Parquet shards from GCS, merges them into one file per section using DuckDB, and stages the merged files for upload. BigQuery exports may produce multiple shards per table (e.g., allele-000000000000.parquet, allele-000000000001.parquet); this step consolidates them into a single allele.parquet.
This step is handled automatically by release-gkm.sh and cannot be run as a standalone script.
Parquet Output¶
The export produces 21 Parquet files — one per dictionary table. Statement and stream passthrough tables (variation, condition, conditionSet, scv, vcv, rcv, evidenceLine) have fully typed columns matching the BigQuery table schema. Key-value tables (sequenceReference, location, allele, gene, submitter, proposition, therapy, therapyGroup, etc.) export as two string columns (key, value).
| Parquet File | Content |
|---|---|
sequenceReference.parquet |
VRS sequence references |
location.parquet |
Genomic locations |
allele.parquet |
VRS alleles |
copyNumberCount.parquet |
Copy number count variants |
copyNumberChange.parquet |
Copy number change variants |
gene.parquet |
Gene MappableConcepts |
variation.parquet |
Categorical variants |
condition.parquet |
Conditions/traits |
conditionSet.parquet |
Condition sets |
submitter.parquet |
Submitters |
proposition.parquet |
SCV propositions |
vcv_proposition.parquet |
VCV propositions |
rcv_proposition.parquet |
RCV propositions |
evidenceLine.parquet |
SCV evidence lines |
vcv_evidenceLine.parquet |
VCV evidence lines |
rcv_evidenceLine.parquet |
RCV evidence lines |
scv.parquet |
SCV statements |
vcv.parquet |
VCV statements |
rcv.parquet |
RCV statements |
See Parquet Files for download URLs and query examples.
Step 4: Upload the Monthly Full to R2¶
upload-gkm-to-r2.sh uploads the assembled full bundle and Parquet files to Cloudflare R2. The full bundle is published once a month, aligned to ClinVar's own monthly VCV releases: when a new ClinVarVCVRelease_YYYY-MM.xml.gz appears at ClinVar's XML index, run-release.sh publishes our clinvar-gkm_YYYY-MM full from the most recent weekly release before ClinVar's monthly cut datetime. The --month-label=YYYY-MM flag sets the target month slot independently of the source release date (e.g. the 2026-07 full is built from our 2026-06-27 release, the last one before ClinVar's _2026-07 cut).
./src/scripts/upload-gkm-to-r2.sh <export_date> <dataset_version> <bundle_file> [--parquet-dir=DIR] [--dry-run]
# Upload full bundle + Parquet files
./src/scripts/upload-gkm-to-r2.sh 2026-06-14 v2_5_0 /tmp/clinvar-gkm-2026-06-14.json.gz \
--parquet-dir=/tmp/clinvar-gkm-2026-06-14-parquet
# Preview without uploading
./src/scripts/upload-gkm-to-r2.sh 2026-06-14 v2_5_0 /tmp/clinvar-gkm-2026-06-14.json.gz --dry-run
The script manages the monthly full slots:
datasets/— monthly full bundles for the current year (clinvar-gkm_yyyy-mm.json.gz) plus a stableclinvar-gkm_00-latest.json.gzdatasets/parquet/{yyyy-mm}/— dated per-section Parquet for each monthly full, plus a stabledatasets/parquet/00-latest/pointing at the newest monthly fullarchives/{yyyy}/— monthly full bundles and dated Parquet month sets (parquet/{yyyy-mm}/) from prior years
There is no weekly full bundle — weekly changes are published as deltas (see Step 5). The monthly upload is unconditional; boundary detection now governs only year rollover — when a new year begins, the prior year's monthly full bundles and dated Parquet month sets are moved to archives/{yyyy}/. After upload, generate-r2-index.sh regenerates index.json, which lists the monthly datasets (bundles + Parquet month sets), the archives, and the deltas.
Step 5: Publish the Weekly Delta¶
release-gkm-delta.sh publishes the per-release delta for every ClinVar release. It exports the delta_<dict> change tables produced by step 8 of the pipeline, assembles a delta bundle (added + updated records, same section structure as the full), merges per-section delta Parquet, builds manifest.json, and uploads the delta tree to R2.
The four internal steps are: export delta tables to GCS, assemble the delta bundle, merge delta Parquet, then build the manifest (build-delta-manifest.py) and upload (upload-gkm-delta-to-r2.sh). The uploader writes:
deltas/<yyyy-mmdd>/— the delta bundle (clinvar-gkm-delta_<yyyy-mmdd>.json.gz),manifest.json, andparquet/<section>.parquetdeltas/00-latest/— a server-side mirror of the most recent delta under stable filenames
build-delta-manifest.py derives each section's added / updated counts and deleted primary-key list from the dataset's gkm_change_log, and records baseline_release, compare_release, pipeline_version, and counts. The uploader then resolves checkpoint_full — the newest monthly full currently in datasets/ — so a consumer knows which full bundle the delta chain replays onto. See Downloads for the manifest shape and the consumer replay model.
R2 bucket CORS policy¶
The "Browse All Releases" file browser on the Downloads page does a cross-origin fetch() of index.json from the docs site. That only works if the public R2 bucket serves a CORS policy allowing GET from other origins — otherwise the browser blocks the read and the widget can't populate (it then shows a network/CORS message with a curl fallback).
The policy is checked in at src/scripts/r2-cors.json (GET/HEAD from * — appropriate for public, read-only data). It is not applied by the release scripts, because their object-scoped r2 upload token cannot read or write bucket configuration (PutBucketCors → AccessDenied). Apply it once with an Admin Read & Write R2 token:
# after configuring an admin token as an aws profile (default name: r2admin)
./src/scripts/apply-r2-cors.sh # or: ./src/scripts/apply-r2-cors.sh <profile>
Or paste r2-cors.json's rules into the Cloudflare dashboard: R2 → clinvar-gkm → Settings → CORS Policy. It only needs to be re-applied if the bucket is recreated or the policy is cleared.
Retired: gkm_json_proc¶
gkm_json_proc previously rendered the statement and dictionary tables into inlined JSON columns. Those render tables were never published — the export assembles the bundle directly from the gkm_dict_* tables — so the procedure has been retired from the hot path. The null/empty stripping it used to perform is now applied during assembly (assemble-gkm-dicts.py), matching the old remove_empty cleanup. gkm_json_proc is not a live pipeline step and not a downstream consumer of the dictionary tables.
Prerequisites¶
- Google Cloud SDK —
bqandgsutilcommands for BigQuery export and GCS operations - AWS CLI — configured with an
r2profile for Cloudflare R2 access - Python 3 — for the assembly script;
orjson(faster JSON) recommended - DuckDB CLI — for merging Parquet shards into single files per section
- BigQuery access — read access to the target dataset in
clingen-dev
Full Example¶
Publish the monthly full bundle for June 14, 2026:
This runs Steps 1–4: export to GCS, assemble JSON bundle, download and merge Parquet, upload the monthly full to R2.
Publish the weekly delta for the same release:
Steps 1 and 2 of the full can also be run individually. Steps 3–4 (Parquet download, shard merging, and upload) are handled internally by release-gkm.sh — use --start-step to resume from a specific step.
# 1. Export dictionary tables to GCS (NDJSON + Parquet)
./src/scripts/export-gkm-dicts.sh clinvar_2026_06_14_v2_5_0 clinvar-gkm gkm-dicts
# 2. Assemble NDJSON into a single JSON bundle
python3 ./src/scripts/assemble-gkm-dicts.py \
gs://clinvar-gkm/gkm-dicts/ \
2026-06-14
# 3-4. Download Parquet, merge shards, upload to R2
./src/scripts/release-gkm.sh 2026-06-14 v2_5_0 --start-step=3