Skip to content

Export & Distribute

The final pipeline step exports the gkm_dict_* tables from BigQuery, assembles them into a single keyed JSON bundle, exports Parquet files directly from BigQuery, and uploads to Cloudflare R2 for public distribution.

The gkm_dict_* tables are the published product — they are built directly by the statement procedures (steps 5–7) and the change-log step (step 8). The retired gkm_json_proc is not part of this path; its inlined JSON-render tables were never published. See Retired: gkm_json_proc.

Distribution follows a full + delta model:

  • The complete monthly full bundle (JSON + Parquet) is published once a month via release-gkm.sh / upload-gkm-to-r2.sh.
  • A weekly delta — added and updated records plus a change manifest — is published for every release via release-gkm-delta.sh / upload-gkm-delta-to-r2.sh.

Workflow

The export and distribution process uses four steps, executed in sequence. The release-gkm.sh wrapper runs all four automatically, or each step can be run individually.

Step 1: Export Dictionaries to GCS

export-gkm-dicts.sh exports all dictionary and statement tables from BigQuery to Google Cloud Storage in two formats:

  • NDJSON — sharded, gzip-compressed files for JSON bundle assembly (gkm-dicts/)
  • Parquet — Snappy-compressed files exported natively via bq extract (gkm-dicts-parquet/)
./src/scripts/export-gkm-dicts.sh <dataset> <gcs_bucket> [prefix] [--parquet-only]
# Export both NDJSON and Parquet
./src/scripts/export-gkm-dicts.sh clinvar_2026_06_14_v2_5_0 clinvar-gkm gkm-dicts

# Export Parquet only (skip NDJSON)
./src/scripts/export-gkm-dicts.sh clinvar_2026_06_14_v2_5_0 clinvar-gkm gkm-dicts --parquet-only

The script exports the following 19 tables:

Table NDJSON Output Parquet Output
gkm_dict_sequence_reference sequenceReference-*.ndjson.gz sequenceReference.parquet
gkm_dict_location location-*.ndjson.gz location.parquet
gkm_dict_allele allele-*.ndjson.gz allele.parquet
gkm_dict_copy_number_count copyNumberCount-*.ndjson.gz copyNumberCount.parquet
gkm_dict_copy_number_change copyNumberChange-*.ndjson.gz copyNumberChange.parquet
gkm_dict_gene gene-*.ndjson.gz gene.parquet
gkm_dict_variation variation-*.ndjson.gz variation.parquet
gkm_dict_condition condition-*.ndjson.gz condition.parquet
gkm_dict_condition_set conditionSet-*.ndjson.gz conditionSet.parquet
gkm_dict_submitter submitter-*.ndjson.gz submitter.parquet
gkm_dict_proposition proposition-*.ndjson.gz proposition.parquet
gkm_dict_evidence_line evidenceLine-*.ndjson.gz evidenceLine.parquet
gkm_dict_vcv_proposition vcv_proposition-*.ndjson.gz vcv_proposition.parquet
gkm_dict_vcv_evidence_line vcv_evidenceLine-*.ndjson.gz vcv_evidenceLine.parquet
gkm_dict_rcv_proposition rcv_proposition-*.ndjson.gz rcv_proposition.parquet
gkm_dict_rcv_evidence_line rcv_evidenceLine-*.ndjson.gz rcv_evidenceLine.parquet
gkm_dict_scv scv-*.ndjson.gz scv.parquet
gkm_dict_vcv vcv-*.ndjson.gz vcv.parquet
gkm_dict_rcv rcv-*.ndjson.gz rcv.parquet

BigQuery EXTRACT shards large NDJSON tables across multiple files automatically. Parquet files are exported as single files per table; BigQuery may auto-shard very large tables with numeric suffixes.

Step 2: Assemble Bundle

assemble-gkm-dicts.py reads all NDJSON shard files and assembles them into a single keyed JSON bundle file.

python3 ./src/scripts/assemble-gkm-dicts.py <source> <date> [--keep-source] [--copy-to-gcs]

<source> is a local path or gs:// URI containing the NDJSON shards. <date> is the ClinVar release date (YYYY-MM-DD); the output path is derived as /tmp/clinvar-gkm-{date}.json.gz. By default, source files are deleted after assembly; use --keep-source to retain them.

python3 ./src/scripts/assemble-gkm-dicts.py \
  gs://clinvar-gkm/gkm-dicts/ \
  2026-06-14

The script assembles 20 bundle sections in a fixed order: sequenceReference, location, allele, copyNumberCount, copyNumberChange, gene, variation, condition, conditionSet, therapy, therapyGroup, submitter, varcond-proposition, vartumor-proposition, vartherapy-proposition, varcustom-proposition, evidenceLine, scv, vcv, rcv. The therapy / therapyGroup sections hold content-addressed (deduplicated) drug therapies referenced by VariantTherapeuticResponseProposition.object via #/therapy/ and #/therapyGroup/. Each section is a keyed object where the key is the record's unique identifier. Propositions from SCV, VCV, and RCV are delivered in four datatype-homogeneous sections keyed by their (subject, object) signature — varcond-proposition (variant×condition), vartumor-proposition (variant×tumorType), vartherapy-proposition (variant×therapy), varcustom-proposition (custom variant×condition); evidence line shards are merged into a single evidenceLine section.

Install orjson for best performance:

pip install orjson

Step 3: Download and Merge Parquet from GCS

release-gkm.sh downloads Parquet shards from GCS, merges them into one file per section using DuckDB, and stages the merged files for upload. BigQuery exports may produce multiple shards per table (e.g., allele-000000000000.parquet, allele-000000000001.parquet); this step consolidates them into a single allele.parquet.

This step is handled automatically by release-gkm.sh and cannot be run as a standalone script.

Parquet Output

The export produces 21 Parquet files — one per dictionary table. Statement and stream passthrough tables (variation, condition, conditionSet, scv, vcv, rcv, evidenceLine) have fully typed columns matching the BigQuery table schema. Key-value tables (sequenceReference, location, allele, gene, submitter, proposition, therapy, therapyGroup, etc.) export as two string columns (key, value).

Parquet File Content
sequenceReference.parquet VRS sequence references
location.parquet Genomic locations
allele.parquet VRS alleles
copyNumberCount.parquet Copy number count variants
copyNumberChange.parquet Copy number change variants
gene.parquet Gene MappableConcepts
variation.parquet Categorical variants
condition.parquet Conditions/traits
conditionSet.parquet Condition sets
submitter.parquet Submitters
proposition.parquet SCV propositions
vcv_proposition.parquet VCV propositions
rcv_proposition.parquet RCV propositions
evidenceLine.parquet SCV evidence lines
vcv_evidenceLine.parquet VCV evidence lines
rcv_evidenceLine.parquet RCV evidence lines
scv.parquet SCV statements
vcv.parquet VCV statements
rcv.parquet RCV statements

See Parquet Files for download URLs and query examples.

Step 4: Upload the Monthly Full to R2

upload-gkm-to-r2.sh uploads the assembled full bundle and Parquet files to Cloudflare R2. The full bundle is published once a month, aligned to ClinVar's own monthly VCV releases: when a new ClinVarVCVRelease_YYYY-MM.xml.gz appears at ClinVar's XML index, run-release.sh publishes our clinvar-gkm_YYYY-MM full from the most recent weekly release before ClinVar's monthly cut datetime. The --month-label=YYYY-MM flag sets the target month slot independently of the source release date (e.g. the 2026-07 full is built from our 2026-06-27 release, the last one before ClinVar's _2026-07 cut).

./src/scripts/upload-gkm-to-r2.sh <export_date> <dataset_version> <bundle_file> [--parquet-dir=DIR] [--dry-run]
# Upload full bundle + Parquet files
./src/scripts/upload-gkm-to-r2.sh 2026-06-14 v2_5_0 /tmp/clinvar-gkm-2026-06-14.json.gz \
  --parquet-dir=/tmp/clinvar-gkm-2026-06-14-parquet

# Preview without uploading
./src/scripts/upload-gkm-to-r2.sh 2026-06-14 v2_5_0 /tmp/clinvar-gkm-2026-06-14.json.gz --dry-run

The script manages the monthly full slots:

  • datasets/ — monthly full bundles for the current year (clinvar-gkm_yyyy-mm.json.gz) plus a stable clinvar-gkm_00-latest.json.gz
  • datasets/parquet/{yyyy-mm}/ — dated per-section Parquet for each monthly full, plus a stable datasets/parquet/00-latest/ pointing at the newest monthly full
  • archives/{yyyy}/ — monthly full bundles and dated Parquet month sets (parquet/{yyyy-mm}/) from prior years

There is no weekly full bundle — weekly changes are published as deltas (see Step 5). The monthly upload is unconditional; boundary detection now governs only year rollover — when a new year begins, the prior year's monthly full bundles and dated Parquet month sets are moved to archives/{yyyy}/. After upload, generate-r2-index.sh regenerates index.json, which lists the monthly datasets (bundles + Parquet month sets), the archives, and the deltas.

Step 5: Publish the Weekly Delta

release-gkm-delta.sh publishes the per-release delta for every ClinVar release. It exports the delta_<dict> change tables produced by step 8 of the pipeline, assembles a delta bundle (added + updated records, same section structure as the full), merges per-section delta Parquet, builds manifest.json, and uploads the delta tree to R2.

./src/scripts/release-gkm-delta.sh <export_date> <dataset_version> [--start-step=N] [--dry-run]
# Publish the weekly delta for a release
./src/scripts/release-gkm-delta.sh 2026-07-06 v2_5_0

The four internal steps are: export delta tables to GCS, assemble the delta bundle, merge delta Parquet, then build the manifest (build-delta-manifest.py) and upload (upload-gkm-delta-to-r2.sh). The uploader writes:

  • deltas/<yyyy-mmdd>/ — the delta bundle (clinvar-gkm-delta_<yyyy-mmdd>.json.gz), manifest.json, and parquet/<section>.parquet
  • deltas/00-latest/ — a server-side mirror of the most recent delta under stable filenames

build-delta-manifest.py derives each section's added / updated counts and deleted primary-key list from the dataset's gkm_change_log, and records baseline_release, compare_release, pipeline_version, and counts. The uploader then resolves checkpoint_full — the newest monthly full currently in datasets/ — so a consumer knows which full bundle the delta chain replays onto. See Downloads for the manifest shape and the consumer replay model.


R2 bucket CORS policy

The "Browse All Releases" file browser on the Downloads page does a cross-origin fetch() of index.json from the docs site. That only works if the public R2 bucket serves a CORS policy allowing GET from other origins — otherwise the browser blocks the read and the widget can't populate (it then shows a network/CORS message with a curl fallback).

The policy is checked in at src/scripts/r2-cors.json (GET/HEAD from * — appropriate for public, read-only data). It is not applied by the release scripts, because their object-scoped r2 upload token cannot read or write bucket configuration (PutBucketCors → AccessDenied). Apply it once with an Admin Read & Write R2 token:

# after configuring an admin token as an aws profile (default name: r2admin)
./src/scripts/apply-r2-cors.sh            # or: ./src/scripts/apply-r2-cors.sh <profile>

Or paste r2-cors.json's rules into the Cloudflare dashboard: R2 → clinvar-gkm → Settings → CORS Policy. It only needs to be re-applied if the bucket is recreated or the policy is cleared.


Retired: gkm_json_proc

gkm_json_proc previously rendered the statement and dictionary tables into inlined JSON columns. Those render tables were never published — the export assembles the bundle directly from the gkm_dict_* tables — so the procedure has been retired from the hot path. The null/empty stripping it used to perform is now applied during assembly (assemble-gkm-dicts.py), matching the old remove_empty cleanup. gkm_json_proc is not a live pipeline step and not a downstream consumer of the dictionary tables.


Prerequisites

  • Google Cloud SDK — bq and gsutil commands for BigQuery export and GCS operations
  • AWS CLI — configured with an r2 profile for Cloudflare R2 access
  • Python 3 — for the assembly script; orjson (faster JSON) recommended
  • DuckDB CLI — for merging Parquet shards into single files per section
  • BigQuery access — read access to the target dataset in clingen-dev

Full Example

Publish the monthly full bundle for June 14, 2026:

./src/scripts/release-gkm.sh 2026-06-14 v2_5_0

This runs Steps 1–4: export to GCS, assemble JSON bundle, download and merge Parquet, upload the monthly full to R2.

Publish the weekly delta for the same release:

./src/scripts/release-gkm-delta.sh 2026-06-14 v2_5_0

Steps 1 and 2 of the full can also be run individually. Steps 3–4 (Parquet download, shard merging, and upload) are handled internally by release-gkm.sh — use --start-step to resume from a specific step.

# 1. Export dictionary tables to GCS (NDJSON + Parquet)
./src/scripts/export-gkm-dicts.sh clinvar_2026_06_14_v2_5_0 clinvar-gkm gkm-dicts

# 2. Assemble NDJSON into a single JSON bundle
python3 ./src/scripts/assemble-gkm-dicts.py \
  gs://clinvar-gkm/gkm-dicts/ \
  2026-06-14

# 3-4. Download Parquet, merge shards, upload to R2
./src/scripts/release-gkm.sh 2026-06-14 v2_5_0 --start-step=3