Pipeline Overview¶
The ClinVar-GKM pipeline transforms ClinVar XML release data into the GKM (Genomic Knowledge Model) schema set (VRS, Cat-VRS, VA-Spec) through a series of BigQuery stored procedures with an external VRS Python processing step.
The two most expensive stages — Variation Identity and VRS Processing — support incremental processing: they recompute only the variations that changed since the prior release and carry the rest forward, driven by the release-to-release diff (dataset_diff_on). The remaining stored procedures currently run as full rebuilds each release.
Pipeline Steps¶
The pipeline executes in the following order. Each step is a BigQuery stored procedure unless otherwise noted.
┌──────────────────────────────┐
│ 1. Variation Identity │ variation_identity[_incremental]
│ Extract & normalize │ → variation_identity table
│ variant data │
└──────────────┬───────────────┘
│
┌──────────────▼───────────────┐
│ 2. VRS Processing │ External: vrs-python
│ Export → VRS Python → │ → gkm_vrs table
│ Import back to BigQuery │
└──────────────┬───────────────┘
│
┌──────────────▼───────────────┐
│ 3. Cat-VRS Generation │ gkm_catvar_proc
│ Canonical alleles & │ → gkm_catvar table
│ categorical variants │
└──────────────┬───────────────┘
│
┌──────────────▼───────────────┐
│ 4. Conditions & Traits │ gkm_scv_condition_proc
│ Map traits, build │ → condition mapping &
│ conditions & condition │ condition set tables
│ sets │
└──────────────┬───────────────┘
│
┌──────────────▼───────────────┐
│ 5. SCV Statements │ gkm_scv_statement_proc
│ Build SCV records, │ → gkm_dict_scv table
│ propositions & statements │
└──────────────┬───────────────┘
│
┌──────────────▼───────────────┐
│ 6. VCV Statements │ gkm_vcv_proc +
│ Aggregate SCVs into │ gkm_vcv_statement_proc
│ variant-level statements │ → gkm_dict_vcv table
└──────────────┬───────────────┘
│
┌──────────────▼───────────────┐
│ 7. RCV Statements │ gkm_rcv_proc +
│ Aggregate SCVs into │ gkm_rcv_statement_proc
│ condition-level statements│ → gkm_dict_rcv table
└──────────────┬───────────────┘
│
┌──────────────▼───────────────┐
│ 8. Change log + deltas │ gkm_change_log +
│ A/U/D per dict + delta │ gkm_delta_build
│ payloads for publishing │ → gkm_change_log, delta_<dict>
└──────────────┬───────────────┘
│
┌──────────────▼───────────────┐
│ 9. Export & Distribute │ export-gkm-dicts.sh
│ Export NDJSON + Parquet │ assemble-gkm-dicts.py
│ to GCS, assemble JSON │ release-gkm.sh
│ bundle, upload to R2 │ → R2 public bucket
└──────────────────────────────┘
Running the Pipeline¶
Single-command run¶
The whole release can be run end-to-end with the run-release.sh orchestrator, which chains five stages: variation identity, the GCS export, vrsification, the transform/load/procedures step (vrs-to-bq-table.sh), and the export/publish step (release-gkm.sh):
./src/scripts/run-release.sh YYYY-MM-DD # incremental (default)
./src/scripts/run-release.sh YYYY-MM-DD --full # full rebuild / reseed
./src/scripts/run-release.sh YYYY-MM-DD --dry-run # run everything but skip the R2 publish (stage 5)
./src/scripts/run-release.sh YYYY-MM-DD --start-step 5 # resume from a stage (1-5)
--full propagates version-invalidation across stages 1–2; use it for the first release or after a variation_identity transform change or a vrsify-pin bump. --dry-run runs the build stages but has the final release stage only print what it would upload — use it for test runs. The vrsify stage (stage 3) requires local SeqRepo / UTA / gene-normalizer services — see src/vrsify/README.md — so on hosts without them, run the BigQuery-side stages with --start-step and run vrsify separately.
The individual stages are documented below.
Step 1: Variation Identity¶
From the BigQuery console — incremental by default, full rebuild when reseeding (see Incremental Rebuild):
-- default: recompute only changed variations, carry the rest forward
CALL `clinvar_ingest.variation_identity_incremental`(CURRENT_DATE(), FALSE);
-- full rebuild: first release, or after a variation_identity transform change
CALL `clinvar_ingest.variation_identity`(CURRENT_DATE(), FALSE);
Step 2: VRS Processing¶
Export, process externally with vrs-python, and load back — incremental by default (only changed variations are vrsified; unchanged gkm_vrs results carry forward). See VRS Processing.
Step 3: Cat-VRS through JSON Output¶
From the BigQuery console:
CALL `clinvar_ingest.gkm_scv_condition_proc`(CURRENT_DATE(), FALSE);
CALL `clinvar_ingest.gkm_scv_statement_proc`(CURRENT_DATE(), FALSE);
CALL `clinvar_ingest.gkm_vcv_proc`(CURRENT_DATE(), FALSE);
CALL `clinvar_ingest.gkm_vcv_statement_proc`(CURRENT_DATE(), FALSE);
CALL `clinvar_ingest.gkm_rcv_proc`(CURRENT_DATE(), FALSE);
CALL `clinvar_ingest.gkm_rcv_statement_proc`(CURRENT_DATE(), FALSE);
-- NOTE: gkm_json_proc is retired (Plan 4) — its JSON-render tables were unpublished;
-- the gkm_dict_* tables are the published product and are built by the procs above.
Step 4: Export & Distribute¶
The gkm_dict_* tables are the published product. The monthly full bundle and the weekly delta are published by separate scripts:
# Monthly full bundle (export, assemble, download Parquet, upload to R2)
./src/scripts/release-gkm.sh 2026-06-14 v2_5_0
# Weekly delta (added + updated records + change manifest)
./src/scripts/release-gkm-delta.sh 2026-06-14 v2_5_0
Steps 1 and 2 can also be run individually. Steps 3–4 (Parquet download, shard merging, and upload) are handled internally by release-gkm.sh — use --start-step to resume from a specific step.
# 1. Export dictionary tables to GCS (NDJSON + Parquet)
./src/scripts/export-gkm-dicts.sh clinvar_2026_06_14_v2_5_0 clinvar-gkm gkm-dicts
# 2. Assemble NDJSON into JSON bundle
python3 ./src/scripts/assemble-gkm-dicts.py gs://clinvar-gkm/gkm-dicts/ 2026-06-14
# 3-4. Download Parquet, merge shards, upload to R2
./src/scripts/release-gkm.sh 2026-06-14 v2_5_0 --start-step=3
See Export for details on each step.
Documentation Tracks¶
The pipeline documentation serves two audiences:
- Pipeline (this section) — documents how data flows through BigQuery stored procedures, including internal table schemas, transformation logic, and step-by-step workflows. Each step is tagged as Pipeline table, JSON artifact, or Internal to indicate its role
- Output Reference — documents the JSON output files from a consumer perspective, covering record structure, field meanings, and usage guidance
Detailed Documentation¶
Each pipeline step has its own documentation page:
- Variation Identity — variant extraction, normalization, VRS class assignment
- VRS Processing — external VRS Python step
- Cat-VRS — categorical variant generation
- Conditions & Traits — condition mapping, traits, condition sets
- SCV Statements — SCV records, propositions, final statements
- VCV Statements — aggregate variant-level VCV statements
- RCV Statements — aggregate condition-level RCV statements
- Export & Distribute — export to GCS, assemble bundle, upload to R2