Skip to content

Pipeline Overview

The ClinVar-GKM pipeline transforms ClinVar XML release data into the GKM (Genomic Knowledge Model) schema set (VRS, Cat-VRS, VA-Spec) through a series of BigQuery stored procedures with an external VRS Python processing step.

The two most expensive stages — Variation Identity and VRS Processing — support incremental processing: they recompute only the variations that changed since the prior release and carry the rest forward, driven by the release-to-release diff (dataset_diff_on). The remaining stored procedures currently run as full rebuilds each release.

Pipeline Steps

The pipeline executes in the following order. Each step is a BigQuery stored procedure unless otherwise noted.

┌──────────────────────────────┐
│ 1. Variation Identity        │  variation_identity[_incremental]
│    Extract & normalize       │  → variation_identity table
│    variant data              │
└──────────────┬───────────────┘
               │
┌──────────────▼───────────────┐
│ 2. VRS Processing            │  External: vrs-python
│    Export → VRS Python →     │  → gkm_vrs table
│    Import back to BigQuery   │
└──────────────┬───────────────┘
               │
┌──────────────▼───────────────┐
│ 3. Cat-VRS Generation        │  gkm_catvar_proc
│    Canonical alleles &       │  → gkm_catvar table
│    categorical variants      │
└──────────────┬───────────────┘
               │
┌──────────────▼───────────────┐
│ 4. Conditions & Traits       │  gkm_scv_condition_proc
│    Map traits, build         │  → condition mapping &
│    conditions & condition    │    condition set tables
│    sets                      │
└──────────────┬───────────────┘
               │
┌──────────────▼───────────────┐
│ 5. SCV Statements            │  gkm_scv_statement_proc
│    Build SCV records,        │  → gkm_dict_scv table
│    propositions & statements │
└──────────────┬───────────────┘
               │
┌──────────────▼───────────────┐
│ 6. VCV Statements            │  gkm_vcv_proc +
│    Aggregate SCVs into       │  gkm_vcv_statement_proc
│    variant-level statements  │  → gkm_dict_vcv table
└──────────────┬───────────────┘
               │
┌──────────────▼───────────────┐
│ 7. RCV Statements            │  gkm_rcv_proc +
│    Aggregate SCVs into       │  gkm_rcv_statement_proc
│    condition-level statements│  → gkm_dict_rcv table
└──────────────┬───────────────┘
               │
┌──────────────▼───────────────┐
│ 8. Change log + deltas       │  gkm_change_log +
│    A/U/D per dict + delta     │  gkm_delta_build
│    payloads for publishing    │  → gkm_change_log, delta_<dict>
└──────────────┬───────────────┘
               │
┌──────────────▼───────────────┐
│ 9. Export & Distribute       │  export-gkm-dicts.sh
│    Export NDJSON + Parquet    │  assemble-gkm-dicts.py
│    to GCS, assemble JSON     │  release-gkm.sh
│    bundle, upload to R2      │  → R2 public bucket
└──────────────────────────────┘

Running the Pipeline

Single-command run

The whole release can be run end-to-end with the run-release.sh orchestrator, which chains five stages: variation identity, the GCS export, vrsification, the transform/load/procedures step (vrs-to-bq-table.sh), and the export/publish step (release-gkm.sh):

./src/scripts/run-release.sh YYYY-MM-DD              # incremental (default)
./src/scripts/run-release.sh YYYY-MM-DD --full       # full rebuild / reseed
./src/scripts/run-release.sh YYYY-MM-DD --dry-run    # run everything but skip the R2 publish (stage 5)
./src/scripts/run-release.sh YYYY-MM-DD --start-step 5   # resume from a stage (1-5)

--full propagates version-invalidation across stages 1–2; use it for the first release or after a variation_identity transform change or a vrsify-pin bump. --dry-run runs the build stages but has the final release stage only print what it would upload — use it for test runs. The vrsify stage (stage 3) requires local SeqRepo / UTA / gene-normalizer services — see src/vrsify/README.md — so on hosts without them, run the BigQuery-side stages with --start-step and run vrsify separately.

The individual stages are documented below.

Step 1: Variation Identity

From the BigQuery console — incremental by default, full rebuild when reseeding (see Incremental Rebuild):

-- default: recompute only changed variations, carry the rest forward
CALL `clinvar_ingest.variation_identity_incremental`(CURRENT_DATE(), FALSE);

-- full rebuild: first release, or after a variation_identity transform change
CALL `clinvar_ingest.variation_identity`(CURRENT_DATE(), FALSE);

Step 2: VRS Processing

Export, process externally with vrs-python, and load back — incremental by default (only changed variations are vrsified; unchanged gkm_vrs results carry forward). See VRS Processing.

Step 3: Cat-VRS through JSON Output

From the BigQuery console:

CALL `clinvar_ingest.gkm_scv_condition_proc`(CURRENT_DATE(), FALSE);
CALL `clinvar_ingest.gkm_scv_statement_proc`(CURRENT_DATE(), FALSE);
CALL `clinvar_ingest.gkm_vcv_proc`(CURRENT_DATE(), FALSE);
CALL `clinvar_ingest.gkm_vcv_statement_proc`(CURRENT_DATE(), FALSE);
CALL `clinvar_ingest.gkm_rcv_proc`(CURRENT_DATE(), FALSE);
CALL `clinvar_ingest.gkm_rcv_statement_proc`(CURRENT_DATE(), FALSE);
-- NOTE: gkm_json_proc is retired (Plan 4) — its JSON-render tables were unpublished;
-- the gkm_dict_* tables are the published product and are built by the procs above.

Step 4: Export & Distribute

The gkm_dict_* tables are the published product. The monthly full bundle and the weekly delta are published by separate scripts:

# Monthly full bundle (export, assemble, download Parquet, upload to R2)
./src/scripts/release-gkm.sh 2026-06-14 v2_5_0

# Weekly delta (added + updated records + change manifest)
./src/scripts/release-gkm-delta.sh 2026-06-14 v2_5_0

Steps 1 and 2 can also be run individually. Steps 3–4 (Parquet download, shard merging, and upload) are handled internally by release-gkm.sh — use --start-step to resume from a specific step.

# 1. Export dictionary tables to GCS (NDJSON + Parquet)
./src/scripts/export-gkm-dicts.sh clinvar_2026_06_14_v2_5_0 clinvar-gkm gkm-dicts

# 2. Assemble NDJSON into JSON bundle
python3 ./src/scripts/assemble-gkm-dicts.py gs://clinvar-gkm/gkm-dicts/ 2026-06-14

# 3-4. Download Parquet, merge shards, upload to R2
./src/scripts/release-gkm.sh 2026-06-14 v2_5_0 --start-step=3

See Export for details on each step.

Documentation Tracks

The pipeline documentation serves two audiences:

  • Pipeline (this section) — documents how data flows through BigQuery stored procedures, including internal table schemas, transformation logic, and step-by-step workflows. Each step is tagged as Pipeline table, JSON artifact, or Internal to indicate its role
  • Output Reference — documents the JSON output files from a consumer perspective, covering record structure, field meanings, and usage guidance

Detailed Documentation

Each pipeline step has its own documentation page: