GA4GH GKS / GKM survey for just-dna-format — 2026-09-28¶
Scope: finding, not deciding. Every claim carries a tag: [M] measured (command run in this session), [D] read in primary docs/source, [I] inferred. Package weights were measured in throwaway venvs in the surveying session's scratchpad, which is not preserved; nothing in the repo was edited during the survey.
Naming first. The workstream rebranded in 2026: the schema set is now GKM (Genomic Knowledge
Models), gks-core is gkm-core, the metaschema processor is ga4gh.gkm.metaschema; "GKS" stays
the name of the GA4GH workstream [D: ga4gh/vrs PR #708 "Rebrand GKS → GKM" in the 2.1.0 release
notes; gh api repos/ga4gh/gks-core resolves to gkm-core; clinvar-gkm PR #89/#92].
0. Headline findings (read these if nothing else)¶
- Our core grain has no GKS counterpart.
VariantRowis keyed(variant_key, genotype);PharmVariantRowis per genotype;DiplotypeRowis per diplotype. VRS 2.x has no Genotype (the 1.3Genotype/Haplotype/VariationSetclasses are absent from the 2.x class list;Haplotypebecame the cis-onlyCisPhasedBlock; in the 2.1.0 docsGenotypeappears only in the 1.3 release notes) and no VA-Spec schema carries a zygosity or genotype field (pathogenicity qualifiers aregeneContext,alleleOrigin,penetrance,modeOfInheritance) [M:gh api repos/ga4gh/vrs/git/trees/2.1.0json class list vsvrs-source.yaml@1.3.0; grep of the 2.1.0 docs; D: VRS 2.0 release notes "renames Haplotype to CisPhasedBlock"; M: grep forzygos|genotype|allelicStateover every VA-Spec JSON schema at 1.0.1 (30 files) and at the 1.1.0 ballot (39 files) → zero hits]. Any GKS export of our annotations is lossy at the genotype axis. - VRS 2.1.0 (released 2026-09-01) changed normalization, and vrs-python has not caught up. The
2.1 minor revises ambiguous-insertion normalization to pick the smallest repeat-unit factor
(was greatest), which changes the digest of affected trial-use Alleles [D:
https://vrs.ga4gh.org/en/2.1.1-ballot.2026-09/releases/2.1.html]. Both
ga4gh.vrs2.3.3 (our pin) and 2.4.0a4 (latest alpha, "update models to … vrs 2.1.0") still iterate_factor_gendescending, first valid cycle wins [M: readga4gh/vrs/normalize.pyin both venvs]; the change is open as vrs-python issue #637 "Add support for switching RLE subunit between largest and smallest" [M:gh api search/issues]. Everyone we compare against today (our enricher, ClinVar-GKM'svrs_output_2_0_1.schema.json, gnomAD perdocs/SCHEMAS.mdground-truth test) is on 2.0.x digests. 270,161 of ClinVar-GKM's 4,460,884 alleles areReferenceLengthExpression— the class where a subset (insertions with a composite repeat length) would move [M: duckdb overallele.parquet; I: which subset exactly moves was not computed]. RM270's candidate 4 (carryvrs_idonto parquets) inherits this pending id move. - A
ga4gh:VA.string does not say which VRS it is. Forchr7:140753336 A>T(BRAF V600E) three ids exist: our/VRS 2.0ga4gh:VA.Otc5ovrw906Ack087o1fhegB4jDRqCAe, vrs-python's VRS 1.3 back-computationga4gh:VA.fZiBjQEolbkL0AxjoTZf4SOkFy9J0ebU, and the ClinGen Allele Registry's/vrAlleleanswerga4gh:VA.HaPTmn-rrjRoZnIVw1I4AZPa6YHa2ojh(a VRS 1.x-era shape:SimpleInterval,SequenceState,sequence_id; the exact minor is not established) [M: all three computed/fetched this session]. CAR's VRS ids are not joinable to ours nor to a 1.3 back-computation; the CAID is the only hop to CAR. Nothing in the repo records the VRS version behind avrs_id[M:grep -rn -i 'vrs_version\|VRS_VERSION'overschema/,enricher/,docs/SCHEMAS.md→ no hits]. - ClinVar-GKM's
allele.parquetis a 1.09 GB, CC0, offline VRS oracle — 4.46 M alleles, each withspdi,hgvs.gandgnomad(VCF-style) expressions beside its VRS 2.0 id. It contains exactly the two RLE ids RM270 minted forrs72613567(ga4gh:VA.Jml7…,4-87310240-T-TA) andrs77944059(ga4gh:VA.ks2G…,2-166204470-GAAAC-G), and our stdlib SNV minter matches its ids byte-for-byte on the two rows sampled [M]. Its value to us is as an independent placement witness for RM267/RM270/RM292, not as a replacement for the ClinVar VCF lane. - The stable Python packages embed last-generation specs.
ga4gh.vrs 2.3.3 → VRS 2.0.1,ga4gh.cat_vrs 0.7.2 → Cat-VRS 1.0.0,ga4gh.va_spec 0.4.4 → VA-Spec 1.0.1[M:*_VERSIONconstants in the installed packages]. The newer line is alpha-only on PyPI:ga4gh.vrs 2.4.0a4 → VRS 2.1.0,ga4gh.cat_vrs 0.8.0a4 → Cat-VRS 1.1.0,ga4gh.va_spec 0.5.0a5 → VA-Spec 1.1.0-snapshot.2026-06.1(not yet the 2026-09 ballot) [M: same constants in a fourth venv]. Adding the Cat-VRS + VA-Spec models on top of what the enricher already carries costs 2 packages, ~144 KB [M]. - VA-Spec 1.1 (ballot 2026-09) is backwards-incompatible at trial use:
EvidenceLinefolded intoStatement,Therapeutic→Therapy,ClinicalVariantProposition→GeneticContextVariantProposition,CohortAlleleFrequencyStudyResult.focusAlleleFrequency→alleleFrequencyetc. [D: va-spec release1.1.0-ballot.2026-09.1notes]. Anything mapped to 1.0.1 now is remapped at 1.1. - Boolean composition of categorical variants is an open, unmerged draft (Cat-VRS PR #248
CompositeCategoricalVariant+CategoricalVariantCriterion, AND/OR +presence: present|absent, no phase). Cat-VRSCategoricalVarianthas no computed digest — ids are system-assigned (clinvar:NNN,civic.mpid:NNN). See §6 for RM28. - Exon coordinates for AlphaGenome need no UTA and no SeqRepo. The MANE Ensembl GTF (8.6 MB gz)
carries 204,815 exons over 19,437 MANE transcripts [M]; the repo already builds a MANE lane.
cool-seq-tool's exon mapper needs UTA (Postgres, CC BY-SA 4.0) on every path [M: read
exon_genomic_coords.py]. See §5.
1. Specifications¶
| Spec | Latest stable | Latest pre-release | Licence | What is new / relevant | Source |
|---|---|---|---|---|---|
| VRS | 2.1.0 (2026-09-01); 2.0.1 (2025-03-20) | 2.1.1-ballot.2026-09.1 (2026-09-16) | Apache-2.0 | 2.1 minor: smallest-factor RLE normalization (digest-moving). 2.1 patch: RelativeAllele, RelativeSequenceLocation, SequenceOffsetLocation (draft, intronic positions). 2.1.1: metaschema rename, abstract classes sealed; no data-model change. There is no VRS 2.2 — the task's "2.1/2.2" premise does not hold [M: release list]. Adjacency/copy number/derivative molecules came with 2.0.0 ("supports structural variation"), not 2.1 |
[M] gh api repos/ga4gh/vrs/releases; [D] https://vrs.ga4gh.org/en/2.1.1-ballot.2026-09/releases/2.1.html |
| VRS class maturity @2.1.0 | Allele, CisPhasedBlock, Adjacency, CopyNumberCount = trial use; CopyNumberChange, DerivativeMolecule, Terminus, RelativeAllele = draft. Digest prefixes: VA, RA, CPB, CN, CX, SL, RSL |
No Genotype, Haplotype, VariationSet in 2.x (present in 1.3 source) |
[M] schema/vrs/json/*@2.1.0 maturity field; vrs-source.yaml grep for prefix: |
||
| Cat-VRS | 1.1.0 (2026-09-01); 1.0.0 (2025-06-11) | 1.1.1-ballot.2026-09.1 | Apache-2.0 | Constraint subtypes (exact names @1.1.0): DefiningAlleleConstraint (trial use), DefiningLocationConstraint (trial use), CopyCountConstraint (trial use), CopyChangeConstraint (draft), FeatureContextConstraint (draft), FunctionConstraint (draft), AdjacencyConstraint (draft). Recipes: CanonicalAllele, ProteinSequenceConsequence (trial use); CategoricalCnv, FunctionVariant, GeneFusion (draft). CategoricalCnv = exactly one DefiningLocationConstraint + one of CopyChange/CopyCount |
[M] gh api repos/ga4gh/cat-vrs/contents/schema/cat-vrs/json/*?ref=1.1.0 |
| VA-Spec | 1.0.1 (2025-07-23) | 1.1.0-ballot.2026-09.1 (2026-09-16) | Apache-2.0 | See profile table below | [M] releases; tree listings at both tags |
| gkm-core (ex gks-core) | 1.2.0 (2026-08-31) | 1.3.0-ballot.2026-09.1 | Apache-2.0 | MappableConcept, ConceptSet, Coding, ConceptMapping, Extension, Entity/Element. 1.3 adds optional ConceptSet.conceptSetType |
[M] releases |
| refget Sequences | v2.0 (approved 2024) | — | — | Our REFGET_GRCh38 constants are refget accessions |
[D] https://ga4gh.github.io/refget/ |
| refget Sequence Collections (seqcol) | v1.0 (approved 2025) | — | — | A content digest for a whole assembly — the only GKS-adjacent thing that could give genome_build a digest rather than a label [I] |
[D] same page |
| VCF VRS annotation | vrs-annotate CLI ships in ga4gh.vrs[extras] (ga4gh/vrs/extras/annotator/vcf.py) |
Apache-2.0 | Needs the [extras] tree (46 pkgs, 172 MB, local SeqRepo) |
[M] installed venv | |
| Phenopackets ↔ VRS 2 | not established | Not checked this session | — |
VA-Spec profiles (what exists, at which maturity)¶
| Profile / class | 1.0.1 | 1.1.0 ballot | Our nearest table |
|---|---|---|---|
VariantPathogenicityProposition (+ ACMG-2015 statement/evidence-line profile) |
trial use | trial use (renamed fields) | variants.csv clin_sig/pathogenic; clinical_assertions.csv |
VariantOncogenicityProposition (+ CCV-2022) |
yes | yes | none (no somatic scope) |
VariantTherapeuticResponseProposition (+ AAC-2017) |
yes | trial use — "response of a neoplasm" | pharm_variants.csv only loosely: germline PGx is out of its stated scope |
VariantDiagnosticProposition, VariantPrognosticProposition |
yes | yes | none |
CohortAlleleFrequencyStudyResult |
yes | trial use; focusCount, locusCount, alleleFrequency, cohort, subCohortFrequency |
frequencies.csv |
ExperimentalVariantFunctionalImpact{Proposition,StudyResult} (the AVE/MAVE profile) |
yes | trial use; functionalImpactScore |
none (no assay table) |
TumorVariantFrequencyStudyResult |
yes | yes | none |
ComputationalVariantFunctionalImpactAnalysisResult |
— | draft (new): focus, transcriptVariationContext, impactScore, impactScoreType, categoricalImpact, impactedFeatureType, impactedFeature, sourceDataSet |
RM23 predictions, AlphaGenome AVI |
VariantMolecularConsequenceProposition |
— | draft (new), transcriptVariationContextQualifier, proteinVariationContextQualifier, geneContextQualifier |
RM236 consequence |
GeneDiseaseValidityProposition |
— | draft (new), modeOfInheritanceQualifier |
gene_validity.csv |
VariantClinicalSignificanceProposition |
— | draft (new) | clin_sig |
DataItem |
— | draft (new) | — |
| PGx / star allele / diplotype profile | none | none | haplotypes.csv, diplotypes.csv, allele_function.csv, activity_phenotype.csv |
| GWAS / polygenic profile | none | none | gwas_effects.csv, pgs.csv |
[M] gh api repos/ga4gh/va-spec/git/trees/<tag>?recursive=1 and per-class maturity/properties;
[M] va-spec issue search for "pharmacogen" returned nothing; "GWAS"/"polygenic" hits are old open
scope issues (#25, #3, #246) with no profile.
2. Implementations and services — measured weight¶
Installed each into its own uv venv (Python 3.12) under scratchpad/venvs/. "Pkgs" = uv pip list
count; size = du -sm site-packages. [extras]-pulling packages fail to build on this host because
hgvs → psycopg2 needs pg_config; re-installed with psycopg2 overridden to psycopg2-binary
(that failure is itself a cost: a system libpq-dev, or an override) [M].
| Package | PyPI stable (latest pre) | Spec embedded | Licence | Pkgs | Size | Needs at runtime | Notes |
|---|---|---|---|---|---|---|---|
ga4gh.vrs (core) — already an enricher dep |
2.3.3 (2.4.0a4) | VRS 2.0.1 | Apache-2.0 | 14 | 11 MB | seqrepo REST for indels (as the enricher does) | pydantic, bioutils, requests, canonicaljson |
ga4gh.vrs[extras] |
2.3.3 | 2.0.1 | Apache-2.0 | 46 | 172 MB | local SeqRepo (13.5 GB), UTA for HGVS | pysam, hgvs, psycopg2, ipython |
ga4gh.cat_vrs |
0.7.2 (0.8.0a4) | Cat-VRS 1.0.0 | Apache-2.0 | 15 | 11 MB | nothing | +1 pkg, 28 KB over core vrs |
ga4gh.va_spec |
0.4.4 (0.5.0a5) | VA-Spec 1.0.1 | Apache-2.0 | 16 | 12 MB | nothing | +1 pkg, 116 KB over cat_vrs; pins cat_vrs~=0.7.1 |
cool-seq-tool |
0.17.0 | — | MIT | 53 | 309 MB | UTA Postgres + local SeqRepo; MANE files via wags-tails (auto-download from NCBI) |
polars (172 MB runtime), boto3, psycopg, pysam, ipython; bundles a 24 MB transcript_mapping.tsv (BioMart-style Ensembl gene/transcript/protein + MANE RefSeq columns [I from its header]) |
gene-normalizer |
0.11.5 | — | MIT | 30 | 46 MB | DynamoDB or Postgres loaded by its ETL | fastapi, uvicorn, boto3 |
variation-normalizer |
0.16.0 | — | MIT | 68 | 421 MB | UTA + SeqRepo + gene-normalizer DB | pulls ga4gh.vrs[extras]; accepts HGVS, "BRAF V600E", gnomAD-VCF strings |
anyvar |
1.0.0 (1.1.0.dev0) | — | Apache-2.0 | 80 | 259 MB | SeqRepo; UTA for HGVS; a SQL store | snowflake-sqlalchemy, sqlalchemy, fastapi, pysam, hgvs. Registration/lookup server around the same ga4gh.vrs normalizer |
metakb |
1.1.0 on PyPI (stale; repo at 2.0.0-a0) | — | MIT | not installed | — | neo4j + all four normalizers + UTA + SeqRepo | server pyproject pins ga4gh.vrs[extras]~=2.4.0-a1, cat_vrs~=0.8.0-a2, va_spec==0.5.0-a0 |
| disease-normalizer / thera-py | 0.13.1 / 0.13.0 | — | MIT | not installed | — | DynamoDB/Postgres | same FastAPI + boto3 shape [D: PyPI requires_dist] |
biocommons.seqrepo |
0.6.11 | — | Apache-2.0 | (inside cool-seq-tool) | — | snapshot 13.5 GB (2024-12-20), ~8 GB first snapshot per README | rsync required [D: seqrepo README] |
| UTA | uta_20241220 |
— | CC BY-SA 4.0 (per meta table row shown in README) |
— | Docker image 174 MB compressed [M: Docker Hub API] | Postgres; dump uta_20241220.pgd.gz |
dump size not established: dl.biocommons.org answers a human-verification page to curl [M] — an automation hazard for a builder |
Sources: [M] curl https://pypi.org/pypi/<pkg>/json, gh api repos/.../releases, the venvs; [D]
github.com/biocommons/uta README; github.com/biocommons/biocommons.seqrepo README.
cool-seq-tool: what the exon mapping actually needs [M: read installed source]¶
| Capability | Module | Local requirement |
|---|---|---|
genomic ↔ transcript segment / exon number (genomic_to_tx_segment, tx_segment_to_genomic) |
mappers/exon_genomic_coords.py |
UTA on every path (self.uta_db.repository() at ~10 call sites) + SeqRepo (chromosome_to_acs, translate_identifier) |
| MANE transcript selection | mappers/mane_transcript.py, sources/mane_transcript_mappings.py |
MANE summary file (MANE_SUMMARY_PATH or wags-tails download) + UTA |
| gene / CDS overlap for a region | mappers/feature_overlap.py |
MANE RefSeq genomic GFF + polars only; SeqRepo only to translate an identifier when the input is a VRS location |
| liftover | mappers/liftover.py |
chain files via env |
Default UTA URL is postgresql://anonymous@localhost:5432/uta?…uta_20241220 [M]. So everything cool-seq-tool
does with exons is UTA; the gene/CDS overlap it does from MANE alone is something a polars read of the
same GFF reproduces without the library [I].
3. Data already published in GKS form¶
| Publisher | What | Spec version | Distribution / size | Cadence | Licence | Status |
|---|---|---|---|---|---|---|
ClinVar-GKM (ClinGen; clingen-data-model/clinvar-gkm, formerly clinvar-gks) |
All ClinVar variations as Cat-VRS CategoricalVariant (CanonicalAllele, CategoricalCnvChange, CategoricalCnvCount, or constraint-less ClinvarNonConstrainedVariant for haplotypes/genotypes/compound hets); VRS Allele/CopyNumberCount/CopyNumberChange; SCV/RCV/VCV VA-Spec statements; conditions; genes |
VRS 2.0.1 output schema; custom Clinvar*Proposition classes (incl. ClinvarDrugResponseProposition, ClinvarRiskFactorProposition) |
Cloudflare R2 pub-f0ad0e0dac0345408dcc95bda20beb42.r2.dev. Monthly full JSON 4.76 GB gz; typed parquet, 20 sections, 13.4 GB total (variation 4.54 GB, scv 2.95 GB, varcond-proposition 1.50 GB, allele 1.09 GB, location 0.60 GB); weekly delta 3.5 MB + manifest.json |
"weekly delta, monthly full". Observed: latest delta release: 2026-08-22, Last-Modified 2026-08-26 — ~5 weeks old on 2026-09-28 |
CC0-1.0 (code and data) | 1.0-rc3 (2026-08-26); a breaking section split (proposition → four var*-proposition) landed mid-RC |
| CIViC | civicpy.exports.civic_gks_record: assertions as VA-Spec statements; molecular profile → CategoricalVariant id civic.mpid:N, mappings read from molecular_profile.variants[0] |
via ga4gh.cat_vrs, ga4gh.va_spec |
produced by civicpy 5.4.0, not a bulk dump | on demand | CIViC data CC0 (recalled, unprobed) | code exists; no published GKS bundle found |
| metakb (VICC) | CIViC + MOA harmonized into VA-Spec | VA-Spec 0.5.0a0 python | web service + neo4j | — | MIT (code) | 2.0.0-a0; skips any CIViC MP whose name is not GENE p.change, and any that normalizes to >1 constraint |
| MaveDB | 3,442,148 mapped variants [M]; VRS alleles; VA-Spec study-result, functional-statement, pathogenicity-statement per mapped variant and per score set (streaming) |
VRS 2.0 / VA-Spec (version per API) | REST api.mavedb.org (v2026.2.7.3) |
continuous | per score set: CC0, CC BY 4.0, CC BY-SA 4.0, CC BY-NC-SA 4.0, "Other" [M: /api/v1/licenses/] |
production |
| gnomAD | VRS allele ids in API / v4.1 VCF | the repo verified SNV ids equal to the live API (docs/SCHEMAS.md § Allele identity) |
— | — | — | already consumed (frequencies.csv.vrs_id) |
| ClinGen Allele Registry | /vrAllele?hgvs= returns a VRS 1.x-era object; CA records carry no VRS id field |
VRS 1.x shape (minor not established) | REST | — | — | ids not joinable to VRS 2 (headline 3) |
| ClinGen LDH / Evidence Repository | LDH states VA-Spec 1.0 as its primary GKS dependency | VA-Spec 1.0 | REST | — | — | [D: search result for clingen-data-model/svcv4-model; LDH about page]. Not probed |
| ClinPGx / PharmGKB / CPIC | nothing found in GKS form | — | — | — | — | Cat-VRS star-allele recipe is only a discussion (#243, progress step 1 unchecked) |
| GWAS Catalog / PGS Catalog | nothing found | — | — | — | — | — |
Sources: [M] clinvar-gkm README, docs/data-access/download.md, docs/pipeline/vrs-processing.md,
releases; HEAD requests against R2; duckdb parquet_schema/count(*) over R2; MaveDB OpenAPI;
civicpy source at griffithlab/civicpy; metakb server/src/metakb/transformers/civic.py.
ClinVar-GKM's variation.parquet also carries, per HGVS expression, molecularConsequence (SO
code/system/name) with maneSelect/manePlus flags [M: parquet schema] — ClinVar's own
consequence calls, CC0, per transcript: a candidate source for RM236.
ClinVar-GKM runs VRS through clinvar-gk-python (vrs-python + variation-normalizer) against a local
SeqRepo + UTA + gene-normalizer, only on the single best expression per variation, and carries
failures as errors [D: docs/pipeline/vrs-processing.md].
4. Mapping onto our tables¶
| Our table (grain) | GKS counterpart | Fit | Gap |
|---|---|---|---|
variants.csv (variant_key, genotype) |
VA-Spec Statement over VariantPathogenicityProposition; subject a VRS Allele or Cat-VRS CanonicalAllele |
partial | no genotype/zygosity anywhere in VRS 2 or VA-Spec; weight, state, direction, negatives have no slot (a Statement has direction/strength/score, which are about evidence, not effect direction) [I] |
resolution.csv (rsid ↔ coord, per-ALT vrs_id) |
VRS Allele + expressions (spdi/hgvs/gnomad) — exactly what ClinVar-GKM's allele.data holds |
close | we key on rsID; VRS keys on content; no VRS version stamp on our side |
frequencies.csv per allele |
CohortAlleleFrequencyStudyResult (trial use) |
close | allele_count→focusCount, allele_number→locusCount, population→cohort/subCohortFrequency; homozygote_count, hemizygote_count, faf95 only via ancillaryResults/qualityMeasures; field names moved between 1.0.1 and 1.1 |
clinical_assertions.csv per (allele, archive record) |
ClinVar-GKM SCV statements (scv.parquet) |
close | ours is a projection of theirs; where review stars sit in theirs was not checked |
clin_sig_concordance.csv / …_authority_calls.csv |
ClinVar-GKM VCV/RCV aggregate statements | partial | our concordance is across authorities; theirs is within ClinVar |
gene_validity.csv |
GeneDiseaseValidityProposition |
close | draft, 1.1 ballot only |
pharm_variants.csv per genotype |
VariantTherapeuticResponseProposition |
poor | scoped to neoplasm response; per-genotype grain absent; ClinVar-GKM uses a custom ClinvarDrugResponseProposition |
haplotypes.csv, allele_function.csv, diplotypes.csv, activity_phenotype.csv |
none (VRS CisPhasedBlock covers the cis variant set of a star allele only) |
none | star-allele recipe not designed; no diplotype |
gwas_effects.csv, pgs.csv |
none | none | — |
expression_effects.csv (eQTL) |
none (ExperimentalVariantFunctionalImpact is assay-based) | none | — |
copynumbers.csv |
VRS CopyNumberCount (trial use); Cat-VRS CopyCountConstraint |
concept fits | ours is not positional (RM65) |
repeat_alleles.csv |
VRS ReferenceLengthExpression (length, repeatSubunitLength) |
one motif only | multi-motif RUS=CAG,TG,CAGG (RM66) has no VRS form [I] |
heteroplasmy.csv |
none | none | — |
studies.csv / literature.csv |
VA-Spec Document (reportedIn), StudyResult |
loose | provenance_quote has no slot |
gene_metrics.csv |
none | none | — |
sources.csv |
VA-Spec DataSet (sourceDataSet), Contribution |
loose | licence terms not modelled |
| AlphaGenome AVI / RM23 predictions | ComputationalVariantFunctionalImpactAnalysisResult |
shape matches RM23's long form | draft |
5. Gene / exon coordinate mapping for AlphaGenome¶
Context from the repo: AVI is one scalar per (position, ALT) with no gene; AlphaGenome's annotation is
GENCODE v46 (docs/probes/ALPHAGENOME_ATLAS.md §1.1, §8); gene_spans.py already reads the MANE
summary for spans under the rule "a span is a query hint, never an attribution".
| Rank by weight | Option | Size | What it answers | What it cannot | Licence | Tier |
|---|---|---|---|---|---|---|
| 1 (lightest) | MANE Ensembl GTF (MANE.GRCh38.v1.5.ensembl_genomic.gtf.gz) as a second file in the existing MANE lane |
8.6 MB gz [M]; 19,363 genes, 19,437 transcripts (74 MANE Plus Clinical), 204,815 exons, 194,202 CDS, 48,592 UTR rows [M]; versioned ENSG/ENST, exon_number, exon_id, db_xref "RefSeq:NM_…" |
exon boundaries and numbers, CDS/UTR, strand, for MANE Select + Plus Clinical | non-coding genes (MANE is protein-coding only; GENCODE v46, AlphaGenome's annotation, includes non-coding genes [I]); any non-MANE isoform; release skew: MANE 1.5 = Ensembl 116 / RefSeq RS_2025_08 [M: README_versions.txt], AlphaGenome = GENCODE v46 (= Ensembl 112) [I from GENCODE numbering] |
NCBI policy, already recorded as MANE_TERMS (@no-named-licence) |
enricher builder (polars) |
| 1b | MANE RefSeq GFF (7.9 MB) / RefSeq GTF (8.7 MB) | same content, RefSeq ids | same | same | same | same |
| 2 | GENCODE v46 basic GTF (gencode.v46.basic.annotation.gtf.gz) |
29 MB gz (comprehensive 49 MB) [M: EBI FTP listing] | the exact annotation AlphaGenome used, all biotypes | nothing MANE-specific unless joined to MANE by ENST | page states only "All the GENCODE project data is open access" [D]; no licence named → @no-named-licence |
enricher builder |
| 3 | Ensembl current GTF (release 116) | not measured (/pub/current_gtf/homo_sapiens/ 404'd) |
all genes/transcripts, matches MANE 1.5's Ensembl release | exact GENCODE v46 parity | "Ensembl imposes no restrictions on access to, or use of, the data" [D: Ensembl disclaimer] | enricher builder |
| 4 (heaviest) | cool-seq-tool + UTA + SeqRepo | 309 MB venv / 53 pkgs [M]; UTA Postgres (image 174 MB; dump size not established); SeqRepo 13.5 GB [D] | any transcript version incl. historical, c.↔g. with exon offsets, MANE selection logic | nothing extra that AVI attribution needs [I] | UTA CC BY-SA 4.0 (→ ShareAlike on derived tables; the compile gate would carry it) | not here: needs Postgres service + 13.5 GB store; conflicts with "builder in polars, runtime in duckdb" and --offline provisioning [I] |
The existing MANE builder already pins the versioned directory and parses README_versions.txt; adding
the GTF is one more file in the same lane [I from mane_build.py docstring].
6. Cat-VRS for meta-conclusions (RM28) — coordinator addendum¶
6.1 Does a CategoricalVariant have a computed GA4GH digest?¶
No. Evidence, all primary:
- Spec source:
cat-vrs-source.yamlat both1.1.0and1.1.1-ballot.2026-09.1contains noga4gh.inherent, noprefix:, nodigest; every classinherits: gkm-core:EntityorConstraint[M: grep of the clone at 1.1.1 andgh apiof the 1.1.0 file]. By contrastvrs-source.yaml@2.1.0declaresprefix:VA/RA/CPB/CN/CX/SL/RSL withinherent:lists [M]. The GA4GH digest machinery only applies to classes that declare those. - Reference implementation:
ga4gh.cat_vrs.models.CategoricalVariant(Entity, BaseModelForbidExtra)— anEntity, not aGa4ghIdentifiableObject; the JSON schema properties areid, name, description, aliases, extensions, type, members, constraints, mappings(nodigest) [M]. - Design record: Cat-VRS models categorical variants as hyperintensional sets, explicitly so
that two catvars with identical properties but different curatorial context stay distinct
("7-14075336-A-T, NM_004333.6…, and rs113488022 … a hyperintensional model allows us to represent
each of these catvars in parallel") [D:
docs/source/appendices/design_decisions.rst]. A content-derived digest would collapse exactly what that decision keeps apart [I]. - Publishers: ClinVar-GKM ids are
clinvar:<VariationID>(all 4,553,639 rows areCategoricalVariant) [M]; civicpy usescivic.mpid:<N>[M]. Both system-assigned. - The
idfield description in ClinVar-GKM docs: "The 'logical' identifier of the Entity in the system of record … may or may not be globally unique outside the system" [D].
Consequence [I]: a Cat-VRS id is not a stable key we could mint ourselves the way we mint
ga4gh:VA.. If we minted one, it would be our own naming scheme wearing a GA4GH class name. The
constituent DefiningAlleleConstraint.allele is a VRS Allele with a digest, so atoms that are
specific alleles do have content ids; the categorical wrapper does not.
6.2 Which RM28 atoms Cat-VRS can express today¶
| RM28 atom | Cat-VRS form | Maturity |
|---|---|---|
A specific allele (VHL S183L at genomic level, an rsID's ALT) |
DefiningAlleleConstraint (+ CanonicalAllele recipe) |
trial use |
A protein change regardless of nucleotide (BRAF V600E) |
ProteinSequenceConsequence recipe |
trial use |
| "MET Amplification", "22q11.2 deletion" | CategoricalCnv = DefiningLocationConstraint + CopyChangeConstraint (gain/loss) or CopyCountConstraint |
recipe draft; CopyCount trial use, CopyChange draft |
| "KRAS mutation" (any variant in a gene) | FeatureContextConstraint (a gene MappableConcept) |
draft |
| Loss-/gain-of-function class | FunctionConstraint / FunctionVariant recipe |
draft |
Fusion (BCR::ABL1) |
AdjacencyConstraint / GeneFusion recipe |
draft |
A star allele (CYP2C9*2) |
none — discussion #243 only | — |
| A diplotype / genotype | none in VRS 2 or Cat-VRS; ClinVar-GKM files CYP2C19*22/*31 etc. as constraint-less ClinvarNonConstrainedVariant keyed only by ClinVar id [M] |
— |
| SNPedia genoset (a set of genotypes) | none | — |
6.3 Does anything in GKS express the combinator?¶
| Need (RM28) | GKS status | Source |
|---|---|---|
| AND / OR over variants | CompositeCategoricalVariant (elements, operator: AND|OR) — Cat-VRS PR #248, open, unmerged, last updated 2026-08-10, not in the 2026-09 ballot |
[M] gh api repos/ga4gh/cat-vrs/pulls/248 + schema at PR head |
NOT / open-world negation (AND NOT KRAS Mutation, "wild type") |
CategoricalVariantCriterion.presence: present|absent — negation only as a criterion on a single catvar, per the design thread (Puthawala: restrict NOT to atoms "for normalization"); NOR discussed then dropped from the schema |
[M] discussion #241, PR #248 |
Phase: trans compound het (VHL S183L AND VHL D126N, "heterozygous compound mutation") |
not expressible. PR #248 has no phase field [M]; the 2026-08-05 comment says the meeting "focused on how this aligns with phasing for additional use cases, such as compound hets and Star Alleles" [D] — no phase design exists. VRS CisPhasedBlock is cis only. ClinVar-GKM maps ClinVar CompoundHeterozygote to a constraint-less catvar [D] |
[M] discussion #241 comments; clinvar-gkm docs/output-reference/cat-vrs.md |
| Cis haplotype | VRS CisPhasedBlock (trial use) — an ordered set of Alleles on one molecule, digest prefix CPB |
[M] |
| Genotype / diplotype (warfarin CYP2C9 diplotype + VKORC1 genotype) | no class in VRS 2.x, Cat-VRS or VA-Spec. VRS 1.3 had Genotype/GenotypeMember (2022 bioRxiv model) and it is absent from 2.x |
[M] class lists; [D] VRS 1.3 release notes |
| How CIViC multi-variant MPs appear in GKS form today | civicpy: one catvar per MP, mappings from variants[0] only; metakb: skips any MP not matching GENE p.change and any with >1 constraint. Neither emits a boolean structure. |
[M] source reads |
A VA-Spec Proposition takes one subject (a Variation or CategoricalVariant), which is why the
Cat-VRS team frames composition as Cat-VRS's job [D: discussion #241 body].
7. Open roadmap items a GKS product would shrink or dissolve¶
| RM | What GKS offers | Effect | Evidence tag |
|---|---|---|---|
RM270 (indel spelling; vrs_id dropped from artifact) |
Confirms candidate 4: VRS justification makes both spellings of rs72613567/rs77944059 one id, and ClinVar-GKM carries the same ids [M]. Caveat: VRS 2.1 smallest-factor rule will move some RLE insertion ids once vrs-python implements it (#637); and no stored vrs_id names its VRS version |
shrinks (the id and a second publisher already exist); adds one question: stamp the VRS version or not | M + D |
| RM267 (Ensembl −1 insertion anchor) | AnyVar/variation-normalizer add nothing over core ga4gh.vrs: normalization detects the wrong event (two ids for S117's rs8176719 pair, already probed in RM270) but cannot repair it, which RM267 already says. ClinVar-GKM's allele.parquet (gnomad expression + VRS id per ClinVar allele) is a second placement the enricher could hold offline, as candidate 4 ("a finding that withholds") needs |
shrinks candidate 4's acquisition problem; does not dissolve | M + I |
| RM292 (a check says what it checked against) | ClinVar-GKM is an independent input (ClinVar's SPDI, not Ensembl's dump) run through the same algorithm (vrs-python). A VRS-id comparison therefore witnesses placement, not the normalizer | gives RM292 a concrete non-self witness; the "self-consistent vs verified" split still applies to the algorithm half | I |
RM236 (consequence/impact) |
VA-Spec VariantMolecularConsequenceProposition (draft) carries a transcript qualifier → validates the per-(variant, transcript) sidecar option (3). ClinVar-GKM variation.parquet publishes SO consequence per HGVS with MANE flags, CC0 |
informs grain decision; offers a CC0 source for ClinVar variants only | M + D |
| RM23 (predictor scores) | ComputationalVariantFunctionalImpactAnalysisResult (draft): one result per (variant, tool, transcript context) with impactScore, impactScoreType, categoricalImpact, impactedFeature — the long form RM23 proposes, with transcript as an optional context |
informs; does not settle acquisition | D |
| RM237 (ClinGen dosage regions) | CategoricalCnv expresses the variant class ("22q11.2 deletion" = location + copy change). No VA-Spec proposition for dosage sensitivity (HI/TS scores) exists |
partial: names the region-CNV subject; the HI/TS claim has no GKS home | M |
| RM65/RM66 (positional repeats/copy number; multi-motif) | VRS CopyNumberCount (trial use) is positional by construction; ReferenceLengthExpression holds one repeat unit. Multi-motif alleles have no VRS form |
RM65: confirms positional shape; RM66: nothing | M + I |
| RM28 (meta-conclusions) | §6: atoms mostly yes; AND/OR/absent in an open draft; trans phase and genotypes no | does not dissolve; the draft is worth tracking | M |
| RM298 (per-row ClinPGx provenance) | VA-Spec Statement.contributions, reportedIn, specifiedBy, StudyResult.sourceDataSet put provenance on every statement |
precedent for per-row source, not a mechanism for our CSV | D |
| RM302 (conclusion registers) | nothing: Statement has description and name; no register notion |
no effect | M |
| RM303 (snapshot contract) | GKS has no snapshot/release descriptor. ClinVar-GKM's delta manifest.json (release, baseline_release, pipeline_version, checkpoint_full, per-section adds/updates/deletes) is a precedent for keys; seqcol v1.0 could give an assembly a content digest |
precedent only | M + D |
8. Where each candidate could live (tier rules) — tensions stated, not decided¶
| Candidate | format | compiler | enricher | not here | Charter tension |
|---|---|---|---|---|---|
Stamp VRS version beside vrs_id (column or manifest field) |
yes (model field) | yes (parquet column) | writes it | — | New optional column is minor-legal (P3/P8); parquet column ~free, derived CSV half (P9) |
Carry vrs_id onto variant parquets (RM270-4) |
— | yes | fills | — | already analysed in RM270 |
ClinVar-GKM allele.parquet witness lane |
— | — | yes (httpx + polars builder; duckdb read) | — | none by dependency; licence CC0 |
| MANE GTF exon lane | — | — | yes | — | none |
| GENCODE v46 GTF lane | — | — | yes | — | no named licence → warns, never gates (@no-named-licence) |
ga4gh.cat_vrs / ga4gh.va_spec models for a GKS export view |
tension: Goal 2 / Non-goals fix format at "pydantic plus cryptography"; these pull ga4gh.vrs → bioutils, requests, canonicaljson (+2 small pkgs). Weight is small; legality would be a Goals amendment regardless |
tension: compiler is "polars/pyyaml/typer … pure-Python"; ga4gh.vrs imports requests (no fetch required for model use [I]), so P2 is not breached by importing, but the dependency list is |
fits: ga4gh.vrs is already a core dep; +144 KB |
or downstream in just-dna-pipelines |
Goal 2 names the format tier's deps exhaustively |
Cat-VRS CategoricalVariant as an RM28 subject |
model could mirror the constraint shapes in pydantic without the package [I] | — | — | — | Mirroring a draft class into an authored table costs full price (P9) and a draft that changes shape forces a P3 question later |
| cool-seq-tool / UTA / SeqRepo / variation-normalizer / AnyVar / metakb | — | — | by weight, possible only as optional extras; Postgres + 13.5 GB store + (UTA) CC BY-SA | yes — service stacks | Non-goals bar orchestration-class deps from format/compiler; the enricher has no bar, but "builder in polars, runtime in duckdb" and --offline provisioning are house rules they would bend |
The Constitution's Goals name just-dna-format's deps as "only pydantic plus cryptography"; any GKS
package there is a Goals amendment, not a weight judgement. The enricher has room by rule; the question
there is P9 price and the lanes' licences.
9. Could not be established¶
- UTA dump size (
dl.biocommons.orgserves a human-verification page to curl). - Ensembl current GTF size (FTP
current_gtfpath returned 404). - Phenopackets ↔ VRS 2 status (not checked).
- CIViC data licence (recalled CC0, not probed this session).
- Which of ClinVar-GKM's 270,161 RLE alleles would change id under the 2.1 smallest-factor rule.
- Whether ClinVar-GKM's weekly cadence resumed after 2026-08-22.
- ClinGen LDH / ERepo GKS output (search result only, no probe).
10. Top 5 adoption candidates, by value / cost¶
| # | Candidate | Value | Cost (concrete work) | Dependency added |
|---|---|---|---|---|
| 1 | ClinVar-GKM allele.parquet as an offline placement witness — a builder slice to (vrs_id, spdi, gnomad_expr, clinvar_variation_id) and a check comparing the Ensembl-snapshot spelling of each resolved indel against it (RM267 candidate 4, RM292 witness, RM270 evidence) |
High: turns RM267's "needs a second placement" into a local read; independent input; CC0 | 2–3 days: 1 day builder (1.09 GB allele.parquet + 0.60 GB location.parquet, whose typed loc_start/loc_end/sequence_reference_id keyed by location_id need no JSON parsing; the gnomad expression sits inside allele.data JSON), 1 day check + SourceRow + finding text, 0.5 day tests on the RM267/RM270 rsIDs. Plus RC risk: the schema moved mid-RC once |
none (httpx + polars already present) |
| 2 | MANE GTF exons into the existing MANE lane for AlphaGenome gene/exon attribution | High for the AlphaGenome work; tiny data | 1–1.5 days: 0.5 day add ensembl_genomic.gtf.gz to mane_build.py → exons.parquet; 0.5 day gene_spans-style lookup (exon number, CDS/UTR, MANE status); tests. Optional +1 day for a GENCODE v46 basic GTF lane if non-coding genes or v46 parity matter |
none |
| 3 | Stamp the VRS version behind every vrs_id (manifest field or column), before RM270 carries ids into parquets |
Medium-high: prevents the headline-3 confusion and absorbs the pending 2.1 RLE change honestly | 0.5–1 day in format + enricher (a constant from ga4gh.vrs.VRS_VERSION and the stdlib minter's declared spec) + tests; a design choice on where it lives |
none |
| 4 | Borrow VA-Spec 1.1 draft shapes as the reference grain for RM23 / RM236 (ComputationalVariantFunctionalImpactAnalysisResult, VariantMolecularConsequenceProposition) — write the mapping into the RM entries, no code |
Medium: settles "which transcript" by precedent; cheap | 2–3 hours of doc work | none |
| 5 | ga4gh.cat_vrs + ga4gh.va_spec in the enricher for a GKS export view of compiled modules (frequencies → CohortAlleleFrequency, clinical assertions → pathogenicity statements) |
Medium for interop, low for our own correctness; genotype grain is lossy | 3–5 days, and better after VA-Spec 1.1 / Cat-VRS 1.1 go stable on PyPI (today alpha-only), since 1.1 renames trial-use fields | ga4gh.cat_vrs, ga4gh.va_spec (+2 pkgs, ~144 KB) |
Not recommended on weight (measured): cool-seq-tool (309 MB, UTA + SeqRepo), variation-normalizer (421 MB, three services), AnyVar (259 MB, SQL store + SeqRepo), metakb (neo4j + all normalizers).
Watch, do not adopt yet: Cat-VRS CompositeCategoricalVariant (PR #248) and star-allele recipe
(#243) for RM28; vrs-python #637 for the 2.1 RLE change.