just_dna_enricher.civic_vcf¶
just_dna_enricher.civic_vcf ¶
The dated accepted_and_submitted VCF — CIViC's own submitted-inclusive publication (RM169).
Why this file exists, and why it is not the builder's input. civic build reads the dated bulk
TSV pair, every row of which is evidence_status = accepted. CIViC also publishes, in the same
dated release directory, <date>-civic_accepted_and_submitted.vcf — so the wider basis is pinnable
and byte-reproducible, and no API read is needed to reach it. RM160 was filed on the belief that it
was; that belief was wrong, and this module is what replaced it.
But the VCF is a strict subset of the TSV, and the subset is not arbitrary. A VCF record needs a
POS, so a CIViC variant with no GRCh37 coordinate cannot appear in one at all. Measured over
01-Aug-2026, the accepted VCF carries 473 germline direction rows on 236 variants where the TSV
carries 533 on 290 — and 52 of the 54 it drops are exactly the unresolvable_identity class, the
records whose identity had to be read out of their names (RM159). Reading the VCF as the row source
would silently discard the hardest-won half of this snapshot.
So the TSV pair stays primary and this file is joined onto it for two things it alone carries:
evidence_statusper evidence item, which is what makes an accepted row and a submitted row distinguishable once both are in the parquet.- The submitted evidence itself — 642 further direction rows on 247 variants over
01-Aug-2026, of which 127 variants are new to the snapshot. The whole build goes 507 rows on 270 variants to 1,149 on 397.
Nothing is placed from a coordinate this file states. Its POS is GRCh37
(##reference=…GRCh37-lite.fa.gz), and lifting it is refused on measurement (RM48, and the class-A
working in docs/probes/CIVIC_UNRESOLVED.md). Every row is placed from a published identifier read out
of the CSQ block, through the same parsers the TSV path uses.
The CSQ block is one entry per evidence item, not per variant. A single VCF record carries a
comma-separated list, each entry naming a variant and the evidence item asserting something about
it, so the join key is (variant_id, evidence_id) and a record with three CSQ entries is three rows.
Its field order is declared in the header rather than assumed — see parse_csq_format.
CivicVcfError ¶
Bases: RuntimeError
The VCF cannot be read: absent, headerless, or missing a CSQ field this module needs.
CivicVcfEntry
dataclass
¶
CivicVcfEntry(
variant_id: int,
evidence_id: int,
status: str,
variant_origin: str,
significance: str,
direction: str,
disease: str,
citation_id: str,
source_type: str,
evidence_level: str,
rating: str,
variant_name: str,
variant_aliases: str,
civic_hgvs: str,
allele_registry_id: str,
gene: str,
molecular_profile_id: str,
molecular_profile_name: str,
)
One CSQ entry — one evidence item's view of one variant.
The evidence-level fields are carried here because nothing else can supply them. Origin, significance, direction, disease, level, rating and the citation are facts about the evidence item, and a submitted item has no row in the evidence TSV — so for those items this file is the only source.
And so is the variant-level half, for the variants the TSV omits. VariantSummaries.tsv is
accepted-only as well, so 112 of the 127 variants the submitted evidence introduces have no row
there at all: no gene, no aliases, no HGVS, no registry id. The same CSQ entry carries all four,
so they are read from here only when the TSV cannot answer, and the resulting rows are stamped
vcf_csq rather than folded into the TSV-derived derivations — a row whose identity came from a
different file is a different provenance claim, and a consumer must be able to see which
(@source-vs-authority, applied to two files of one release). The other 15 join the TSV normally
and take an ordinary derivation, which is why the stamp counts rows rather than the whole widening.
Nothing here is read from the record's POS. The VCF is GRCh37 and lifting it stays refused (RM48); all 112 place on a published identifier — 57 by ClinGen CAID, 40 by rs-number, 14 by a GRCh38 accession, 1 by both — through the same parsers the TSV path uses, measured over the emitted parquet rather than over the file.
The vocabulary differs from the TSV's: the VCF is SCREAMING_CASE (RARE_GERMLINE,
PREDISPOSITION, SUPPORTS) where the TSV is title case (Rare Germline, Predisposition,
Supports). One normalizer, applied at the boundary, and both sides' raw tokens are tested
(@one-normalizer-two-spellings).
parse_csq_format ¶
The CSQ sub-field names, in the order the file declares them.
Read rather than hardcoded. The header carries Format: A|B|C, and taking it from there is what
keeps this parser correct when CIViCpy adds a field — a positional constant would keep parsing and
return the wrong column, which is the failure mode with no symptom.
Source code in enricher/src/just_dna_enricher/civic_vcf.py
read_vcf_entries ¶
Every evidence CSQ entry in the file, with the status CIViC assigned it.
Assertion entries are skipped: CIViC Entity Type is evidence or assertion, and an assertion
is a different table with a different key. Reading both under one name would let an assertion's
status stand in for an evidence item's.
Source code in enricher/src/just_dna_enricher/civic_vcf.py
assert_vocabulary_covers ¶
Every vocabulary token in the file has a TSV spelling.
Redundant with _title raising per row, and kept because it states the property as an equality
over a walked set rather than as a side effect of parsing (@registry-completeness): a caller can
put the question directly, and a test can prove the guard without contriving a bad row.
Source code in enricher/src/just_dna_enricher/civic_vcf.py
status_by_evidence ¶
evidence_id → status.
One entry per evidence item is the normal case, but a multi-variant profile repeats an item once per member variant, so the mapping is many-to-one and must agree with itself. A disagreement is raised rather than resolved: it would mean the file states two statuses for one item, and picking one would publish a coin-flip as the source's word.
Source code in enricher/src/just_dna_enricher/civic_vcf.py
summarize ¶
Evidence items per status, over the whole file — the denominator a count here is against.