just_dna_compiler.hints¶
just_dna_compiler.hints ¶
Authoring hints — answer questions about authored cells without ever writing one (0.5).
The third piece of the authoring surface, beside draft (append rows into an existing table) and
scaffold (create tables from nothing). This one writes nothing at all: CSV text in, a report
out. Applying anything is the human's move, or draft.append_rows'.
Why a hint may not simply fill the cell it knows the answer to. COMPILER.md's Class 2,
validate-by-redundancy, is "where most real authoring bugs are caught", and every check in it
compares two independently-authored things. Fill chrom/start from Ensembl and the compiler's
rsid↔coordinate check compares Ensembl against Ensembl; fill doi from PubMed and
literature._doi_conflicts compares PubMed against PubMed. The check does not merely become
tautological — for an rsid-only row resolution._verify never runs at all, so the row moves from
honestly unverified to verified against whatever filled it, and the compile now reports success.
Auto-filling does not bend a convention; it deletes a validation class.
The codebase already made this argument once, for one field, in literature: Crossref is asked about
the authored DOI in preference to the derived one, because checking the registry's own DOI "would
be circular — it exists by construction". This module generalizes that to every cell.
So hints are partitioned by what the value is, not by where it came from:
- normalized — a re-spelling of the author's own cell by a rule the model already applies on
load. Safe to render, and arguably must be: the model does it anyway, silently.
DiplotypeRowswapshaplotype_a/haplotype_bwithout saying so. - derived — computed from other authored cells on the same row. Not rendered: it adds no
information and spends a redundancy (
p_value→p_value_numis cross-checked at compile). - advisory — an external fact. Reported, never rendered, whether or not a checker consumes it
today;
refusalnames why.
csv_out therefore differs from the input only by normalized alterations, and every difference is
listed. Nothing here needs the network: that half lives in the enricher, which supplies the advisory
facts through the same report shape.
Alteration
dataclass
¶
Alteration(
row: int,
column: str,
before: str,
after: str,
kind: str,
applied: bool,
source: str = "model",
refusal: str | None = None,
note: str = "",
)
One cell that differs between an input row and the emitted one, or that a source disagrees with.
applied is the whole contract: it is True only for normalized, and csv_out reflects exactly
the applied ones. A caller can therefore diff the report instead of trusting the prose.
Finding
dataclass
¶
Something worth telling the author about a cell or a row. Never fails a build.
Two coordinates, because the two audiences count differently (S18). row is a 0-based index
into the data rows, which is what a caller indexing csv_out or a parsed list needs. line is
the 1-based line number in the file, counting the header, which is what the author's editor
shows and what validate_spec/compile_module already report (line 2 [source]: …). Neither was
stated, and the only way to learn which one row was, was to author a broken row and count — on a
short table an off-by-one still lands on a real row, so it misdirects instead of obviously breaking.
line is set by whatever builds the finding, since only that code knows whether the input carried
a header at all; inspect_rows fills it for every row-scoped finding. It stays None for a
table-scoped finding, exactly as row does.
FieldOptions
dataclass
¶
FieldOptions(
column: str,
vocabulary: str,
options: list[str],
closed: bool,
notes: dict[str, str] | None = None,
)
The values a column accepts, for a picker.
notes is the per-member prose for the vocabularies whose member name cannot carry the whole
rule, None for the rest — a picker offering largest beside largest_alt has nowhere else to
say which of them counts the reference element. It rides here because it rides on the marker
(base.vocabulary), so the surface an author reaches when their row was rejected knows what
describe and the whole-schema reference know.
HintReport
dataclass
¶
HintReport(
csv_name: str,
header: list[str],
rows_in: int,
csv_out: list[str],
alterations: list[Alteration] = list(),
findings: list[Finding] = list(),
options: list[FieldOptions] = list(),
requirements: dict[str, Any] = dict(),
)
CSV text in, the same CSV out plus everything known about it.
Invariants, all pinned by tests: csv_out has the same row count and order as the input; only
normalized alterations are applied; every applied change appears in alterations.
to_json ¶
The machine form. Lists keep insertion order — a report that reorders is not diffable.
Source code in compiler/src/just_dna_compiler/hints.py
TableKey
dataclass
¶
TableKey(
columns: tuple[str, ...],
rule: str,
stamped: tuple[str, ...] = (),
fallback: tuple[str, ...] = (),
)
What makes two rows of one kind the same row, in columns an author can actually see (S48).
draft.natural_key answers this per row, as a tuple of values, so a tool holding only a table
name could not ask it which columns those values came from — and _TABLE_DUPE_KEYS, even reached
into, yields lambdas that name no columns at all. Every consumer wanting the column names was
therefore hand-keeping a string, and a hand-kept string went stale the day 0.6 deprecated a column
inside one: the reporting consumer was telling authors to key copynumbers.csv on modifier_cn,
whose own description reads DEPRECATED since 0.6, removed at 1.0.
columns is in the author's own spelling wherever one exists: a derived _KEY_FIELDS member is
mapped back to the authored column it coalesces, never dropped. CopyNumberRow keys on
effective_modifier_copy_number, a property over two columns, and it appears here as
modifier_copy_number — the preferred spelling, so this surface can never hand an author the
deprecated half of a pair. Dropping the member instead, which is what the compiler's own
grounding-remedy sentence does, would say two rows differing only in modifier dosage are the same
row. That is a wrong answer where the drop is invisible, so it is the one thing this must not do.
stamped names the members the compiler fills rather than the author — variant_key on the
two variant-keyed kinds. They stay in columns because they really are part of the key; they are
flagged because an author cannot type one, and a surface that presented them as fillable columns
would be describing a cell that does not exist in their CSV.
fallback is the key used when every column of columns is null, and one derived table has
one: gene_validity.csv keys on the source's own assertion_id, and on the gene's grain when the
source published none. It is () everywhere else, and a consumer that ignores it is right about
seven of the eight tables — which is exactly why it is a field rather than a footnote (S51). A
fallback is only meaningful where the primary key can be absent, so a kind declaring one whose
primary column is required is a contradiction, and a test asserts it away rather than this
dataclass — the check needs the model, which a TableKey does not carry.
derived_model_for ¶
The row model for a machine-produced CSV, or DraftError for a name this format does not read.
The counterpart to draft.model_for, which answers for authored kinds only. Between them they
cover every CSV the compiler reads, and a tool that dispatches on a filename needs both: an author
reading resolution.csv or frequencies.csv — files they must read and must never hand-finish —
was getting "not an authored table of this format", which is true and useless (S47).
Raises DraftError rather than returning None so the failure matches model_for's, on the house
rule that a generic rejection is a dead end where a specific one is a fix: the message names the
authored route, because asking this function for variants.csv is the likely mistake and the
answer to it is one call away.
Source code in compiler/src/just_dna_compiler/hints.py
key_fields ¶
What makes two rows of this table kind the same row, or None for a kind with no declared key.
Answers for authored kinds and for the machine-produced tables alike, so a tool dispatching on a
filename gets one route. None is the honest answer for a derived table that declares no key —
withheld rather than invented, on the house three-valued rule.
Consistent with draft.natural_key by construction, not by agreement: both read the model's own
_KEY_FIELDS, and a test pins natural_key(row) against getattr over these columns for every
keyed kind. natural_key returns None for the binning kinds because their duplicate rule is
overlap, not equality; this function still names their grouping columns and says
rule="overlap", which is the distinction a tool needs to explain either one.
The machine-produced tables answer here too, and their key is the one the writing pass merges
on (S51). It was readable only as a dict-key expression inside the body of each enricher pass,
so every consumer re-deriving a sidecar had to guess it — and the guess a reporting consumer
shipped, the required members of the fact-field tuple, was measurably coarse on two of the seven:
it dropped disease_id from gene_validity.csv and variation_id from clinical_assertions.csv,
the two tables where one subject legitimately carries several rows. Each pass now keys its
existing dict off the model's declared tuple, so the pass and this surface cannot disagree.
Reading rule matters as much as reading the columns: resolution.csv's key is a subject, not
a unique row, and reporting it as equality would be a wrong answer rather than a coarse one.
Source code in compiler/src/just_dna_compiler/hints.py
describe_table ¶
Everything a tool needs to help someone fill one table kind, without any input from them.
Columns with their types, requiredness and allowed values; the alternative identity groups; the natural key two rows are the same row by. All generated from the live models.
Whatever the whole-schema authoring_reference() says about a column, this says too — this
is the command an author filling one table runs, so a fact that reaches only the other surface
reaches the reader least likely to be reading it. Two were missing, both because this column dict
is assembled here rather than shared with the other surface, and a guard compared only
field_category. The per-member prose a closed vocabulary can carry: source_element's eight
members printed with no way to tell which of them counts the reference element, which is the one
question the vocabulary was added to answer (D1-4). And category: required alone is
pydantic's two-way answer, false for MeasureBinRow.measure_kind and unresolved, which an
author must still write out carrying their defaults. required stays beside it — a published key
is not removed, and it is insufficient rather than wrong.
Source code in compiler/src/just_dna_compiler/hints.py
field_options ¶
The pick-lists for one table kind — the "valid options to select from" half of authoring.
Source code in compiler/src/just_dna_compiler/hints.py
inspect_rows ¶
Inspect authored CSV text: validate each cell, preview the model's own normalizations, and report duplicate keys and bin incoherence. Pure — nothing on disk is read or written.
csv_text may carry a header or not; without one the model's authored column order is assumed,
which is what a caller pasting a single row from a template has.