Skip to content

just_dna_compiler.hints

just_dna_compiler.hints

Authoring hints — answer questions about authored cells without ever writing one (0.5).

The third piece of the authoring surface, beside draft (append rows into an existing table) and scaffold (create tables from nothing). This one writes nothing at all: CSV text in, a report out. Applying anything is the human's move, or draft.append_rows'.

Why a hint may not simply fill the cell it knows the answer to. COMPILER.md's Class 2, validate-by-redundancy, is "where most real authoring bugs are caught", and every check in it compares two independently-authored things. Fill chrom/start from Ensembl and the compiler's rsid↔coordinate check compares Ensembl against Ensembl; fill doi from PubMed and literature._doi_conflicts compares PubMed against PubMed. The check does not merely become tautological — for an rsid-only row resolution._verify never runs at all, so the row moves from honestly unverified to verified against whatever filled it, and the compile now reports success. Auto-filling does not bend a convention; it deletes a validation class.

The codebase already made this argument once, for one field, in literature: Crossref is asked about the authored DOI in preference to the derived one, because checking the registry's own DOI "would be circular — it exists by construction". This module generalizes that to every cell.

So hints are partitioned by what the value is, not by where it came from:

  • normalized — a re-spelling of the author's own cell by a rule the model already applies on load. Safe to render, and arguably must be: the model does it anyway, silently. DiplotypeRow swaps haplotype_a/haplotype_b without saying so.
  • derived — computed from other authored cells on the same row. Not rendered: it adds no information and spends a redundancy (p_value → p_value_num is cross-checked at compile).
  • advisory — an external fact. Reported, never rendered, whether or not a checker consumes it today; refusal names why.

csv_out therefore differs from the input only by normalized alterations, and every difference is listed. Nothing here needs the network: that half lives in the enricher, which supplies the advisory facts through the same report shape.

Alteration dataclass

Alteration(
    row: int,
    column: str,
    before: str,
    after: str,
    kind: str,
    applied: bool,
    source: str = "model",
    refusal: str | None = None,
    note: str = "",
)

One cell that differs between an input row and the emitted one, or that a source disagrees with.

applied is the whole contract: it is True only for normalized, and csv_out reflects exactly the applied ones. A caller can therefore diff the report instead of trusting the prose.

Finding dataclass

Finding(
    row: int | None,
    column: str | None,
    level: str,
    message: str,
    line: int | None = None,
)

Something worth telling the author about a cell or a row. Never fails a build.

Two coordinates, because the two audiences count differently (S18). row is a 0-based index into the data rows, which is what a caller indexing csv_out or a parsed list needs. line is the 1-based line number in the file, counting the header, which is what the author's editor shows and what validate_spec/compile_module already report (line 2 [source]: …). Neither was stated, and the only way to learn which one row was, was to author a broken row and count — on a short table an off-by-one still lands on a real row, so it misdirects instead of obviously breaking.

line is set by whatever builds the finding, since only that code knows whether the input carried a header at all; inspect_rows fills it for every row-scoped finding. It stays None for a table-scoped finding, exactly as row does.

FieldOptions dataclass

FieldOptions(
    column: str,
    vocabulary: str,
    options: list[str],
    closed: bool,
    notes: dict[str, str] | None = None,
)

The values a column accepts, for a picker.

notes is the per-member prose for the vocabularies whose member name cannot carry the whole rule, None for the rest — a picker offering largest beside largest_alt has nowhere else to say which of them counts the reference element. It rides here because it rides on the marker (base.vocabulary), so the surface an author reaches when their row was rejected knows what describe and the whole-schema reference know.

HintReport dataclass

HintReport(
    csv_name: str,
    header: list[str],
    rows_in: int,
    csv_out: list[str],
    alterations: list[Alteration] = list(),
    findings: list[Finding] = list(),
    options: list[FieldOptions] = list(),
    requirements: dict[str, Any] = dict(),
)

CSV text in, the same CSV out plus everything known about it.

Invariants, all pinned by tests: csv_out has the same row count and order as the input; only normalized alterations are applied; every applied change appears in alterations.

to_json

to_json() -> str

The machine form. Lists keep insertion order — a report that reorders is not diffable.

Source code in compiler/src/just_dna_compiler/hints.py
def to_json(self) -> str:
    """The machine form. Lists keep insertion order — a report that reorders is not diffable."""
    return json.dumps(
        {
            "csv_name": self.csv_name,
            "header": self.header,
            "rows_in": self.rows_in,
            "csv_out": self.csv_out,
            "alterations": [asdict(a) for a in self.alterations],
            "findings": [asdict(f) for f in self.findings],
            "options": [asdict(o) for o in self.options],
            "requirements": self.requirements,
        },
        indent=2,
    )

TableKey dataclass

TableKey(
    columns: tuple[str, ...],
    rule: str,
    stamped: tuple[str, ...] = (),
    fallback: tuple[str, ...] = (),
)

What makes two rows of one kind the same row, in columns an author can actually see (S48).

draft.natural_key answers this per row, as a tuple of values, so a tool holding only a table name could not ask it which columns those values came from — and _TABLE_DUPE_KEYS, even reached into, yields lambdas that name no columns at all. Every consumer wanting the column names was therefore hand-keeping a string, and a hand-kept string went stale the day 0.6 deprecated a column inside one: the reporting consumer was telling authors to key copynumbers.csv on modifier_cn, whose own description reads DEPRECATED since 0.6, removed at 1.0.

columns is in the author's own spelling wherever one exists: a derived _KEY_FIELDS member is mapped back to the authored column it coalesces, never dropped. CopyNumberRow keys on effective_modifier_copy_number, a property over two columns, and it appears here as modifier_copy_number — the preferred spelling, so this surface can never hand an author the deprecated half of a pair. Dropping the member instead, which is what the compiler's own grounding-remedy sentence does, would say two rows differing only in modifier dosage are the same row. That is a wrong answer where the drop is invisible, so it is the one thing this must not do.

stamped names the members the compiler fills rather than the author — variant_key on the two variant-keyed kinds. They stay in columns because they really are part of the key; they are flagged because an author cannot type one, and a surface that presented them as fillable columns would be describing a cell that does not exist in their CSV.

fallback is the key used when every column of columns is null, and one derived table has one: gene_validity.csv keys on the source's own assertion_id, and on the gene's grain when the source published none. It is () everywhere else, and a consumer that ignores it is right about seven of the eight tables — which is exactly why it is a field rather than a footnote (S51). A fallback is only meaningful where the primary key can be absent, so a kind declaring one whose primary column is required is a contradiction, and a test asserts it away rather than this dataclass — the check needs the model, which a TableKey does not carry.

derived_model_for

derived_model_for(csv_name: str) -> type[BaseModel]

The row model for a machine-produced CSV, or DraftError for a name this format does not read.

The counterpart to draft.model_for, which answers for authored kinds only. Between them they cover every CSV the compiler reads, and a tool that dispatches on a filename needs both: an author reading resolution.csv or frequencies.csv — files they must read and must never hand-finish — was getting "not an authored table of this format", which is true and useless (S47).

Raises DraftError rather than returning None so the failure matches model_for's, on the house rule that a generic rejection is a dead end where a specific one is a fix: the message names the authored route, because asking this function for variants.csv is the likely mistake and the answer to it is one call away.

Source code in compiler/src/just_dna_compiler/hints.py
def derived_model_for(csv_name: str) -> type[BaseModel]:
    """The row model for a machine-produced CSV, or `DraftError` for a name this format does not read.

    The counterpart to `draft.model_for`, which answers for authored kinds only. Between them they
    cover every CSV the compiler reads, and a tool that dispatches on a filename needs both: an author
    reading `resolution.csv` or `frequencies.csv` — files they must read and must never hand-finish —
    was getting *"not an authored table of this format"*, which is true and useless (S47).

    Raises `DraftError` rather than returning `None` so the failure matches `model_for`'s, on the house
    rule that a generic rejection is a dead end where a specific one is a fix: the message names the
    authored route, because asking this function for `variants.csv` is the likely mistake and the
    answer to it is one call away.
    """
    model = DERIVED_TABLE_MODELS.get(csv_name)
    if model is None:
        if csv_name in DRAFTABLE:
            raise DraftError(
                f"{csv_name!r} is an authored table, not a machine-produced one — "
                f"use model_for({csv_name!r}) instead."
            )
        raise DraftError(
            f"{csv_name!r} is not a machine-produced table of this format. "
            f"Known: {sorted(DERIVED_TABLE_MODELS)}"
        )
    return model

key_fields

key_fields(csv_name: str) -> TableKey | None

What makes two rows of this table kind the same row, or None for a kind with no declared key.

Answers for authored kinds and for the machine-produced tables alike, so a tool dispatching on a filename gets one route. None is the honest answer for a derived table that declares no key — withheld rather than invented, on the house three-valued rule.

Consistent with draft.natural_key by construction, not by agreement: both read the model's own _KEY_FIELDS, and a test pins natural_key(row) against getattr over these columns for every keyed kind. natural_key returns None for the binning kinds because their duplicate rule is overlap, not equality; this function still names their grouping columns and says rule="overlap", which is the distinction a tool needs to explain either one.

The machine-produced tables answer here too, and their key is the one the writing pass merges on (S51). It was readable only as a dict-key expression inside the body of each enricher pass, so every consumer re-deriving a sidecar had to guess it — and the guess a reporting consumer shipped, the required members of the fact-field tuple, was measurably coarse on two of the seven: it dropped disease_id from gene_validity.csv and variation_id from clinical_assertions.csv, the two tables where one subject legitimately carries several rows. Each pass now keys its existing dict off the model's declared tuple, so the pass and this surface cannot disagree.

Reading rule matters as much as reading the columns: resolution.csv's key is a subject, not a unique row, and reporting it as equality would be a wrong answer rather than a coarse one.

Source code in compiler/src/just_dna_compiler/hints.py
def key_fields(csv_name: str) -> TableKey | None:
    """What makes two rows of this table kind the same row, or `None` for a kind with no declared key.

    Answers for authored kinds and for the machine-produced tables alike, so a tool dispatching on a
    filename gets one route. `None` is the honest answer for a derived table that declares no key —
    withheld rather than invented, on the house three-valued rule.

    Consistent with `draft.natural_key` by construction, not by agreement: both read the model's own
    `_KEY_FIELDS`, and a test pins `natural_key(row)` against `getattr` over these columns for every
    keyed kind. `natural_key` returns `None` for the binning kinds because their duplicate rule is
    *overlap*, not equality; this function still names their grouping columns and says
    `rule="overlap"`, which is the distinction a tool needs to explain either one.

    **The machine-produced tables answer here too, and their key is the one the writing pass merges
    on** (S51). It was readable only as a dict-key expression inside the body of each enricher pass,
    so every consumer re-deriving a sidecar had to guess it — and the guess a reporting consumer
    shipped, *the required members of the fact-field tuple*, was measurably coarse on two of the seven:
    it dropped `disease_id` from `gene_validity.csv` and `variation_id` from `clinical_assertions.csv`,
    the two tables where one subject legitimately carries several rows. Each pass now keys its
    `existing` dict off the model's declared tuple, so the pass and this surface cannot disagree.

    Reading `rule` matters as much as reading the columns: `resolution.csv`'s key is a **subject**, not
    a unique row, and reporting it as `equality` would be a wrong answer rather than a coarse one.
    """
    try:
        model: type[BaseModel] = model_for(csv_name)
    except DraftError:
        model = derived_model_for(csv_name)
    declared = getattr(model, "_KEY_FIELDS", None)
    if not declared:
        return None
    columns = tuple(_authored_spelling(model, name) for name in declared)
    authored = set(authored_field_names(model))
    rule = getattr(model, "_KEY_RULE", None)
    if rule is None:
        rule = "overlap" if issubclass(model, MeasureBinRow) else "equality"
    return TableKey(
        columns=columns,
        rule=rule,
        stamped=tuple(c for c in columns if c not in authored),
        fallback=tuple(
            _authored_spelling(model, name) for name in getattr(model, "_KEY_FALLBACK_FIELDS", ())
        ),
    )

describe_table

describe_table(csv_name: str) -> dict[str, Any]

Everything a tool needs to help someone fill one table kind, without any input from them.

Columns with their types, requiredness and allowed values; the alternative identity groups; the natural key two rows are the same row by. All generated from the live models.

Whatever the whole-schema authoring_reference() says about a column, this says too — this is the command an author filling one table runs, so a fact that reaches only the other surface reaches the reader least likely to be reading it. Two were missing, both because this column dict is assembled here rather than shared with the other surface, and a guard compared only field_category. The per-member prose a closed vocabulary can carry: source_element's eight members printed with no way to tell which of them counts the reference element, which is the one question the vocabulary was added to answer (D1-4). And category: required alone is pydantic's two-way answer, false for MeasureBinRow.measure_kind and unresolved, which an author must still write out carrying their defaults. required stays beside it — a published key is not removed, and it is insufficient rather than wrong.

Source code in compiler/src/just_dna_compiler/hints.py
def describe_table(csv_name: str) -> dict[str, Any]:
    """Everything a tool needs to help someone fill one table kind, without any input from them.

    Columns with their types, requiredness and allowed values; the alternative identity groups; the
    natural key two rows are the same row by. All generated from the live models.

    **Whatever the whole-schema `authoring_reference()` says about a column, this says too** — this
    is the command an author filling *one* table runs, so a fact that reaches only the other surface
    reaches the reader least likely to be reading it. Two were missing, both because this column dict
    is assembled here rather than shared with the other surface, and a guard compared only
    `field_category`. The **per-member prose** a closed vocabulary can carry: `source_element`'s eight
    members printed with no way to tell which of them counts the reference element, which is the one
    question the vocabulary was added to answer (D1-4). And **`category`**: `required` alone is
    pydantic's two-way answer, false for `MeasureBinRow.measure_kind` and `unresolved`, which an
    author must still write out carrying their defaults. `required` stays beside it — a published key
    is not removed, and it is insufficient rather than wrong."""
    model = model_for(csv_name)
    vocabularies = field_vocabularies(model)
    requirements = authoring_requirements(csv_name)
    return {
        "csv": csv_name,
        "model": model.__name__,
        "columns": [
            {
                "name": name,
                "required": model.model_fields[name].is_required(),
                "category": field_category(model, name),
                "description": model.model_fields[name].description,
                **(
                    {
                        "vocabulary": vocabularies[name]["name"],
                        "options": vocabularies[name]["options"],
                        "closed": vocabularies[name]["closed"],
                        **(
                            {"notes": dict(vocabularies[name]["notes"])}
                            if vocabularies[name].get("notes")
                            else {}
                        ),
                    }
                    if name in vocabularies
                    else {}
                ),
                "redundancy_bearing": REDUNDANCY_BEARING.get(name),
            }
            for name in authored_field_names(model)
        ],
        "requirements": requirements,
        # The docstring above has promised "the natural key two rows are the same row by" since 0.5
        # and the dict did not carry it, so every caller hand-kept the string instead and one of them
        # went stale (S48). `None` for a kind with no declared key, never an empty tuple: no key and a
        # key of no columns are different claims.
        "key": (
            {
                "columns": list(table_key.columns),
                "rule": table_key.rule,
                "stamped": list(table_key.stamped),
                "fallback": list(table_key.fallback),
            }
            if (table_key := key_fields(csv_name)) is not None
            else None
        ),
    }

field_options

field_options(csv_name: str) -> list[FieldOptions]

The pick-lists for one table kind — the "valid options to select from" half of authoring.

Source code in compiler/src/just_dna_compiler/hints.py
def field_options(csv_name: str) -> list[FieldOptions]:
    """The pick-lists for one table kind — the "valid options to select from" half of authoring."""
    return [
        FieldOptions(
            column=column,
            vocabulary=marker["name"],
            options=marker["options"],
            closed=marker["closed"],
            notes=dict(marker["notes"]) if marker.get("notes") else None,
        )
        for column, marker in field_vocabularies(model_for(csv_name)).items()
    ]

inspect_rows

inspect_rows(csv_name: str, csv_text: str) -> HintReport

Inspect authored CSV text: validate each cell, preview the model's own normalizations, and report duplicate keys and bin incoherence. Pure — nothing on disk is read or written.

csv_text may carry a header or not; without one the model's authored column order is assumed, which is what a caller pasting a single row from a template has.

Source code in compiler/src/just_dna_compiler/hints.py
def inspect_rows(csv_name: str, csv_text: str) -> HintReport:
    """Inspect authored CSV text: validate each cell, preview the model's own normalizations, and
    report duplicate keys and bin incoherence. Pure — nothing on disk is read or written.

    `csv_text` may carry a header or not; without one the model's authored column order is assumed,
    which is what a caller pasting a single row from a template has.
    """
    model = model_for(csv_name)
    fieldnames = authored_field_names(model)
    rows, header, header_lines, ragged = _parse(csv_text, fieldnames)

    report = HintReport(
        csv_name=csv_name,
        header=header,
        rows_in=len(rows),
        csv_out=[],
        options=field_options(csv_name),
        requirements=authoring_requirements(csv_name),
    )

    # Reported BEFORE the per-row validation whose message it explains: a shifted row's type error
    # names the column to the right of the real mistake, so the author has to see the field count
    # first or they go and "fix" a cell that was correct.
    _report_ragged(ragged, len(header), report)

    emitted: list[dict[str, str]] = []
    parsed: list[BaseModel | None] = []
    for index, raw in enumerate(rows):
        cells = dict(raw)
        # The stub columns are handed to `_validate_row` rather than left implicit in the call order:
        # both layers see the same cell, and this is what lets the second one stay quiet about it.
        stubbed = _check_placeholders(index, cells, report)
        instance = _validate_row(index, cells, model, report, stubbed=stubbed)
        parsed.append(instance)
        if instance is not None:
            cells = _apply_normalizations(index, cells, instance, model, header, report)
        emitted.append(cells)

    _flag_advisory_columns(emitted, model, report)
    _check_duplicate_keys(parsed, report)
    _check_bins(parsed, model, report)
    _check_conclusions(parsed, model, header_lines, report)

    # Every row-scoped finding gains its file line number in one place, rather than each producer
    # remembering to pass it — the header offset is known here and nowhere else.
    report.findings = [
        f if f.row is None else replace(f, line=f.row + header_lines + 1) for f in report.findings
    ]
    report.csv_out = [",".join(header)] + [_render_line(cells, header) for cells in emitted]
    return report