just_dna_compiler.draft¶
just_dna_compiler.draft ¶
Drafting — append validated rows into an authored CSV without ever clobbering one (0.5).
The mechanism half of the drafting story. A source that publishes a table (CPIC's star-allele
definitions, a ClinVar gene slice) can hand its rows to append_rows and they land in the authored
CSV a human then owns; the fetching half lives in just-dna-enricher, which is the only tier allowed
to go to the network. This module is pure: rows in, CSV appended, report out.
It lives in the compiler because the compiler already does this. reverse_module writes authored
CSVs from parquet, and _TABLE_DUPE_KEYS already defines what makes two rows the same row. Building a
second, parallel notion of either in the enricher is how they drift.
Append-only, at row granularity. A file-level "refuse if it exists" rule self-defuses after the first gene and makes a multi-gene module unbuildable, so the granularity is the row:
- a row whose natural key is not present is appended;
- a row whose key IS present is never rewritten — it is reported as
already_present, or asdifferswhen the incoming cells disagree with the authored ones, and the run continues.
That is the line between this and the enricher-co-authoring idea the roadmap parks: appending rows a
source publishes leaves content_signature a function of the authored bytes exactly as before, while
mutating a cell a human wrote would make the content identity depend on a network fetch. Drift on
rows that already exist is the cross-check passes' job to report, not this module's to fix.
Existing cells are never touched. New rows are appended to the open file; the table is rewritten
whole only when the header must grow or when placement is delegated, and a rewrite re-reads existing
rows as text (_render_existing), so it can never reformat 1.0 into 1. Rows gain an empty
cell in a new column — value-neutral, and content_signature ignores unset optional columns by
construction.
Where a row lands: the end by default, or its group when asked (group_by, 0.5.1). Appending at
the end makes a re-drafted file chronological rather than logical, so a gene's rows scatter as a
module grows; delegated insertion puts a row after the last member of its block instead. The tool
chooses where, never what — there is deliberately no at=N, because an index is the caller
deciding and buys nothing a text editor does not. Moving a row's line number is safe: a pure reorder
moves artifact.digest but leaves content_signature untouched (it is order-independent), the
round-trip fixed point still holds, and duplicate keys are rejected so order can disambiguate
nothing. DraftReport.shifted names every row whose line moved.
DraftError ¶
Bases: RuntimeError
A draft could not be attempted — an unknown table kind, or an unreadable existing file.
RowOutcome
dataclass
¶
What happened to one incoming row.
PartialRow
dataclass
¶
PartialRow(
model: type[BaseModel],
cells: dict[str, Any],
stubbed: tuple[str, ...],
match_on: tuple[str, ...],
)
Cells a source published, plus the columns only a human can decide.
match_on is what makes a re-draft safe. A partial row cannot be keyed the usual way — its
natural key runs through a column that is still a placeholder — so sameness is decided on the
columns that are filled. Once the human replaces the stub, a re-run matches on those same
columns and reports already_present instead of appending the stub again.
rendered ¶
The CSV cells for this row: what the source gave, the placeholder where it could not.
Source code in compiler/src/just_dna_compiler/draft.py
validation_errors ¶
What is wrong with the cells the source did publish.
The stubbed columns are validated by omission: the row is built without them and any error located on one is discarded. That avoids the alternative — a per-column table of plausible dummy values — which would be exactly the hand-kept list this module keeps abolishing.
Source code in compiler/src/just_dna_compiler/draft.py
model_for ¶
The row model for an authored CSV name, or DraftError for one this format does not define.
Source code in compiler/src/just_dna_compiler/draft.py
natural_key ¶
The identity that decides whether two rows are the same row, or None when the kind has none.
Reuses the compiler's own _TABLE_DUPE_KEYS so an append can never produce a row the compiler
would then reject as a duplicate. The binning kinds return None on purpose: their duplicate rule
is overlap, not equality (validate_bins), and two bins can conflict while sharing no key — so
they are appended and the overlap is caught at compile, where it can actually be judged.
Source code in compiler/src/just_dna_compiler/draft.py
blank_template ¶
A header-only CSV for one table kind, with columns in the model's own field order.
Generated from the model, never a hand-kept list — the same drift-proof route
reference.authoring_reference() takes. Starting a new table currently means copying a header out
of the docs, which is exactly the thing that goes stale.
Compiler-managed columns are left out (base.authored_field_names): variant_key and
authored_ident are stamped at load and never written back by reverse_module, so offering them
would invite an author to fill a column the compiler overwrites — and authored_ident is a list,
so a rendered cell would not even reload.
Source code in compiler/src/just_dna_compiler/draft.py
required_fields ¶
The columns an author must fill for this kind — everything without a default.
Field-local only, and deliberately unchanged (it is shipped API). It does not answer "what
must I fill to get a valid row": see authoring_requirements, which adds the columns that have a
default but reject an empty cell, and the alternative identity groups a model validator enforces.
Source code in compiler/src/just_dna_compiler/draft.py
authoring_requirements ¶
What an author actually has to supply for one table kind, machine-readably.
Three parts, because requiredness has three shapes here and only the first is visible to
pydantic's is_required():
always— columns with no default;any_of— alternative identity groups, any ONE of which satisfies the row (rsid, orchrom+start). Enforced by a model validator, so no per-field flag can express it;defaulted—{column: rendered default}for columns that have a default but reject an empty cell. A template writes these out; seefield_category.
Source code in compiler/src/just_dna_compiler/draft.py
stub_template ¶
A header plus rows stub rows: the sentinel where a human must decide, defaults written out.
The point of the sentinel over a blank cell is that an unreplaced stub cannot compile —
vocab.reject_template_placeholders refuses it by name and row, in both modes. A blank required
cell would also fail, but a blank optional one would silently become a real row asserting
nothing, and the binning kinds' unresolved sentinel could not be used for this at all: it is
real data designed to compile.
For a binning kind the mandatory unresolved companion row is emitted too, because a binning
table without one is incomplete by contract and the author would otherwise meet that rule as a
compile error about a row they never wrote.
Source code in compiler/src/just_dna_compiler/draft.py
append_rows ¶
append_rows(
spec_dir: Path,
csv_name: str,
rows: list[BaseModel],
*,
group_by: Sequence[str] = (),
dry_run: bool = False,
before_commit: Callable[[], None] | None = None,
) -> DraftReport
Append rows to spec_dir/csv_name, skipping any whose natural key is already there.
Creates the file when absent. Returns a DraftReport describing every incoming row, so a caller
can show what a run would do (dry_run=True writes nothing and reports the same thing).
group_by turns the append into a delegated insertion: each new row lands after the last
existing row sharing those columns (its gene, its haplotype) instead of at the end of the file.
The tool picks the position; the caller never supplies an index. Existing cells are still never
rewritten — only their line number can move, and DraftReport.shifted names every row it did.
before_commit is handed straight to layout.atomic_writer, so it runs after this table's bytes
are down and before the rename (RM232). It is what lets a drafter's licence row land inside
the commit of the table it licenses, the way an enrichment pass's already does (S98, RM231): the
enricher records the row only once a table has rows in it, so without this the drafted rows were
on disk first and a refused merge left them with nothing to say what licensed them.
It fires only when this call actually writes, which is the grain the caller wants: a run whose
every row is already_present or differs adds nothing to this table and reaches no writer, so
no licence row is claimed for a table this run did not change. A drafter appending several tables
passes the same callable to each — the merge is never-clobber, so N firings record one row, and
binding it to only the first or only the last would leave a table committed unlicensed whenever
that particular one was the no-op.
Source code in compiler/src/just_dna_compiler/draft.py
384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 | |
append_partial_rows ¶
append_partial_rows(
spec_dir: Path,
csv_name: str,
partials: list[PartialRow],
*,
group_by: Sequence[str] = (),
dry_run: bool = False,
before_commit: Callable[[], None] | None = None,
) -> DraftReport
Append rows a source could only partly fill, leaving the rest as stubs a human must replace.
Same promises as append_rows: never rewrites a cell, never removes a row, and a row already
covered is reported rather than duplicated. The difference is only how sameness is decided — see
PartialRow.match_on — because a row whose key column is a placeholder has no usable key yet.
before_commit behaves exactly as it does in append_rows, and for the same reason (RM232).
Source code in compiler/src/just_dna_compiler/draft.py
593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 | |
group_of ¶
The block a row belongs to, or None when it declares no group (then it goes to the end).
Source code in compiler/src/just_dna_compiler/draft.py
place_rows ¶
place_rows(
existing: list[dict[str, str]],
incoming: list[dict[str, str]],
group_by: Sequence[str],
) -> tuple[list[dict[str, str]], list[int]]
Merge incoming into existing, each row after the last member of its group.
Returns the final row list and the original indices of existing rows whose line moved. With no
group_by this is a plain append and nothing shifts. Incoming rows keep their relative order
within a group, so a provider's own ordering survives.