Skip to content

licensing.csv

One row is one data source, at one layer of the module, with the terms it was used under — keyed (source, layer). The same source consulted at two layers is two rows, and that is not duplication: the terms can genuinely differ.

It is the only table the compile licence gate reads, and the gate keys on this file and nothing else. A module drafted entirely from one source once carried no licensing.csv at all and compiled as though unrestricted. That is why every pass that consults a source writes its row — and why a pass that contributed nothing writes none.

It is a machine-produced fact table that a human starts from a template, which is the odd combination, and the reason it carries a placeholder guard: a scaffold's unreplaced <<REPLACE>> in a terms cell must be unable to compile, rather than publishing a module whose licence is a placeholder.

declared_use is a third axis with three states, not a mode. Unstated is not the same as commercial or non-commercial, and a source with unknown commercial terms warns — it never gates. A host's terms are also not its contents' terms: the host's row is the floor and each record may override it, so one permissive hosting statement does not license what it hosts.

layer names what a source fed, not which table it fed. Every authored table is annotation — the layer where a curated claim is expressed and a derivative work genuinely exists — whatever domain the table covers, so a PGS Catalog row filed against pgs.csv reads as annotation-layer content in the gate's own error message, and a row about a score panel is not a mis-filing. The other members are the machine-produced fact sidecars, which carry what a source reports. There is deliberately no per-table member: VALID_SOURCE_LAYERS is a wire vocabulary, so a new one is a format change rather than a label, and the last request for one was refused on the reporter's own argument (S82, RM147).

Either spelling is a key. licensing.csv is current and sources.csv deprecated until 1.0; write to the file you read, and a module carrying both is an error rather than a merge.

The artifact keeps the old name, and that is not an oversight — this page shows both. The identity card above says this table becomes sources.parquet, and the manifest block is ModuleManifest.sources. Both stay spelled that way for all of 0.x, because both are inside artifact.digest: renaming them would move the identity of every module that has one, which is a major-version change, while renaming the authored file is free. So the two halves moved at different times on purpose. An author writes licensing.csv; a consumer reading a compiled module looks for sources.parquet and manifest.sources, and finds nothing under the new name. Asking whether to rename the parquet too is asking to break every published digest.

Identity

Row model SourceRow (just_dna_format.sources)
Becomes sources.parquet
Authored or derived both — authored, and also written by an enricher pass
Draftable yes — draft can append rows
Drafted by no drafting provider targets this table — draft writes no row of it
Written and checked by every pass that consults a source — a pass contributing nothing writes no row, so this table has no single producer
Natural key source, layer
Fact signature yes — its sidecar carries one
In the attestation binding no
Also accepted as sources.csv — deprecated, removed at 1.0

Columns

Column Type Required Values Meaning
source str required The source identifier, joining to the open source column on the other fact tables (e.g. 'clinpgx', 'cpic', 'pharmvar', 'ensembl', 'gnomad'). Open, like every source column — a closed vocabulary here would have to be revised every time a link is added.
layer str required one of: annotation, clinical_assertion, expression_effect, frequency, gene_metrics, gene_validity, gwas_effect, literature, resolution Which layer this source contributed to (VALID_SOURCE_LAYERS). Only 'annotation' — the module's own authored tables — carries a derivative-work obligation; the fact sidecars report facts, not expression.
license str | None optional Licence identifier or name, e.g. 'CC-BY-SA-4.0'. Deliberately an open string rather than a closed SPDX vocabulary: SPDX evolves, and several of these sources are an SPDX licence PLUS a bespoke clause, which no single identifier expresses.
license_url str | None optional Where the terms were read from
license_sha256 str | None optional sha256: over the licence text as read, pinning the terms to the same moment as the data. Re-enriching recomputes it, so an upstream policy change becomes a finding rather than a silent pass.
attribution str | None optional The credit line the licence requires — one lookup, not a reconstruction
notice str | None optional Any use restriction the licence text states that is not captured by the flags below (e.g. PharmVar's 'not intended for direct diagnostic use or medical decision-making'). Carried with the data because that is where it is needed.
share_alike bool | None optional Whether the licence imposes a ShareAlike/copyleft obligation on derivatives. None means UNKNOWN, never false.
commercial_use bool | None optional Whether commercial use/sale is permitted. Orthogonal to share_alike — CC BY-SA, CC BY-NC and CC BY-NC-SA are three different combinations. None means UNKNOWN and must never be rendered as false: a source whose terms could not be established has not been shown to permit anything.
redistribution bool | None optional Whether the terms permit passing the data on to a third party at all. A THIRD axis, not a shade of commercial_use: an academic-use-only source (OMIM, dbNSFP) permits neither sale nor redistribution, while CC BY-NC forbids sale and expressly allows sharing — recording the first as merely non-commercial understates it, and a module that embeds it cannot be published at all, free or not. None means UNKNOWN, never false.
declared_use str | None optional one of: commercial, non_commercial, unstated The use declared when the data was fetched (VALID_DECLARED_USE). A claim about the user, not about the licence — which is why it is a separate axis from the three flags.
dataset str | None optional Which release the data came from, e.g. 'clinpgx_2026-07-05'
fetched_at str | None optional ISO-8601 UTC timestamp, second resolution (e.g. '2026-08-03T02:03:23Z'). Canonicalized on load; records when this row was last written by a pass, not when the source published anything
draft_digest str | None optional Set by a drafting provider: a hash of the drafted table projected onto the column a cross-check later compares against this source, so that check can tell a value still copied from the source from one a human has since edited (RM73). Recomputed, never trusted as a claim — an unrecognized or absent value simply means the check runs.

Generated from the row model at build time — reference.authoring_reference(), the same answer describe_table gives an authoring tool. Nothing on this page is hand-kept.