Coding corpus and coverage registries
Choose the right Grid training artifact, preserve its evidence boundary, and keep evaluation data private.
Coding corpus and coverage registries
The public corpus separates five kinds of evidence. They should not be flattened into one undifferentiated training file.
| Artifact | What it contains | Safe use |
|---|---|---|
| Controlled language features | 165 compiler-traced IDs with status, kind, a public-document path, implementation owners, and named conformance evidence | Coverage planning, labels, and gap reports |
| Feature coverage | Reviewed source-presence tags and an exact uncovered-feature inventory | Corpus slicing and gap prioritization |
| Canonical programs | Complete, compiler-checked models paired with authoring prompts | Source-generation supervision |
| Verified behaviors | Receipt-derived outputs tied to the exact parent source | Expected-output supervision |
| Function cases | 1,134 upstream executable cases in direct-call, formula, or full-source form | Function retrieval, coverage analysis, and fixture-expectation training |
The corpus manifest binds the supervised program and behavior files to their schemas, source export, compiler evidence, Runtime receipts, controlled feature registry, and license. The function-case manifest separately records the fixture tree, runner, catalog, counts, and per-file identities for the larger function suite.
Controlled features versus capability labels
A controlled language feature is a stable compiler-owned ID such as array_comprehensions, rule_blocks, or pipe_operator. Use these IDs when measuring syntax coverage or defining a capability-specific dataset slice.
Existing capabilities fields are broader review labels. Some describe a function family, fixture, product surface, domain, or provider rather than a language construct. For example, aggregations, fake-clock, and map-surface are useful labels, but they are not compiler feature IDs. Keep both taxonomies and map a program to controlled IDs only after reviewing the actual source; do not infer equivalence from similar spelling.
The registry has two explicit export boundaries. The full language export supplies corpus semantics
and its feature matrix; the narrower feature-authority export supplies the compiler-generated JSON
registry, traceability report, and every referenced public document. Both current snapshots come
from the same clean Grid 0.65.0 candidate commit. Generation requires all 165 language-export IDs to remain
present with the same status in the authority matrix and compiler JSON. There are currently no
authorityAdditions. The generator copies the authority JSON metadata into a deterministic
derivative while retaining separate byte and provenance identities for the two export paths.
The generated downstream coverage report is the reviewed example-to-feature map. Its tags mean only that a construct is visibly present in that checked-in source; they are independent of the upstream compiler metadata.
The raw feature-authority manifest,
authority matrix,
compiler registry, and
traceability report expose the core inputs. The
authority license must be byte-identical to the
Grid 0.65.0 candidate corpus license. Every registry publicDoc is also included in the authority manifest and
served beneath /corpus/language-feature-authority/ at its repository-relative path, for example
the assignments authority page.
Path labels alone are not accepted as provenance: the generator verifies every byte count and
SHA-256 identity against the authority manifest. The portable downstream check authenticates those
artifact bytes against the checked-in manifest; it does not authenticate Git ancestry or prove that
the recorded commit and tree exist. Clean-checkout state and Git IDs are exporter assertions. For
stronger provenance, reproduce the export from the recorded Grid Core commit and compare every
identity.
Freshness is snapshot-level only. This page's last_verified date records the site integration
check, not a semantic verification date for every feature. The authority manifest records one Grid
version, source commit, tree, and deterministic commit timestamp (sourceCommittedAt) for the
complete snapshot; the derived registry exposes that timestamp as the feature-taxonomy recency
point. It does not provide per-feature introduction or last-verified commits.
Why function cases remain structured
Direct function cases carry runner-equivalent typed projections of arguments and expected matchers. Formula cases carry an authored expression, and source cases carry an authored multi-binding program; those strings are preserved verbatim.
The function-case manifest includes a hashed runner contract. It fixes NOW at the suite's declared
instant, disables PROJ network access, records the RESULT = {formula} wrapper and RESULT lookup,
and defines the default tolerance and error-code normalization used by the authoritative runner.
Each record binds that contract digest. Error expectations expose every accepted normalized code,
including the runner's NA/N/A equivalence, instead of asking consumers to guess matcher behavior.
It does not turn every direct call into invented Grid source. That conversion is safe only after a lossless Grid-literal serializer and a compiled-runtime equivalence check exist. Until then, use direct cases for typed invocation behavior and retrieval; use formula and source cases when authored syntax is required.
The import is marked preserved-not-reexecuted: its targets are fixture-authored expectations that
the pinned upstream runner defines, not new Runtime observations captured by this site. Use
verified-behaviors.jsonl when training specifically on receipt-observed outputs.
The current suite contains 1,035 direct calls, 93 formulas, and six full-source cases. Direct calls cover 655 canonical catalog functions after the two exercised aliases are resolved. Case records are training data under the included Grid Core license; review that license before distributing trained weights.
Consumer checklist
- Verify the artifact digest against its manifest before reading records.
- Validate every JSON or JSONL record with the published schema.
- Keep
canonical-programsandverified-behaviorsjoined by lineage and in the same split. - Preserve function-case mode, typed values, matcher tolerances and accepted codes, source file, source index, and runner-contract digest.
- Report coverage separately for controlled language features, canonical functions, diagnostics, runtime classes, and user tasks.
- Record every additional dataset in the complete training inventory used for evaluation leakage checks.
Treat record-level feature tags as reviewed source-presence annotations. They are useful coverage labels, but they do not by themselves prove execution, mapping completeness, or freshness after a source record changes.
Private evaluation boundary
No validation or test examples are published here. The public corpus schemas require split: "train"; private suites live outside the repository and use the public contracts under corpus/evaluation/.
Hold out independently authored prompts and unfamiliar capability combinations. Keep all derivatives of one source family in the same split. Freeze the test set before training, and retire a validation case after its prompt or answer has been inspected and used for tuning.
The boundary verifier rejects supplied private-suite paths inside the repository, including symlink
re-entry, and checks the public corpus for non-training records. Its repository scan is structural:
it blocks protected paths and unapproved files in corpus/evaluation/, but cannot recognize a secret
deliberately renamed elsewhere. Contamination checks reduce accidental overlap, but hashes and token
similarity cannot prove that a model never encountered semantically equivalent material elsewhere.
Each private case uses one fixed public system message, so case material cannot be hidden from the leakage comparison in that slot. A training inventory maps every logical record to exactly one declared source family; public corpus families are derived deterministically from their lineage IDs. The suite also selects exact clean compiler and Runtime builds from the reviewed provenance registries and pins generation settings. Its runner digest is a reproducibility assertion, not authentication; the operator still owns private runner trust.
Run npm run verify:evaluation-boundary for the public custody check. For a private run, pass the
external suite manifest and complete training inventory to scripts/verify-evaluation-boundary.mjs.
After evaluation, add an external detailed result to verify one complete validation or test split.
Only test results may add a publishable aggregate. The verifier recomputes exact-source and exact-JSON
decisions plus all totals and rates; execution-dependent scorers still rely on the suite-pinned
trusted runner and therefore remain private rather than entering a public aggregate. Near-duplicate
signals also block verification while no review-disposition protocol exists. If any task partition
is smaller than the privacy cell size, every task breakdown is suppressed so the small residual
cannot be inferred by subtraction.
Schemas
- Canonical program schema
- Verified behavior schema
- Controlled language feature schema
- Controlled language feature ID schema
- Controlled language feature coverage schema
- Function case schema
- Private evaluation policy
- Private suite schema and case schema
- Training inventory schema
- Model adapter request and response schemas
- Private detailed result and publishable aggregate schemas