Browse documentation
Docs/Integrate

Coding corpus and coverage registries

Choose the right Grid training artifact, preserve its evidence boundary, and keep evaluation data private.

Coding corpus and coverage registries

The public corpus separates five kinds of evidence. They should not be flattened into one undifferentiated training file.

Artifact What it contains Safe use
Controlled language features 165 compiler-traced IDs with status, kind, a public-document path, implementation owners, and named conformance evidence Coverage planning, labels, and gap reports
Feature coverage Reviewed source-presence tags and an exact uncovered-feature inventory Corpus slicing and gap prioritization
Canonical programs Complete, compiler-checked models paired with authoring prompts Source-generation supervision
Verified behaviors Receipt-derived outputs tied to the exact parent source Expected-output supervision
Function cases 1,134 upstream executable cases in direct-call, formula, or full-source form Function retrieval, coverage analysis, and fixture-expectation training

The corpus manifest binds the supervised program and behavior files to their schemas, source export, compiler evidence, Runtime receipts, controlled feature registry, and license. The function-case manifest separately records the fixture tree, runner, catalog, counts, and per-file identities for the larger function suite.

Controlled features versus capability labels

A controlled language feature is a stable compiler-owned ID such as array_comprehensions, rule_blocks, or pipe_operator. Use these IDs when measuring syntax coverage or defining a capability-specific dataset slice.

Existing capabilities fields are broader review labels. Some describe a function family, fixture, product surface, domain, or provider rather than a language construct. For example, aggregations, fake-clock, and map-surface are useful labels, but they are not compiler feature IDs. Keep both taxonomies and map a program to controlled IDs only after reviewing the actual source; do not infer equivalence from similar spelling.

The registry has two explicit export boundaries. The full language export supplies corpus semantics and its feature matrix; the narrower feature-authority export supplies the compiler-generated JSON registry, traceability report, and every referenced public document. Both current snapshots come from the same clean Grid 0.65.0 candidate commit. Generation requires all 165 language-export IDs to remain present with the same status in the authority matrix and compiler JSON. There are currently no authorityAdditions. The generator copies the authority JSON metadata into a deterministic derivative while retaining separate byte and provenance identities for the two export paths.

The generated downstream coverage report is the reviewed example-to-feature map. Its tags mean only that a construct is visibly present in that checked-in source; they are independent of the upstream compiler metadata.

The raw feature-authority manifest, authority matrix, compiler registry, and traceability report expose the core inputs. The authority license must be byte-identical to the Grid 0.65.0 candidate corpus license. Every registry publicDoc is also included in the authority manifest and served beneath /corpus/language-feature-authority/ at its repository-relative path, for example the assignments authority page. Path labels alone are not accepted as provenance: the generator verifies every byte count and SHA-256 identity against the authority manifest. The portable downstream check authenticates those artifact bytes against the checked-in manifest; it does not authenticate Git ancestry or prove that the recorded commit and tree exist. Clean-checkout state and Git IDs are exporter assertions. For stronger provenance, reproduce the export from the recorded Grid Core commit and compare every identity.

Freshness is snapshot-level only. This page's last_verified date records the site integration check, not a semantic verification date for every feature. The authority manifest records one Grid version, source commit, tree, and deterministic commit timestamp (sourceCommittedAt) for the complete snapshot; the derived registry exposes that timestamp as the feature-taxonomy recency point. It does not provide per-feature introduction or last-verified commits.

Why function cases remain structured

Direct function cases carry runner-equivalent typed projections of arguments and expected matchers. Formula cases carry an authored expression, and source cases carry an authored multi-binding program; those strings are preserved verbatim.

The function-case manifest includes a hashed runner contract. It fixes NOW at the suite's declared instant, disables PROJ network access, records the RESULT = {formula} wrapper and RESULT lookup, and defines the default tolerance and error-code normalization used by the authoritative runner. Each record binds that contract digest. Error expectations expose every accepted normalized code, including the runner's NA/N/A equivalence, instead of asking consumers to guess matcher behavior.

It does not turn every direct call into invented Grid source. That conversion is safe only after a lossless Grid-literal serializer and a compiled-runtime equivalence check exist. Until then, use direct cases for typed invocation behavior and retrieval; use formula and source cases when authored syntax is required.

The import is marked preserved-not-reexecuted: its targets are fixture-authored expectations that the pinned upstream runner defines, not new Runtime observations captured by this site. Use verified-behaviors.jsonl when training specifically on receipt-observed outputs.

The current suite contains 1,035 direct calls, 93 formulas, and six full-source cases. Direct calls cover 655 canonical catalog functions after the two exercised aliases are resolved. Case records are training data under the included Grid Core license; review that license before distributing trained weights.

Consumer checklist

  1. Verify the artifact digest against its manifest before reading records.
  2. Validate every JSON or JSONL record with the published schema.
  3. Keep canonical-programs and verified-behaviors joined by lineage and in the same split.
  4. Preserve function-case mode, typed values, matcher tolerances and accepted codes, source file, source index, and runner-contract digest.
  5. Report coverage separately for controlled language features, canonical functions, diagnostics, runtime classes, and user tasks.
  6. Record every additional dataset in the complete training inventory used for evaluation leakage checks.

Treat record-level feature tags as reviewed source-presence annotations. They are useful coverage labels, but they do not by themselves prove execution, mapping completeness, or freshness after a source record changes.

Private evaluation boundary

No validation or test examples are published here. The public corpus schemas require split: "train"; private suites live outside the repository and use the public contracts under corpus/evaluation/.

Hold out independently authored prompts and unfamiliar capability combinations. Keep all derivatives of one source family in the same split. Freeze the test set before training, and retire a validation case after its prompt or answer has been inspected and used for tuning.

The boundary verifier rejects supplied private-suite paths inside the repository, including symlink re-entry, and checks the public corpus for non-training records. Its repository scan is structural: it blocks protected paths and unapproved files in corpus/evaluation/, but cannot recognize a secret deliberately renamed elsewhere. Contamination checks reduce accidental overlap, but hashes and token similarity cannot prove that a model never encountered semantically equivalent material elsewhere.

Each private case uses one fixed public system message, so case material cannot be hidden from the leakage comparison in that slot. A training inventory maps every logical record to exactly one declared source family; public corpus families are derived deterministically from their lineage IDs. The suite also selects exact clean compiler and Runtime builds from the reviewed provenance registries and pins generation settings. Its runner digest is a reproducibility assertion, not authentication; the operator still owns private runner trust.

Run npm run verify:evaluation-boundary for the public custody check. For a private run, pass the external suite manifest and complete training inventory to scripts/verify-evaluation-boundary.mjs. After evaluation, add an external detailed result to verify one complete validation or test split. Only test results may add a publishable aggregate. The verifier recomputes exact-source and exact-JSON decisions plus all totals and rates; execution-dependent scorers still rely on the suite-pinned trusted runner and therefore remain private rather than entering a public aggregate. Near-duplicate signals also block verification while no review-disposition protocol exists. If any task partition is smaller than the privacy cell size, every task breakdown is suppressed so the small residual cannot be inferred by subtraction.

Schemas