Choose a Backbone topology
Keep ordinary Backbone as the default or deliberately admit exact-node HTTP v2 and DDIL with the required state, trust, recovery, and readiness controls.
Choose a Backbone topology
Grid 0.65.0 has three startup-frozen Backbone shapes. They share one Runtime Host and one LIR evaluator, but they have different authority, cost, and operating contracts.
| Shape | Selection | Use it when |
|---|---|---|
| Ordinary Backbone | GRID_BACKBONE_HTTP_TRANSPORT absent or v1 |
The deployment needs local durable jobs, control records, ordinary reactive delivery, current peer sync, or Node connector ingress. This remains the default. |
| Exact-node HTTP v2 | GRID_BACKBONE_HTTP_TRANSPORT=v2-exact-node-mtls plus lifecycle-v3 receipt selection |
Exact admitted peers require TLS 1.3 mutual authentication, field-complete HMAC requests, durable replay and outbound recovery, and authenticated retirement. |
| Distributed DDIL authority | HTTP v2 plus GRID_BACKBONE_DISTRIBUTED_MODE=ddil-authority-v1 |
A deployment supplies monotonic trust history and needs bounded disconnected authority exchange and final sticky-input convergence. |
The selection is coherent or startup fails. Grid does not partially start a requested v2 or DDIL topology, and it does not fall back to v1 when a selected assurance capability is unavailable.
Stay on the ordinary topology unless the contract requires more
An application posting model or connector requests to the Node API is not an exact-node peer. Existing ordinary clients keep their public API boundary. Adding TLS in front of an ordinary helper also does not grant the replay, identity, durable receipt, or retirement semantics of v2.
The ordinary topology avoids v2 policy parsing, TLS identities, replay journals, outbound queues, DDIL state, and their workers. Do not select v2 only because the package contains it.
Admit exact-node v2 deliberately
A selected v2 deployment needs all of the following before its listener can become ready:
- exact peer identities, TLS 1.3 mutual-authentication material, key epochs, and field-complete HMAC policy;
- lifecycle-v3 receipts and a closed inventory of admitted routes and representations;
- isolated durable directories for replay, semantic receipts, retained core, outbound state, and joint-generation authority;
- startup-frozen capacity and retention budgets;
- a one-time first-activation declaration for an empty state directory, followed by resume declarations on every later startup;
- recovery and operator procedures for poison, capacity refusal, key retirement, source retirement, and uncertain durability.
The core selectors include:
GRID_BACKBONE_MODE=disk
GRID_BACKBONE_DELIVERIES_FSYNC_INTERVAL_MS=0
GRID_BACKBONE_FROZEN_FANOUT_RECEIPTS=1
GRID_BACKBONE_FROZEN_FANOUT_CUTOVER=post-drain-v3
GRID_BACKBONE_HTTP_TRANSPORT=v2-exact-node-mtls
GRID_BACKBONE_HTTP_V2_REPLAY_CUTOVER=initialize-empty-v2|resume-v2
GRID_BACKBONE_HTTP_V2_OUTBOUND_CUTOVER=initialize-empty-v2|resume-v2
These lines are a selection outline, not a complete configuration. Use the exact 0.65.0 Runtime Environment contract for the required paths, inventories, policies, certificates, keys, and bounds. Never reuse initialize-empty-v2 for a progressed directory; startup rejects reinitialization because it would discard durable authority.
Add DDIL only with external authority
DDIL is an additional authority selection over v2. It requires separate authority, journal, and application-projection state plus a deployment-owned monotonic trust service and retained history. Its final application shape is intentionally narrow: a canonical sticky-input upsert or tombstone for one named input. It does not admit source replacement, computed targets, Frames, connector effects, or arbitrary mutation.
Mission dispatch and defense control are separately constructed owners. The v2 or DDIL selector does not create an operator boundary, identity provider, HSM, key service, conflict policy, receiver, or external-effect guarantee for them.
Read readiness as selected authority
For ordinary deployments, keep using the Node /healthz, /readyz, and /metrics boundaries described in Run Grid in production. For selected v2 and DDIL, readiness also depends on recovery of the selected durable authority. DDIL additionally requires a healthy peer-sync pass.
Treat refusal classes differently:
- Replay-window saturation can return retryable
503withretryAfterMsuntil the durable horizon advances. - Replay-journal or key-retirement capacity can return non-retryable
507and needs operator capacity recovery. - Ambiguous durability, corruption, rollback, or a post-durability publication failure can poison the selected owner and require restart recovery.
Follow the returned code and selected health state. Restarting every capacity refusal can delay recovery and does not create authority.
Roll out and recover
Before first activation, prove configuration rejection, fresh initialization, restart resume, certificate and key rotation, replay refusal, capacity recovery, and restored-state readiness on the exact packaged build. A backup is useful only if restoration preserves the joint durable authority and the deployment can prove that the restored generation is the one it is entitled to resume.
Do not claim automatic failover, consensus, or exactly-once external effects. The deployment still owns trust anchors, stable identities, retention, conflict policy, operator authorization, and receiver idempotency.