What you will build
You will turn Grid's production boundary into an explicit promotion gate: process exposure, credential separation, durable-state ownership, readiness, metrics, compatibility evidence, a fail-closed readiness test, and a rollback drill. The included gate and responses let you test the decision logic without disrupting a server.
Passing the local fixture is not proof that a deployment is production-ready. Final completion requires evidence from the exact staged release and its actual storage, secrets, network, and rollback mechanisms.
Before you begin
You need:
- an existing staging Grid Linux Server deployment;
- operator access to its network, workload configuration, secret references, state mounts, logs, metrics, and load balancer;
curl,jq, and Node.js 20 or newer on an operator workstation;- the four assets above saved in one working directory;
- a public API token for the staged deployment.
export GRID_URL="https://your-staging-node.example"
export GRID_TOKEN="replace-with-your-staging-api-token"
The supported server shape contains Node plus Runtime Host. Only Node is public. Runtime Host is a private execution plane reached through GRID_RUNTIME_URL and an independent RPC credential. This tutorial intentionally gives no image pull or installation command because a generally available distribution coordinate is not part of the current public contract.
Read Run Grid in production before changing a deployment.
1. Validate the non-secret deployment inventory
The included inventory names authorities and storage variables, but contains no secret values or environment-specific paths:
jq -e '
.schemaVersion == 1 and
.processes.node.public == true and
.processes.runtimeHost.public == false and
(.credentialVariables | unique | length) == 4 and
.containsSecretValues == false
' deployment-inventory.json
jq '{processes, credentialVariables, state}' deployment-inventory.json
Deterministic local checkpoint: the first command prints true; Node is public, Runtime Host is private, four distinct credential variable names are present, and every state entry is review-required.
Copy this inventory into your private deployment records and replace each review-required classification with one of:
- durable and backed up;
- ephemeral and safely regenerated;
- capability disabled and path unused.
Do not add token values to this public tutorial asset or to source control.
2. Exercise the readiness gate locally
Run the gate against the healthy fixture:
node readiness-gate.mjs readyz-ok.json
Expected checkpoint and exit status 0:
READINESS OK: 1 probe(s)
The gate requires top-level status: "ok", at least one probe, and every returned probe to have ok: true. It does not assume the server returns only a Runtime Host probe.
3. Make one deliberate readiness change
Run the same unchanged gate against the degraded fixture:
node readiness-gate.mjs readyz-degraded.json
Expected failure and nonzero exit status:
READINESS BLOCKED: status=degraded failing=runtime-host
This is the deliberate change: only the observed readiness input changed. The promotion decision failed closed.
Minimal repair for the local exercise is to use the current healthy response—not to edit the gate until degraded state passes:
node readiness-gate.mjs readyz-ok.json
In a real incident, repair the failing probe or withdraw the instance from traffic. Never rewrite a captured degraded response to look healthy.
4. Verify the live process boundary
From outside the deployment network, confirm the Node endpoint is reachable:
curl --fail-with-body "$GRID_URL/healthz" | jq
curl --fail-with-body "$GRID_URL/readyz" \
| tee /tmp/grid-readyz-live.json \
| jq '{status, probes}'
node readiness-gate.mjs /tmp/grid-readyz-live.json
/healthz is liveness only. It does not prove Runtime Host readiness. The load balancer must use /readyz; the process supervisor should use /healthz.
From an authorized position inside the private network, verify Runtime Host is not bound or routed publicly and that Node reaches it using the private GRID_RUNTIME_URL. Record the network-policy or service-discovery evidence in the deployment review; do not expose its RPC token in command output.
5. Prove credential separation
Inspect secret references, not secret values, in the workload configuration. Confirm distinct authorities for:
| Variable | Authority |
|---|---|
GRID_API_TOKEN |
ordinary public API routes |
GRID_RUNTIME_RPC_TOKEN |
private Node-to-Runtime-Host RPC |
GRID_ADMIN_TOKEN |
optional administrative operations |
GRID_INGRESS_TOKEN |
connector ingress |
A missing authority can mean the capability is disabled or authentication is configured differently; document the exact release behavior. It must not mean one credential was copied into every role.
Failure and repair checkpoint: make one request to an ordinary protected API route using a deliberately invalid disposable value—not another real authority—and retain the 401 response. Repeat with the correct public API token and expect success:
curl -sS -o /tmp/grid-auth-failure.json -w '%{http_code}\n' \
-H 'Authorization: Bearer deliberately-invalid' \
"$GRID_URL/api/model-summaries"
curl --fail-with-body \
-H "Authorization: Bearer $GRID_TOKEN" \
"$GRID_URL/api/model-summaries" | jq 'type'
If the invalid request succeeds, stop: either authentication is intentionally disabled for that environment or the boundary is misconfigured. Resolve that policy before promotion.
6. Classify and test durable state
For every enabled path in the durable storage inventory, record:
- the mounted path and owning process;
- durable, ephemeral, or regenerated classification;
- backup mechanism and last successful timestamp;
- restore destination and last exercised restore;
- recovery-point and recovery-time expectations.
At minimum, review backbone state, resident HTTP models, model event archives, model ledgers, agent-run ledgers, AI caches, and training caches. Not every capability requires every directory.
Restore a recent backup into an isolated staging location and start the exact candidate release against that restored copy using your deployment's supported procedure. A backup that has never been restored is not promotion evidence.
7. Capture health, metrics, and compatibility evidence
curl --fail-with-body "$GRID_URL/readyz" > /tmp/grid-readyz-before.json
curl --fail-with-body "$GRID_URL/metrics" > /tmp/grid-metrics-before.txt
curl --fail-with-body \
-H "Authorization: Bearer $GRID_TOKEN" \
"$GRID_URL/api/capabilities" \
| tee /tmp/grid-capabilities.json \
| jq '{schemaVersion, productVersion, catalogSha256}'
Store the product version and catalogSha256 digest with the staged release record. Compare them with the previous approved deployment, then run application compatibility tests against the public routes your clients actually use.
Metrics and alert rules must distinguish at least overload, unavailability, and timeout. Preserve Retry-After, public error code, readiness probes, model ID and source hash, connector identity, approximate failure time, and deployment revision when investigating. Never log bearer or RPC credentials.
8. Run the staged failure and rollback drill
Use your platform's approved staging controls; this guide does not invent a universal drain or restart command.
- Remove one staged instance from the load balancer.
- Confirm no new client traffic reaches it.
- Trigger the deployment's normal drain operation and confirm
/readyzstops admitting traffic. - Exercise a controlled Runtime Host unavailability on that withdrawn instance.
- Confirm Node reports degraded readiness and the gate from step 2 blocks promotion.
- Restore Runtime Host connectivity and wait for
/readyzto returnok. - Roll forward to the candidate build against restored state.
- Roll back using the exact prior build, storage, and secret references.
- Re-run public-route compatibility tests before returning the instance to service.
Capture commands, timestamps, build digests, probe responses, and outcomes in the private release record. The supplied degraded JSON demonstrates gate behavior only; it cannot substitute for this drill.
You are done when
- Node is the only public process and Runtime Host uses a private RPC boundary.
- The four authorities are separate or explicitly disabled, with no secret values in artifacts.
- Every enabled state path has a tested durability and restore classification.
- The exact staged
/readyzresponse passes the unchanged gate, while degraded input fails it. - Product version, catalog SHA-256, metrics, compatibility results, and logs are attached to the release record.
- Drain, failure, recovery, upgrade, and rollback have been exercised on the candidate release.
Use Troubleshoot Grid for status-specific diagnosis and API conventions for client retry and concurrency behavior.