← All guided builds

Guided build · 20

Promote a model to production

Turn process, credential, storage, health, logging, capacity, compatibility, and rollback evidence into one fail-closed promotion gate.

You will finish with: A staged production-readiness packet with a deterministic local gate and an evidence checklist for the exact release.

35 minAdvancedSource checked for Grid 0.61.0Reviewed 2026-08-26
Related canonical example07-treasury-control-plane.grid
Get Grid
Deployment inventorydeployment-inventory.jsonA non-secret checklist of process, credential, and storage ownership.
Readiness gatereadiness-gate.mjsA small fail-closed gate for captured readiness responses.
Healthy readiness fixturereadyz-ok.jsonA deterministic passing readiness response.
Degraded readiness fixturereadyz-degraded.jsonA deliberate failing response for the promotion-block exercise.
On this page

What you will build

You will turn Grid's production boundary into an explicit promotion gate: process exposure, credential separation, durable-state ownership, readiness, metrics, compatibility evidence, a fail-closed readiness test, and a rollback drill. The included gate and responses let you test the decision logic without disrupting a server.

Passing the local fixture is not proof that a deployment is production-ready. Final completion requires evidence from the exact staged release and its actual storage, secrets, network, and rollback mechanisms.

Before you begin

You need:

  • an existing staging Grid Linux Server deployment;
  • operator access to its network, workload configuration, secret references, state mounts, logs, metrics, and load balancer;
  • curl, jq, and Node.js 20 or newer on an operator workstation;
  • the four assets above saved in one working directory;
  • a public API token for the staged deployment.
export GRID_URL="https://your-staging-node.example"
export GRID_TOKEN="replace-with-your-staging-api-token"

The supported server shape contains Node plus Runtime Host. Only Node is public. Runtime Host is a private execution plane reached through GRID_RUNTIME_URL and an independent RPC credential. This tutorial intentionally gives no image pull or installation command because a generally available distribution coordinate is not part of the current public contract.

Read Run Grid in production before changing a deployment.

1. Validate the non-secret deployment inventory

The included inventory names authorities and storage variables, but contains no secret values or environment-specific paths:

jq -e '
  .schemaVersion == 1 and
  .processes.node.public == true and
  .processes.runtimeHost.public == false and
  (.credentialVariables | unique | length) == 4 and
  .containsSecretValues == false
' deployment-inventory.json

jq '{processes, credentialVariables, state}' deployment-inventory.json

Deterministic local checkpoint: the first command prints true; Node is public, Runtime Host is private, four distinct credential variable names are present, and every state entry is review-required.

Copy this inventory into your private deployment records and replace each review-required classification with one of:

  • durable and backed up;
  • ephemeral and safely regenerated;
  • capability disabled and path unused.

Do not add token values to this public tutorial asset or to source control.

2. Exercise the readiness gate locally

Run the gate against the healthy fixture:

node readiness-gate.mjs readyz-ok.json

Expected checkpoint and exit status 0:

READINESS OK: 1 probe(s)

The gate requires top-level status: "ok", at least one probe, and every returned probe to have ok: true. It does not assume the server returns only a Runtime Host probe.

3. Make one deliberate readiness change

Run the same unchanged gate against the degraded fixture:

node readiness-gate.mjs readyz-degraded.json

Expected failure and nonzero exit status:

READINESS BLOCKED: status=degraded failing=runtime-host

This is the deliberate change: only the observed readiness input changed. The promotion decision failed closed.

Minimal repair for the local exercise is to use the current healthy response—not to edit the gate until degraded state passes:

node readiness-gate.mjs readyz-ok.json

In a real incident, repair the failing probe or withdraw the instance from traffic. Never rewrite a captured degraded response to look healthy.

4. Verify the live process boundary

From outside the deployment network, confirm the Node endpoint is reachable:

curl --fail-with-body "$GRID_URL/healthz" | jq
curl --fail-with-body "$GRID_URL/readyz" \
  | tee /tmp/grid-readyz-live.json \
  | jq '{status, probes}'
node readiness-gate.mjs /tmp/grid-readyz-live.json

/healthz is liveness only. It does not prove Runtime Host readiness. The load balancer must use /readyz; the process supervisor should use /healthz.

From an authorized position inside the private network, verify Runtime Host is not bound or routed publicly and that Node reaches it using the private GRID_RUNTIME_URL. Record the network-policy or service-discovery evidence in the deployment review; do not expose its RPC token in command output.

5. Prove credential separation

Inspect secret references, not secret values, in the workload configuration. Confirm distinct authorities for:

Variable Authority
GRID_API_TOKEN ordinary public API routes
GRID_RUNTIME_RPC_TOKEN private Node-to-Runtime-Host RPC
GRID_ADMIN_TOKEN optional administrative operations
GRID_INGRESS_TOKEN connector ingress

A missing authority can mean the capability is disabled or authentication is configured differently; document the exact release behavior. It must not mean one credential was copied into every role.

Failure and repair checkpoint: make one request to an ordinary protected API route using a deliberately invalid disposable value—not another real authority—and retain the 401 response. Repeat with the correct public API token and expect success:

curl -sS -o /tmp/grid-auth-failure.json -w '%{http_code}\n' \
  -H 'Authorization: Bearer deliberately-invalid' \
  "$GRID_URL/api/model-summaries"

curl --fail-with-body \
  -H "Authorization: Bearer $GRID_TOKEN" \
  "$GRID_URL/api/model-summaries" | jq 'type'

If the invalid request succeeds, stop: either authentication is intentionally disabled for that environment or the boundary is misconfigured. Resolve that policy before promotion.

6. Classify and test durable state

For every enabled path in the durable storage inventory, record:

  • the mounted path and owning process;
  • durable, ephemeral, or regenerated classification;
  • backup mechanism and last successful timestamp;
  • restore destination and last exercised restore;
  • recovery-point and recovery-time expectations.

At minimum, review backbone state, resident HTTP models, model event archives, model ledgers, agent-run ledgers, AI caches, and training caches. Not every capability requires every directory.

Restore a recent backup into an isolated staging location and start the exact candidate release against that restored copy using your deployment's supported procedure. A backup that has never been restored is not promotion evidence.

7. Capture health, metrics, and compatibility evidence

curl --fail-with-body "$GRID_URL/readyz" > /tmp/grid-readyz-before.json
curl --fail-with-body "$GRID_URL/metrics" > /tmp/grid-metrics-before.txt
curl --fail-with-body \
  -H "Authorization: Bearer $GRID_TOKEN" \
  "$GRID_URL/api/capabilities" \
  | tee /tmp/grid-capabilities.json \
  | jq '{schemaVersion, productVersion, catalogSha256}'

Store the product version and catalogSha256 digest with the staged release record. Compare them with the previous approved deployment, then run application compatibility tests against the public routes your clients actually use.

Metrics and alert rules must distinguish at least overload, unavailability, and timeout. Preserve Retry-After, public error code, readiness probes, model ID and source hash, connector identity, approximate failure time, and deployment revision when investigating. Never log bearer or RPC credentials.

8. Run the staged failure and rollback drill

Use your platform's approved staging controls; this guide does not invent a universal drain or restart command.

  1. Remove one staged instance from the load balancer.
  2. Confirm no new client traffic reaches it.
  3. Trigger the deployment's normal drain operation and confirm /readyz stops admitting traffic.
  4. Exercise a controlled Runtime Host unavailability on that withdrawn instance.
  5. Confirm Node reports degraded readiness and the gate from step 2 blocks promotion.
  6. Restore Runtime Host connectivity and wait for /readyz to return ok.
  7. Roll forward to the candidate build against restored state.
  8. Roll back using the exact prior build, storage, and secret references.
  9. Re-run public-route compatibility tests before returning the instance to service.

Capture commands, timestamps, build digests, probe responses, and outcomes in the private release record. The supplied degraded JSON demonstrates gate behavior only; it cannot substitute for this drill.

You are done when

  • Node is the only public process and Runtime Host uses a private RPC boundary.
  • The four authorities are separate or explicitly disabled, with no secret values in artifacts.
  • Every enabled state path has a tested durability and restore classification.
  • The exact staged /readyz response passes the unchanged gate, while degraded input fails it.
  • Product version, catalog SHA-256, metrics, compatibility results, and logs are attached to the release record.
  • Drain, failure, recovery, upgrade, and rollback have been exercised on the candidate release.

Use Troubleshoot Grid for status-specific diagnosis and API conventions for client retry and concurrency behavior.

Build statusReached the expected checkpoint?