Browse documentation
Docs/Operate

Run Grid in production

Deploy the Node API and private Runtime Host with explicit network, storage, credential, health, and upgrade boundaries.

Run Grid in production

Supported server shape

The active server output is the Linux Server container. Its normal process set is grid,grid-runtime-host:

  • Node serves the UI, public API, and control plane.
  • Runtime Host owns execution and resident model state.

Expose Node to clients. Keep Runtime Host private to the deployment network.

This guide intentionally does not include a pull or install command. The source proves that a Linux Server container exists, but it does not prove a public image coordinate, licensing path, or generally available distribution channel.

Network and credentials

Setting Purpose
PORT Node API port; default 3000
HOST Node bind address; default 0.0.0.0
GRID_RUNTIME_URL Node-to-Runtime-Host URL; local default http://127.0.0.1:3001
GRID_RUST_RUNTIME_HTTP_ADDR Runtime Host bind address
GRID_API_TOKEN Public API bearer token
GRID_RUNTIME_RPC_TOKEN Shared private Node/Runtime-Host credential
GRID_ADMIN_TOKEN Separate optional admin authority
GRID_INGRESS_TOKEN Connector-ingress authority

Give each authority a distinct secret. Do not place the Runtime Host or its private credential in browser-accessible configuration.

Durable storage

Inventory the state used by your enabled capabilities and mount it on durable storage:

GRID_BACKBONE_STATE_DIR=/data/grid/backbone
GRID_BACKBONE_MODE=disk
GRID_RUNTIME_HTTP_MODELS_DIR=/data/grid/http-models
GRID_MODEL_EVENT_ARCHIVE_DIR=/data/grid/model-events
GRID_MODEL_LEDGER_DIR=/data/grid/model-ledger
GRID_AGENT_RUN_LEDGER_DIR=/data/grid/agent-runs
GRID_RUST_AI_CACHE_DIR=/data/grid/ai-cache
GRID_TRAIN_CACHE_DIR=/data/grid/train-cache

Not every deployment enables every directory. The GRID_BACKBONE_STATE_DIR entry above describes ordinary local Backbone state; selected exact-node v2 and DDIL require additional isolated stores whose joint authority must be recovered together. Record which paths are durable, ephemeral, or regenerated before launch. A publicly reviewed universal backup/restore recipe is not part of this guide; prove recovery for the exact topology and packaged build before depending on snapshots.

Choose the Backbone topology before launch

Ordinary in-process and v1 Backbone remains the default. Exact-node HTTP v2 and distributed DDIL are explicit startup selections with additional identity, policy, durable-state, recovery, capacity, and readiness requirements. They do not change formula meaning and they do not turn ordinary Node API clients into v2 peers.

Read Choose a Backbone topology before setting GRID_BACKBONE_HTTP_TRANSPORT or GRID_BACKBONE_DISTRIBUTED_MODE. A partial selected configuration fails closed during startup; do not configure a fallback to the ordinary topology.

Start and readiness sequence

  1. Start Runtime Host on the private network.
  2. Start Node with GRID_RUNTIME_URL and the matching private RPC credential.
  3. Wait for /readyz to return 200 and status: "ok".
  4. Only then add the instance to the load balancer.

Health and metrics

  • GET /healthz: liveness; returns 503 only when unhealthy.
  • GET /readyz: traffic readiness; returns 503 when degraded, unhealthy, or draining.
  • GET /metrics: Prometheus text.
  • GET /api/health: compatibility-only; use the dedicated probes for new deployments.

Representative readiness response:

{
  "status": "ok",
  "uptime_seconds": 123,
  "probes": [
    {
      "name": "runtime-host",
      "kind": "ready",
      "ok": true,
      "metadata": { "status": "ok", "ready": true }
    }
  ]
}

Additional registered probes may appear; do not assume the array contains only one entry.

Selected v2 readiness additionally requires recovery of its durable authority. DDIL also requires a healthy peer-sync pass. A live Node process is never sufficient evidence that either selected topology is ready.

Upgrades

Before promoting a new build:

  1. Compare /api/capabilities product version and catalog digest in staging.
  2. Execute application compatibility tests against the documented public routes.
  3. Confirm the new build reads a restored copy of production state.
  4. Exercise a rollback using the deployment's actual storage and secrets mechanism.
  5. Drain traffic and confirm /readyz reflects draining before termination.
  6. For selected v2 or DDIL, restart from the staged durable state using the required resume declarations and verify selected authority health before promotion.

Production checklist

  • Node is the only publicly exposed Grid process.
  • Runtime Host has a private bind address and independent RPC credential.
  • API, admin, and ingress authorities use separate secrets.
  • Every required state directory is deliberately durable or deliberately ephemeral.
  • The load balancer uses /readyz; the process supervisor uses /healthz.
  • Metrics are scraped and alerts distinguish overload, unavailability, and timeout.
  • Upgrade, backup, restore, and rollback have been exercised on the exact release.