Run Grid in production
Deploy the Node API and private Runtime Host with explicit network, storage, credential, health, and upgrade boundaries.
Run Grid in production
Supported server shape
The active server output is the Linux Server container. Its normal process set is grid,grid-runtime-host:
- Node serves the UI, public API, and control plane.
- Runtime Host owns execution and resident model state.
Expose Node to clients. Keep Runtime Host private to the deployment network.
This guide intentionally does not include a pull or install command. The source proves that a Linux Server container exists, but it does not prove a public image coordinate, licensing path, or generally available distribution channel.
Network and credentials
| Setting | Purpose |
|---|---|
PORT |
Node API port; default 3000 |
HOST |
Node bind address; default 0.0.0.0 |
GRID_RUNTIME_URL |
Node-to-Runtime-Host URL; local default http://127.0.0.1:3001 |
GRID_RUST_RUNTIME_HTTP_ADDR |
Runtime Host bind address |
GRID_API_TOKEN |
Public API bearer token |
GRID_RUNTIME_RPC_TOKEN |
Shared private Node/Runtime-Host credential |
GRID_ADMIN_TOKEN |
Separate optional admin authority |
GRID_INGRESS_TOKEN |
Connector-ingress authority |
Give each authority a distinct secret. Do not place the Runtime Host or its private credential in browser-accessible configuration.
Durable storage
Inventory the state used by your enabled capabilities and mount it on durable storage:
GRID_BACKBONE_STATE_DIR=/data/grid/backbone
GRID_BACKBONE_MODE=disk
GRID_RUNTIME_HTTP_MODELS_DIR=/data/grid/http-models
GRID_MODEL_EVENT_ARCHIVE_DIR=/data/grid/model-events
GRID_MODEL_LEDGER_DIR=/data/grid/model-ledger
GRID_AGENT_RUN_LEDGER_DIR=/data/grid/agent-runs
GRID_RUST_AI_CACHE_DIR=/data/grid/ai-cache
GRID_TRAIN_CACHE_DIR=/data/grid/train-cache
Not every deployment enables every directory. The GRID_BACKBONE_STATE_DIR entry above describes ordinary local Backbone state; selected exact-node v2 and DDIL require additional isolated stores whose joint authority must be recovered together. Record which paths are durable, ephemeral, or regenerated before launch. A publicly reviewed universal backup/restore recipe is not part of this guide; prove recovery for the exact topology and packaged build before depending on snapshots.
Choose the Backbone topology before launch
Ordinary in-process and v1 Backbone remains the default. Exact-node HTTP v2 and distributed DDIL are explicit startup selections with additional identity, policy, durable-state, recovery, capacity, and readiness requirements. They do not change formula meaning and they do not turn ordinary Node API clients into v2 peers.
Read Choose a Backbone topology before setting GRID_BACKBONE_HTTP_TRANSPORT or GRID_BACKBONE_DISTRIBUTED_MODE. A partial selected configuration fails closed during startup; do not configure a fallback to the ordinary topology.
Start and readiness sequence
- Start Runtime Host on the private network.
- Start Node with
GRID_RUNTIME_URLand the matching private RPC credential. - Wait for
/readyzto return200andstatus: "ok". - Only then add the instance to the load balancer.
Health and metrics
GET /healthz: liveness; returns503only when unhealthy.GET /readyz: traffic readiness; returns503when degraded, unhealthy, or draining.GET /metrics: Prometheus text.GET /api/health: compatibility-only; use the dedicated probes for new deployments.
Representative readiness response:
{
"status": "ok",
"uptime_seconds": 123,
"probes": [
{
"name": "runtime-host",
"kind": "ready",
"ok": true,
"metadata": { "status": "ok", "ready": true }
}
]
}
Additional registered probes may appear; do not assume the array contains only one entry.
Selected v2 readiness additionally requires recovery of its durable authority. DDIL also requires a healthy peer-sync pass. A live Node process is never sufficient evidence that either selected topology is ready.
Upgrades
Before promoting a new build:
- Compare
/api/capabilitiesproduct version and catalog digest in staging. - Execute application compatibility tests against the documented public routes.
- Confirm the new build reads a restored copy of production state.
- Exercise a rollback using the deployment's actual storage and secrets mechanism.
- Drain traffic and confirm
/readyzreflects draining before termination. - For selected v2 or DDIL, restart from the staged durable state using the required resume declarations and verify selected authority health before promotion.
Production checklist
- Node is the only publicly exposed Grid process.
- Runtime Host has a private bind address and independent RPC credential.
- API, admin, and ingress authorities use separate secrets.
- Every required state directory is deliberately durable or deliberately ephemeral.
- The load balancer uses
/readyz; the process supervisor uses/healthz. - Metrics are scraped and alerts distinguish overload, unavailability, and timeout.
- Upgrade, backup, restore, and rollback have been exercised on the exact release.