Troubleshoot Grid
Diagnose readiness, authentication, overload, Runtime Host, Backbone v2 and DDIL, compilation, concurrency, realtime, and connector failures.
Troubleshoot Grid
Start with liveness and readiness
curl -fsS "$GRID_URL/healthz"
curl -fsS "$GRID_URL/readyz" | jq
curl -fsS "$GRID_URL/metrics" | head
/healthz proves the Node process can answer HTTP. It does not prove Runtime Host is ready. Use /readyz before routing traffic and inspect every failing probe.
API request fails with 401
Verify the request targets the Node API and uses the credential for that surface. Public API, admin, connector-ingress, and private Runtime Host credentials are separate.
API request fails with 429
Honor Retry-After. The request may have reached an endpoint-specific rate limit or the Runtime Host admission queue may be saturated. Add jitter and retry only when the operation is safe to repeat.
API request fails with 503
A response with code: RUNTIME_HOST_UNAVAILABLE means Node could not reach Runtime Host. Check the process, its private URL, and the shared runtime credential. Connector overload can also return 503 with overloaded: true, backpressure details, and Retry-After.
API request fails with 504
RUNTIME_TIMEOUT means the private runtime request exceeded its deadline. Inspect saturation, request size, and model cost before increasing client timeouts.
Model is not found
Confirm the model ID with GET /api/model-summaries. Ordinary missing-model errors map to 404 with RESOURCE_NOT_FOUND.
Source does not compile
Deploy and source-replacement errors can include a parseDiagnostic object with a Grid diagnostic code, message, source span, snippet, and raw compiler message. Preserve and display that diagnostic rather than reducing it to a generic 400.
Source or input write conflicts
Refetch source or symbol state, then repeat with the new baseSourceHash or expectedRequestRev. Do not overwrite blindly. Exact public code mapping can vary by route in this release, so use the status plus returned error body.
Realtime stream appears connected but stops changing
The server emits a named heartbeat every five seconds. Recreate a silent stream, request the default initial snapshot, and refetch visible symbols or ranges. There is no Last-Event-ID replay contract.
Connector event was accepted but nothing changed
Inspect unbound, applied, and stream.batchStatus:
unbound > 0: no matching delivery binding.applied == 0: no local resident target write was applied.batchStatus == "unsubscribed": no durable Stream subscriber is registered.deduped > 0orbatchStatus == "duplicate": the retry was already admitted.
Selected Backbone v2 or DDIL is not ready
Confirm that the deployment selected a complete topology before investigating individual requests. Exact-node v2 requires lifecycle-v3 receipts, mTLS identities, HMAC key epochs and policy, closed route and representation inventories, isolated durable stores, retention bounds, and correct first-activation or resume declarations. DDIL additionally requires its authority, journal, projection, monotonic trust history, and peer-sync configuration.
Use the selected health response to distinguish recovery, replay-window pressure, journal or key-retirement capacity, source-retirement backlog, authority uncertainty, and poison. A retryable 503 with retryAfterMs is different from a non-retryable 507. Poison or ambiguous durability can require restart recovery; capacity pressure may require time or an operator capacity change. Do not weaken the topology or clear durable state to make readiness green.
If a progressed v2 directory is configured with initialize-empty-v2, correct it to the required resume selection. Never reinitialize progressed authority.
Capture before escalating
Record:
- Grid product version and catalog digest from
/api/capabilities; - the complete
/readyzresponse; - HTTP status,
Retry-After, code, and explanatory text; - model ID and source hash, never secret values;
- connector source ID, sequence, and idempotency key;
- approximate failure time and deployment revision.