Browse documentation
Docs/Functionality

Train and Use Calibrated Predictions

Use the Predict surface when a workbook needs a trained prediction function. For XGBoost regression, you can additionally request a calibrated upper P90: an endpoint targeting 90% coverage…

Train and Use Calibrated Predictions

Use the Predict surface when a workbook needs a trained prediction function. For XGBoost regression, you can additionally request a calibrated upper P90: an endpoint targeting 90% coverage across comparable future rows. It is not a 90% guarantee for each individual row, a confidence interval for a parameter, or a complete probability distribution.

Train from a workbook

  1. Save the workbook and add a Predict surface. Give the trained model a name, select its feature data and target column, and confirm which columns are inputs rather than labels. Keep the feature order consistent at inference.
  2. Select XGBoost, then Regression. Check Calibrate an automatic upper P90 bound if you need the endpoint. Leaving it unchecked keeps ordinary point-only training. The checkbox selects automatic calibration with a seeded random holdout and probability 0.9.
  3. Train and inspect the completed job, metrics, and predictive-uncertainty information. Automatic calibration can produce a usable point model without a bound; a successful training job alone does not establish P90 capability.
  4. In the deployment view, bind a worksheet feature range to obtain copyable formulas. The upper-P90 formula appears only when the selected immutable registry version actually admits that endpoint. File-only training does not invent a worksheet feature range for you.
  5. Evaluate the returned exact model@version before changing the registry's pin. Unversioned calls use the pinned version, which may still be an older point-only model. A pin change is an explicit registry operation, not a consequence of asking for a bound.

For example, once weekly-demand@2 exists and expects the three features in A2:C2 in that order:

forecast = PREDICT("weekly-demand@2", A2:C2)
upper_p90 = PREDICTION_BOUND("weekly-demand@2", A2:C2, 0.9, "upper")
calibration = MODEL_UNCERTAINTY("weekly-demand@2")

MODEL_UNCERTAINTY reports availability, supported probabilities/directions, method, split and seed, train/calibration/test counts, dataset identity, held-out observed coverage, and limitations. It does not run inference. Observed coverage is evidence about that held-out data, not proof that future data will retain the same distribution.

Choose the calibration policy

Training requests expose options beyond the checkbox:

Option Behavior
Omitted or mode: "off" Ordinary point training, without calibration splits or residual storage.
mode: "auto" Attempt calibration; unsupported families/tasks or insufficient rows can yield point-only training with an unavailability reason. Malformed requests and invalid data still fail.
mode: "required" Fail training if the requested calibrated capability cannot be produced.
calibrationSplit: "random" Deterministic seeded split of raw rows before fitting preprocessing or the model.
calibrationSplit: "temporal" Preserve input row order for training, calibration, then test holdout. Supply chronologically ordered data yourself.

Selected probabilities must be unique finite numbers strictly between zero and one, at most 16. An omitted or empty list selects [0.9]. The producer currently supplies upper endpoints only; it does not interpolate other levels or derive a lower bound from an upper one.

The current P90 split needs at least 110 finite labeled rows: at least 99 for calibration, 10 for training, and one for testing. This is a minimum for admission, not a recommendation for adequate predictive quality. Larger data uses larger holdouts; probabilities closer to one may require more rows to support a finite endpoint. Preprocessing is fitted on training rows only.

Register an externally trained artifact

For native XGBoost, use Upload & register artifact with both the booster and its exported metadata sidecar. Calibrated artifacts use .ubj or .xgb.json with the matching metadata format. Legacy .bst remains point-only. Save the workbook before uploading.

Grid binds the uploaded bytes to content-addressed Files paths and verifies the booster and sidecar together. A user-entered claim that a model is calibrated is not sufficient. Missing, modified, or mismatched metadata must be corrected by registering the exact valid artifact/sidecar pair; do not work around rejection by copying capability fields into a request.

When no bound is available

Symptom Next step
Point prediction works but P90 is absent Inspect the exact version's MODEL_UNCERTAINTY; check task, calibration policy, row counts, and the unavailability reason.
MODEL_PREDICTIVE_CALIBRATION_INSUFFICIENT_ROWS Supply more finite labeled rows or intentionally choose point-only training.
MODEL_PREDICTIVE_UNCERTAINTY_UNAVAILABLE Select a calibrated version or retrain/register one. A legacy model does not gain calibration from a newer Grid installation.
MODEL_PREDICTION_BOUND_UNSUPPORTED_LEVEL or MODEL_PREDICTION_BOUND_UNSUPPORTED_DIRECTION Use an exact admitted probability and direction, or train a new version that supports the requested upper endpoint.
Metadata/hash admission failure Restore or register the correctly bound original bytes; do not weaken verification.

Do not silently replace an unavailable bound with the point prediction. That would change the meaning of a downstream risk limit. Workbook CONF_* and PROB_* operations propagate authored input distributions and do not synthesize calibration for external models. See external-function semantics.