keel
Model

Retry and fallback

Follow classified model failures through bounded retries and fallback.

The chain makes retry decisions from failure classes. An adapter reports what happened; the runtime decides whether to try it once more or move to the next tier.

Tier 1: transient
  ↳ retry Tier 1 once
  ↳ fails again → Tier 2
Tier 1: unavailable
  ↳ skip retry
  ↳ try Tier 2
All tiers fail → chain-exhausted → failed turn
Retry policy is based on failure class. Partial-content failure follows a separate cut path.

Classify the failure

FailureChain behavior
transient, or an unclassified thrown errorRetry the same tier once, then advance if it fails again.
overloadedAdvance immediately.
unavailableAdvance immediately.
auth or bad-requestAdvance immediately.
refusedAdvance immediately.
A result marked failure: "empty" or "error-pre"Retry once, then advance if repeated.

Provider adapters map errors into these classes. Inspect the reported class and message rather than assuming every network or HTTP failure follows the same policy. The policy is fixed in runChainStep; it is separate from a tool's configurable retries.

The retry counter is one per tier per step, shared across the retryable classes: a transient throw and an empty or error-pre result draw from the same single retry. An empty result followed by a transient error advances to the next tier; it does not earn a second same-tier attempt.

When the turn carries tools, a tier that does not declare supportsTools is skipped without being called. The skip is loud: a fallback chain event with detail.reason: "no-tool-support". See Tool capabilities.

The turn record's fallbacksTaken lists the tiers that failed or were skipped — the tiers the chain consumed, not the tier it landed on.

Preserve the request on fallback

The fallback receives the same opening context, available tool specifications, and current-turn transcript. It can see the tool results already obtained. The chain advances through its configured tiers and keeps that selection for later steps in the turn.

Different tiers can still interpret the request differently. Test fallback behavior against your application's requirements, not just whether it returns text.

Do not replay partial output blindly

Streaming sharpens this rule. A streaming tier's failure point decides the policy: a throw before any content is classified and follows the table above — nothing reached the wire, so a retry or fallback is safe. A throw after content becomes a result with finish: "error-post": text deltas have already landed on the wire, and replaying the request would duplicate them. The chain passes that result to the turn instead of retrying it as a clean request, and the runtime cuts the turn with the typed stream-error reason.

A length result is also distinct from a provider exception. It follows the turn's bounded continuation behavior; it is not an automatic reason to switch tiers. See Turn.

Handle exhaustion as a failure

If no tier can serve the step, the chain returns ok: false with a chain-exhausted defect. The runtime settles the turn as failed with that cause — at every call site, including continuation rounds after a truncation and deliverable nudge rounds. The real cause is never rebranded as a truncation or a missing deliverable, and a generated apology is not substituted for the result.

The turn records refunded: true on this path. That flag is application-facing runtime bookkeeping; it does not perform a payment-provider refund.

Inspect in Studio

Use deterministic adapters to produce a transient failure followed by a successful retry, immediate fallback, and exhaustion. In Studio, inspect MODEL for retry, fallback, step, and exhausted chain events, then inspect TURN for the final outcome.

The custom adapter guide shows how to inject these failures without relying on an actual provider outage.

Next: Tool capabilities.