Retry and fallback
Follow classified model failures through bounded retries and fallback.
The chain makes retry decisions from failure classes. An adapter reports what happened; the runtime decides whether to try it once more or move to the next tier.
Tier 1: transient
↳ retry Tier 1 once
↳ fails again → Tier 2Tier 1: unavailable
↳ skip retry
↳ try Tier 2Classify the failure
| Failure | Chain behavior |
|---|---|
transient, or an unclassified thrown error | Retry the same tier once, then advance if it fails again. |
overloaded | Advance immediately. |
unavailable | Advance immediately. |
auth or bad-request | Advance immediately. |
refused | Advance immediately. |
A result marked failure: "empty" or "error-pre" | Retry once, then advance if repeated. |
Provider adapters map errors into these classes. Inspect the reported class
and message rather than assuming every network or HTTP failure follows the
same policy. The policy is fixed in runChainStep; it is separate from a
tool's configurable retries.
The retry counter is one per tier per step, shared across the retryable
classes: a transient throw and an empty or error-pre result draw from the
same single retry. An empty result followed by a transient error advances to
the next tier; it does not earn a second same-tier attempt.
When the turn carries tools, a tier that does not declare supportsTools is
skipped without being called. The skip is loud: a fallback chain event with
detail.reason: "no-tool-support". See
Tool capabilities.
The turn record's fallbacksTaken lists the tiers that failed or were
skipped — the tiers the chain consumed, not the tier it landed on.
Preserve the request on fallback
The fallback receives the same opening context, available tool specifications, and current-turn transcript. It can see the tool results already obtained. The chain advances through its configured tiers and keeps that selection for later steps in the turn.
Different tiers can still interpret the request differently. Test fallback behavior against your application's requirements, not just whether it returns text.
Do not replay partial output blindly
Streaming sharpens this rule. A streaming tier's failure point decides the
policy: a throw before any content is classified and follows the table
above — nothing reached the wire, so a retry or fallback is safe. A throw
after content becomes a result with finish: "error-post": text deltas
have already landed on the wire, and replaying the request would duplicate
them. The chain passes that result to the turn instead of retrying it as a
clean request, and the runtime cuts the turn with the typed stream-error
reason.
A length result is also distinct from a provider exception. It follows the
turn's bounded continuation behavior; it is not an automatic reason to switch
tiers. See Turn.
Handle exhaustion as a failure
If no tier can serve the step, the chain returns ok: false with a
chain-exhausted defect. The runtime settles the turn as failed with that
cause — at every call site, including continuation rounds after a truncation
and deliverable nudge rounds. The real cause is never rebranded as a
truncation or a missing deliverable, and a generated apology is not
substituted for the result.
The turn records refunded: true on this path. That flag is application-facing
runtime bookkeeping; it does not perform a payment-provider refund.
Inspect in Studio
Use deterministic adapters to produce a transient failure followed by a
successful retry, immediate fallback, and exhaustion. In Studio,
inspect MODEL for retry, fallback, step, and exhausted chain events,
then inspect TURN for the final outcome.
The custom adapter guide shows how to inject these failures without relying on an actual provider outage.
Next: Tool capabilities.
