Limits and usage
Understand chain limits, normalized finishes, and token and cost records.
A chain must accommodate the same request on every fallback. Keel checks declared allocations when you define the chain or agent, then estimates the actual request before every model step.
context: {
maxInputTokens: 30_000,
outputReserveTokens: 4_000
}KeelError: context allocation
does not fit tier "fallback"Separate model capacity from request budgets
| Value | Meaning |
|---|---|
provider.capabilities | Known or explicitly supplied model capabilities. |
provider.windowTokens | Selected combined input and output ceiling. |
provider.maxOutputTokens | Selected maximum output per request. |
context.maxInputTokens | Application allocation for the complete input. |
context.outputReserveTokens | Output reserved and passed to the provider. |
Known models resolve capabilities automatically. Application ceilings may be smaller, but cannot exceed the resolved capacity. See Connecting providers for custom models.
Reject incompatible chains at startup
const chain = defineChain({
id: "assistant-chain",
tiers: [primary, fallback],
context: {
maxInputTokens: 30_000,
outputReserveTokens: 4_000,
},
});If the fallback has a 32,000-token window, this declaration throws a
KeelError: the required 34,000 tokens exceed that tier's capacity. The
message names the chain and limiting tier, with the required and available
limits. Reduce the allocation or choose a larger fallback.
The output reserve must also fit each tier's selected maxOutputTokens.
An agent may declare its own context allocation; defineAgent validates
that override against every tier. Without an override, it inherits the
chain's allocation.
This is configuration-time validation when the definition executes, before serving turns — and it happens exactly once. The runtime does not re-validate the declared allocation on every step; only the per-request fit estimate below runs each step. TypeScript checks model IDs and configuration shapes; numeric capacity comparisons run in JavaScript, including values loaded at startup.
Check every request
Without an explicit allocation, the chain uses the smallest tier window and
smallest output ceiling. With an allocation, the input must also stay within
maxInputTokens. Keel checks the opening assembly and then checks each
normalized request, including tool schemas, delivery instructions, accumulated
results, and recovery nudges.
An overflow returns fit-overflow before calling a provider. It does not
retry or advance to a fallback. The shipped adapters receive the output
reserve as the request's output limit, so input checks and requested output
use the same allocation.
These are character-based estimates, not exact provider token counts. Provider framing and tokenization can differ. See Budgets and fit for the estimate's limits.
Grant a turn-level output budget
TurnBudget.outputTokens is an enforced, cumulative output-token grant for
the whole turn, spanning every model request it makes:
- Each request's output reserve is clamped to whatever remains of the grant, so a request can never ask the provider for more output than the turn has left.
- When the grant is exhausted before the model finishes — including before a
continuation or deliverable-recovery round — the turn settles as
failedwith abudget-cutdefect. There is no silent overrun and no synthetic completion.
The grant counts the turn's accumulated output usage, provider-reported where available and estimated otherwise (see below).
Read the normalized finish
| Finish | Meaning for the runtime |
|---|---|
stop | The model stopped; the turn still checks whether its output satisfies completion requirements. |
tool-calls | Process the requested tools and continue the model loop. |
length | Apply the turn's continuation policy or settle truncation. |
error-post | Cut the turn because the stream failed after content. |
An adapter step is not the same thing as a completed turn. A returned stop
can still fail the close gate if the required output is missing.
Inspect usage accounting
Adapters can return usage with input and output token counts. The runtime
accumulates those figures for returned steps. When a step reports no usage,
the runtime estimates it and marks the record's usageEstimated: the input
estimate uses the same full-request accounting as the fit gate
(estimateRequestTokens over the system sections, brief, tool schemas, and
transcript), and the output estimate uses the step's emitted text length.
These records are useful for understanding a run, but they are not an invoice reconciliation system. Failed attempts that return no usage cannot contribute provider-reported token charges to the record.
Supply account pricing
Adapters have no built-in price table. Supply your account's rates explicitly:
// Options passed to your provider factory; values come from app configuration.
pricing: {
inputPerMTok: accountRates.inputPerMillion,
outputPerMTok: accountRates.outputPerMillion,
},The runtime computes each priced step's cost from its usage. If no used tier
has pricing, costUsd stays null. If only some used tiers have pricing, the
record only totals those priced steps; configure all tiers for meaningful
comparisons. The two-rate calculation does not model every possible provider
billing category.
Inspect in Studio
Compare MODEL usage with TURN records after a normal run and a fallback run. Check whether usage is estimated and whether every used tier has configured pricing before interpreting a cost total.
Next: Custom adapters.
