keel
Model

Limits and usage

Understand chain limits, normalized finishes, and token and cost records.

A chain must accommodate the same request on every fallback. Keel checks declared allocations when you define the chain or agent, then estimates the actual request before every model step.

One allocation · every tier
context: {
  maxInputTokens: 30_000,
  outputReserveTokens: 4_000
}
Primary
34,000 required / 128,000 available
Fits the window
Fallback
34,000 required / 32,000 available
Exceeds window by 2,000
KeelError: context allocation
does not fit tier "fallback"
A fallback must support the same allocation as the primary. This mismatch is caught when the chain is defined.

Separate model capacity from request budgets

ValueMeaning
provider.capabilitiesKnown or explicitly supplied model capabilities.
provider.windowTokensSelected combined input and output ceiling.
provider.maxOutputTokensSelected maximum output per request.
context.maxInputTokensApplication allocation for the complete input.
context.outputReserveTokensOutput reserved and passed to the provider.

Known models resolve capabilities automatically. Application ceilings may be smaller, but cannot exceed the resolved capacity. See Connecting providers for custom models.

Reject incompatible chains at startup

const chain = defineChain({
  id: "assistant-chain",
  tiers: [primary, fallback],
  context: {
    maxInputTokens: 30_000,
    outputReserveTokens: 4_000,
  },
});

If the fallback has a 32,000-token window, this declaration throws a KeelError: the required 34,000 tokens exceed that tier's capacity. The message names the chain and limiting tier, with the required and available limits. Reduce the allocation or choose a larger fallback.

The output reserve must also fit each tier's selected maxOutputTokens. An agent may declare its own context allocation; defineAgent validates that override against every tier. Without an override, it inherits the chain's allocation.

This is configuration-time validation when the definition executes, before serving turns — and it happens exactly once. The runtime does not re-validate the declared allocation on every step; only the per-request fit estimate below runs each step. TypeScript checks model IDs and configuration shapes; numeric capacity comparisons run in JavaScript, including values loaded at startup.

Check every request

Without an explicit allocation, the chain uses the smallest tier window and smallest output ceiling. With an allocation, the input must also stay within maxInputTokens. Keel checks the opening assembly and then checks each normalized request, including tool schemas, delivery instructions, accumulated results, and recovery nudges.

An overflow returns fit-overflow before calling a provider. It does not retry or advance to a fallback. The shipped adapters receive the output reserve as the request's output limit, so input checks and requested output use the same allocation.

These are character-based estimates, not exact provider token counts. Provider framing and tokenization can differ. See Budgets and fit for the estimate's limits.

Grant a turn-level output budget

TurnBudget.outputTokens is an enforced, cumulative output-token grant for the whole turn, spanning every model request it makes:

  • Each request's output reserve is clamped to whatever remains of the grant, so a request can never ask the provider for more output than the turn has left.
  • When the grant is exhausted before the model finishes — including before a continuation or deliverable-recovery round — the turn settles as failed with a budget-cut defect. There is no silent overrun and no synthetic completion.

The grant counts the turn's accumulated output usage, provider-reported where available and estimated otherwise (see below).

Read the normalized finish

FinishMeaning for the runtime
stopThe model stopped; the turn still checks whether its output satisfies completion requirements.
tool-callsProcess the requested tools and continue the model loop.
lengthApply the turn's continuation policy or settle truncation.
error-postCut the turn because the stream failed after content.

An adapter step is not the same thing as a completed turn. A returned stop can still fail the close gate if the required output is missing.

Inspect usage accounting

Adapters can return usage with input and output token counts. The runtime accumulates those figures for returned steps. When a step reports no usage, the runtime estimates it and marks the record's usageEstimated: the input estimate uses the same full-request accounting as the fit gate (estimateRequestTokens over the system sections, brief, tool schemas, and transcript), and the output estimate uses the step's emitted text length.

These records are useful for understanding a run, but they are not an invoice reconciliation system. Failed attempts that return no usage cannot contribute provider-reported token charges to the record.

Supply account pricing

Adapters have no built-in price table. Supply your account's rates explicitly:

// Options passed to your provider factory; values come from app configuration.
pricing: {
  inputPerMTok: accountRates.inputPerMillion,
  outputPerMTok: accountRates.outputPerMillion,
},

The runtime computes each priced step's cost from its usage. If no used tier has pricing, costUsd stays null. If only some used tiers have pricing, the record only totals those priced steps; configure all tiers for meaningful comparisons. The two-rate calculation does not model every possible provider billing category.

Inspect in Studio

Compare MODEL usage with TURN records after a normal run and a fallback run. Check whether usage is estimated and whether every used tier has configured pricing before interpreting a cost total.

Next: Custom adapters.