Budgets and fit
Understand section limits, visible clipping, and configuration and per-request fit checks.
Context budgets make size decisions explicit. Read grants control rendered state; framework section defaults identify other overruns; the fit gate checks the opening assembly and each subsequent request against the declared input allocation and the chain's window; and the turn's output-token grant caps what the model may spend answering.
Declare the context allocation
Supply context to defineChain, or to defineAgent for an agent override:
context: {
maxInputTokens: 30_000,
outputReserveTokens: 2_000,
},The combined allocation must fit every tier, and the output reserve must fit
every tier's output ceiling. A mismatch throws KeelError when the definition
executes. This catches a smaller fallback before the application serves turns.
maxInputTokens covers the complete request, including tool schemas and the
growing transcript. It does not replace individual section budgets.
Know the section defaults
| Section | Current budget in estimated tokens |
|---|---|
sys.time and sys.rules | 64 |
| Each app ambient section | Its declared budgetTokens, default 64 |
prompt | 4,096 |
| Each skill | 1,024 |
| Each tool description | 256 |
| Each state render | Its read grant's budgetTokens |
history | 8,192 |
Current user brief | 8,192 |
Read grants are configurable on the agent, and ambient section budgets on
the app's defineAmbient declaration. The other values above are current
assembler defaults, not fields exposed in an agent's configuration.
Distinguish clipping from warnings
If a state reader produces too much text, the assembler clips it and includes
a notice naming the read budget. The render event has clipped: true, and a
warning is emitted. The reader receives its budget as an argument, so prefer
a useful summary before this fallback is needed.
Other section overruns are recorded in Assembly.overBudget and emitted as
warnings. Their text is not silently shortened. A section warning alone does
not fail the turn if the total opening context still fits.
Clipping uses character counts and a notice string. Extremely small grants can be smaller than the notice itself; do not use clipping as a substitute for a usable read budget. The total fit check still evaluates the result.
Apply the opening fit check
The opening gate applies two conditions:
estimated assembly tokens + output reserve <= chain window
estimated assembly tokens <= context.maxInputTokensFor a chain window of 8,000 tokens, an output reserve of 2,000, and a
maxInputTokens of 6,000:
| Opening assembly estimate | Total with reserve | Result |
|---|---|---|
| 5,800 | 7,800 | Fits both conditions. |
| 6,200 | 8,200 | Refused with fit-overflow — both conditions fail. |
Either condition failing refuses the turn. On refusal, no model tier is
called. The defect includes totalTokens, outputReserve,
maxInputTokens, and windowTokens, so the application can report a
concrete cause. This check happens before model execution, after assembly
and its reader functions have already run.
Check the growing request
Before each model step, Keel estimates the normalized request: system text, user brief, tool schemas, transcript, and any forced-tool instruction. That includes delivery instructions and tool results added after opening assembly, as well as continuation and recovery nudges. The per-step gate has a third condition:
estimated request input <= context.maxInputTokens
estimated request input + output reserve <= chain window
output reserve <= every tier's output ceilingA request whose output reserve exceeds any tier's maxOutputTokens is also
a fit-overflow — the reserve could not be honored on that tier.
A refusal returns fit-overflow before any adapter runs. It is a capacity
failure, so the chain does not retry or switch providers — and a mid-turn
fit-overflow hard-fails the whole turn with that cause. This holds even
inside continuation and deliverable-nudge rounds: a pending deliverable
recovery does not soften the failure or rebrand it.
Without an explicit context allocation, the input ceiling is the smallest chain window minus the smallest tier output ceiling.
Grant the turn's output tokens
Fit checks reason about capacity; TurnBudget.outputTokens is a spend
grant, and it is enforced:
const outcome = await runtime.runTurn({
agent: assistantAgent,
ambient,
budget: { steps: 6, outputTokens: 4_000 },
envelope,
});The grant is cumulative across the turn's model requests. Before each
request, the output reserve is clamped to what remains of the grant, so a
turn near its limit makes smaller requests rather than overrunning. When
the grant is exhausted before the model finishes — including before a
continuation or nudge round — the turn settles as a failed budget-cut,
never a synthetic completion. Leaving outputTokens unset leaves the
reserve governed by the chain's declared allocation alone.
Understand the estimate's boundary
Keel estimates section sizes as Math.ceil(text.length / 4) and request sizes
from the JSON-serialized normalized request using that same ratio. Request
checks include schema and transcript content, but do not implement each
provider's tokenizer or exact message framing.
A passing estimate therefore cannot guarantee provider acceptance. Leave headroom for your content and model, especially for languages or data where four characters per token is optimistic. A provider rejection still follows the chain's classified failure policy.
Fix the source of an overrun
| Large input | Useful adjustment |
|---|---|
| State render | Select fewer fields or summarize inside the reader. |
| Stored history | Reduce historyWindow and inspect unusually long messages. |
| Prompt or skills | Remove repeated instructions and unnecessary examples. |
| Tool results | Return the fields needed for the next decision. |
| Output reserve | Review configured tier limits alongside the output your task needs. |
Increasing a declared window only helps when the actual chosen model supports it. Do not change that number merely to silence a refusal.
Inspect in Studio
In Studio, inspect CONTEXT for each section's used and budget figures. A clipped render should show its notice and warning. For a controlled overflow, expand a prompt until the opening total exceeds the configured window, then confirm a failed fit gate and no model step.
Shorten the offending source and repeat the same request. Keep an overflow test in your suite when a prompt or reader change could reintroduce the overrun.
Next: Agents.
