keel
Context

Budgets and fit

Understand section limits, visible clipping, and configuration and per-request fit checks.

Context budgets make size decisions explicit. Read grants control rendered state; framework section defaults identify other overruns; the fit gate checks the opening assembly and each subsequent request against the declared input allocation and the chain's window; and the turn's output-token grant caps what the model may spend answering.

Declared window: 8,000 tokens
Context 6,000
+ 2,000
8,000 total → fits
Context 6,500
+ 2,000
8,500 total → fit-overflow · no model call
Illustrative estimates, not provider tokenization. The request estimate is checked before every model step.

Declare the context allocation

Supply context to defineChain, or to defineAgent for an agent override:

context: {
  maxInputTokens: 30_000,
  outputReserveTokens: 2_000,
},

The combined allocation must fit every tier, and the output reserve must fit every tier's output ceiling. A mismatch throws KeelError when the definition executes. This catches a smaller fallback before the application serves turns.

maxInputTokens covers the complete request, including tool schemas and the growing transcript. It does not replace individual section budgets.

Know the section defaults

SectionCurrent budget in estimated tokens
sys.time and sys.rules64
Each app ambient sectionIts declared budgetTokens, default 64
prompt4,096
Each skill1,024
Each tool description256
Each state renderIts read grant's budgetTokens
history8,192
Current user brief8,192

Read grants are configurable on the agent, and ambient section budgets on the app's defineAmbient declaration. The other values above are current assembler defaults, not fields exposed in an agent's configuration.

Distinguish clipping from warnings

If a state reader produces too much text, the assembler clips it and includes a notice naming the read budget. The render event has clipped: true, and a warning is emitted. The reader receives its budget as an argument, so prefer a useful summary before this fallback is needed.

Other section overruns are recorded in Assembly.overBudget and emitted as warnings. Their text is not silently shortened. A section warning alone does not fail the turn if the total opening context still fits.

Clipping uses character counts and a notice string. Extremely small grants can be smaller than the notice itself; do not use clipping as a substitute for a usable read budget. The total fit check still evaluates the result.

Apply the opening fit check

The opening gate applies two conditions:

estimated assembly tokens + output reserve <= chain window
estimated assembly tokens <= context.maxInputTokens

For a chain window of 8,000 tokens, an output reserve of 2,000, and a maxInputTokens of 6,000:

Opening assembly estimateTotal with reserveResult
5,8007,800Fits both conditions.
6,2008,200Refused with fit-overflow — both conditions fail.

Either condition failing refuses the turn. On refusal, no model tier is called. The defect includes totalTokens, outputReserve, maxInputTokens, and windowTokens, so the application can report a concrete cause. This check happens before model execution, after assembly and its reader functions have already run.

Check the growing request

Before each model step, Keel estimates the normalized request: system text, user brief, tool schemas, transcript, and any forced-tool instruction. That includes delivery instructions and tool results added after opening assembly, as well as continuation and recovery nudges. The per-step gate has a third condition:

estimated request input <= context.maxInputTokens
estimated request input + output reserve <= chain window
output reserve <= every tier's output ceiling

A request whose output reserve exceeds any tier's maxOutputTokens is also a fit-overflow — the reserve could not be honored on that tier.

A refusal returns fit-overflow before any adapter runs. It is a capacity failure, so the chain does not retry or switch providers — and a mid-turn fit-overflow hard-fails the whole turn with that cause. This holds even inside continuation and deliverable-nudge rounds: a pending deliverable recovery does not soften the failure or rebrand it.

Without an explicit context allocation, the input ceiling is the smallest chain window minus the smallest tier output ceiling.

Grant the turn's output tokens

Fit checks reason about capacity; TurnBudget.outputTokens is a spend grant, and it is enforced:

const outcome = await runtime.runTurn({
  agent: assistantAgent,
  ambient,
  budget: { steps: 6, outputTokens: 4_000 },
  envelope,
});

The grant is cumulative across the turn's model requests. Before each request, the output reserve is clamped to what remains of the grant, so a turn near its limit makes smaller requests rather than overrunning. When the grant is exhausted before the model finishes — including before a continuation or nudge round — the turn settles as a failed budget-cut, never a synthetic completion. Leaving outputTokens unset leaves the reserve governed by the chain's declared allocation alone.

Understand the estimate's boundary

Keel estimates section sizes as Math.ceil(text.length / 4) and request sizes from the JSON-serialized normalized request using that same ratio. Request checks include schema and transcript content, but do not implement each provider's tokenizer or exact message framing.

A passing estimate therefore cannot guarantee provider acceptance. Leave headroom for your content and model, especially for languages or data where four characters per token is optimistic. A provider rejection still follows the chain's classified failure policy.

Fix the source of an overrun

Large inputUseful adjustment
State renderSelect fewer fields or summarize inside the reader.
Stored historyReduce historyWindow and inspect unusually long messages.
Prompt or skillsRemove repeated instructions and unnecessary examples.
Tool resultsReturn the fields needed for the next decision.
Output reserveReview configured tier limits alongside the output your task needs.

Increasing a declared window only helps when the actual chosen model supports it. Do not change that number merely to silence a refusal.

Inspect in Studio

In Studio, inspect CONTEXT for each section's used and budget figures. A clipped render should show its notice and warning. For a controlled overflow, expand a prompt until the opening total exceeds the configured window, then confirm a failed fit gate and no model step.

Shorten the offending source and repeat the same request. Keep an overflow test in your suite when a prompt or reader change could reintroduce the overrun.

Next: Agents.