History and transcripts
Distinguish stored conversation history from the current turn’s tool loop.
A model can need both the previous conversation and the work already done in the current turn. Keel supplies these through two different paths.
Opening history
store.listMessages(thread, 6)
→ "user: …\nassistant: …"Current turn
model → tool call
tool → result
model → next stepSelect recent stored messages
Set historyWindow on the agent:
historyWindow: 6,This selects the last six stored messages, not six user-assistant pairs
and not a token count. The default is six; zero excludes stored history.
That zero rule is part of the Store contract, not a convenience: an
adapter's listMessages must return none for window <= 0, never
everything. A custom store that slices with slice(-0) would leak the whole
thread into an agent that declined history.
The assembler joins selected messages into text:
user: Please summarize my notes.
assistant: Your notes cover the Friday draft.This history comes from the thread's store. The current implementation uses message roles and text; it does not automatically replay structured message parts as full provider-native history.
Know when history is written
History grows on exactly two events, both on root turns:
- A user message is appended only when a root turn opens with a request envelope. It is stored before assembly.
- An assistant message is appended only when a root turn settles
completed. Failed, cut, stopped, and truncated turns write no assistant message.
Brief-driven root turns (no envelope) and delegated child turns never append to history. A researcher a turn delegates to leaves no trace in the thread's conversation — its output returns to its caller, not to the store.
Distinguish history from the current brief
The current request has its own user assembly section and is passed to the
adapter as brief. Because an envelope's user message is stored before
assembly, the selected stored history can also include that message; it is
not guaranteed to contain only messages preceding the current request.
Inspect the actual assembled history when diagnosing repetition. A history window is a storage selection rule, not automatic deduplication or summarization.
Follow the current-turn transcript
After a model step, the runtime records its text and calls. Tool results are then added to the transcript supplied to the next step:
assistant → calls set-tone({ tone: "detailed" })
tools → returns { tone: "detailed" }
assistant → continues with the result availableA normalized transcript item can be an assistant step, a group of tool results, or a recovery nudge. Provider adapters replay this transcript on each request, including when the chain falls back to another tier.
This is how the model learns a tool's result within a turn. The runtime does not rebuild every opening state view after each tool call.
Control growth deliberately
Lowering historyWindow reduces selected stored messages, but one large
message can still be expensive. It also does not limit the current-turn
transcript, which grows as tools and model steps run.
Keep tool results focused, limit the work granted to a turn, and inspect
large messages. Before each model step, Keel estimates the complete request
including the accumulated transcript. An overrun fails with fit-overflow
before the provider call; this is an estimate, not exact provider tokenization.
Inspect in Studio
Run a multi-step interaction in Studio. Compare CONTEXT's
opening history with MODEL's tool calls and results. Then start a new
turn with a smaller historyWindow and compare the opening history again.
Use a fresh thread or a controlled seed so you know which messages were
stored before each run.
Next: Budgets and fit.
