keel
Harness

Overview

Control what a turn may execute, how it recovers, and when it can finish.

01Assemble
prompt + state views
+ tools + history
02Call model
estimate + reserve
→ chain.generate
03Validate
known tool? allowed?
input schema passes?
04Execute
tool result
or explicit failure
05Continue
result → transcript
↳ next model step
06Settle
close gate
→ typed outcome
Correction returns through the same checks · recovery stays bounded
The model proposes the next action. The harness decides whether that action may run and whether the turn may finish.

The harness is the runtime around the model. It assembles input, checks proposed operations, executes tools, returns feedback, and decides how the turn ends. The model chooses what to attempt; application declarations and runtime checks determine what is permitted.

Why a harness?

A model can request a nonexistent tool, supply invalid arguments, or stop without returning the result its caller needs. A prompt can describe the right behavior, but execution needs checks of its own.

Keel makes those checks part of the turn. Invalid arguments are refused before tool execution. Corrected calls pass through the same validation. A missing required deliverable prevents completion even if the model says it finished.

Declare the work and its limits

const outcome = await runtime.runTurn({
  agent: assistant,
  ambient,
  envelope,
  budget: { steps: 8, timeMs: 30_000 },
});

if (outcome.kind === "completed") {
  present(outcome.payload);
} else {
  handleOutcome(outcome);
}

Here, runtime, assistant, and the request context come from your app; present and handleOutcome are application handlers. Branch on the outcome, not on whether any text arrived.

Make failure observable

Suppose an agent delegates research and the child never submits its required findings. The child fails with deliverable-missing. The caller receives that failure and either continues with feedback or fails according to its onError policy. A child finishing does not complete its caller.

The harness checks declared contracts. It does not prove factual accuracy or that a schema-valid result is useful. Cover those requirements with your app's own behavior tests — see Testing.

Work with the harness

GuideWhat you'll learn
Gates and invariantsSeparate configuration errors from checks during a turn.
Budgets and cancellationBound work, recovery, and nested execution.
Failure and recoveryTrace rejected calls, retries, correction feedback, and failures.
Completion and deliverablesDeclare what must exist before a turn can complete.
Tracing and inspectionExplain a run from its events and final record.

Model owns provider selection and fallback. Delegation owns the caller/child relationship. Harness explains how those mechanisms affect execution and the turn's outcome.

Inspect in Studio

Run a normal request and a deliberate failure. In TURN RECORDS, compare the outcomes; in MODEL, follow the operations that led to them. Use WIRE to distinguish visible output from proof of completion.

Next: Gates and invariants.