keel

Testing

Fake a tier behind the public ModelAdapter interface, keep the runtime real — and let the framework's own behavior suites carry the rest.

A keel app is deterministic everywhere except the model. The framework's testing story follows from that split: keel verifies its own contracts with per-module behavior suites, and your app verifies its behavior by scripting the one nondeterministic piece — the tier — while every gate, budget, and state rule runs for real.

Fake a tier, keep the runtime real

Testing a keel app needs no testing package. A tier is anything that implements the public ModelAdapter interface, so a deterministic fake is a few lines of your own code — and the whole arrangement fits in one test:

import assert from "node:assert/strict";
import { createRuntime, defineAgent, defineChain, PartRegistry, type ModelAdapter } from "@keel-dev/core";

const fakeTier: ModelAdapter = {
  id: "fake",
  windowTokens: 32_000,
  maxOutputTokens: 4_000,
  generate: () => ({ events: [{ type: "text", text: "Found it." }], finish: "stop" }),
};

const agent = defineAgent({
  id: "assistant",
  chain: defineChain({ id: "assistant-chain", tiers: [fakeTier] }),
  prompt: "Answer honestly.",
  reads: [], tools: [], emits: [],
});

const runtime = createRuntime({
  registry: new PartRegistry(),
  budgets: { "turn.wallClock": { limitMs: 30_000 }, "transport.envelopeBytes": { limit: 16_384 } },
});

const outcome = await runtime.runTurn({
  agent,
  ambient: { userId: "u-1", threadId: "t-1", scope: {} },
  envelope: { threadId: "t-1", clientTurnId: "test-1", message: "hello", attachments: [] },
  budget: { steps: 4 },
});

assert.equal(outcome.kind, "completed");
assert.ok(runtime.events.some((e) => e.kind === "gate" && e.gate === "close" && e.verdict === "pass"));

Because the fake sits behind the same interface as a live provider, the test exercises the real runtime — real gates, real budgets, real chain policy — with only the text generation scripted. A fake that returns a call event drives the real tool executor; a fake that throws a TierError proves your chain's retry and fallback policy; a fake with stream exercises the live drain. For time-budget tests, inject your own clock: Clock is a one-method interface ({ now(): number }), so a manually advanced implementation passed as RuntimeConfig.clock makes a wall-clock test run in microseconds.

Assert behavior, not just settlement: check the outcome's kind, then the evidence — outcome.payload, the records from await runtime.store.records(), and the bus events in runtime.events. For a search that finds nothing, check that the answer says so; for a state write, check the stored block. Completion alone proves neither.

Fakes stay out of production

A fake tier belongs in tests — never in a deployed chain, where scripted content answering real users would be a synthetic success. The Model chapter carries the same warning from the production side.

Run the suites

From the repository root:

pnpm --filter @keel-dev/core test
pnpm --filter example-assistant test

Bun runs these tests. The first checks the framework's own contracts; the second checks the reference application's behavior. A scaffolded app runs its own tests with npm test.

How the framework tests itself

The framework's tests live in packages/keel/test/, one folder per module — state, wire, context, model, agent, turn, delegation — and each suite pins the behavior its module promises, using the same technique this chapter recommends: scripted ModelAdapter fakes driving the real runtime.

SuiteWhat it pins
state/The Store contract — block roundtrips, per-thread keying, listMessages window semantics — and block open/reduce/repair behavior.
wire/Part registration, fixtures, marker promotion, and the reserved outcome part.
context/Assembly order, section budgets, and render clipping.
model/Chain retry and fallback classes, capacity validation, and provider request shapes.
agent/Build-time declaration checks — undeclared reads, tool/tier compatibility, byproduct bindings.
turn/The four gates, budgets, deliverables, streaming, tools, and exhaustion — each refusal driven end to end through runTurn.
delegation/Caller/child ownership: grants, forwarding, and failure feedback.

test/turn/gates.test.ts is a good first read: an oversized context settles failed with fit-overflow and zero steps used; an open marker settles marker-open; narration-only output settles empty. Every check is a real turn with a scripted tier — no live provider, no flakiness.

Add a regression test

When a bug appears, reproduce it with a fake tier or an injected failure, assert the observable result that matters to the caller, and keep the test next to your app (test/, as in examples/assistant/test/ and the scaffolded starter's test/app.test.ts). The assistant's workspace-search.test.ts shows the shape: run a turn, assert the root outcome, find the child's record, and check the wire events for what did — and did not — reach the client.

Mocks test the runtime; evals test the prompts

A scripted model verifies how the runtime responds to known model actions. Whether your prompts reliably produce those actions is a separate question — cover it with live-model evaluation.

That completes the tour. Head back to the Introduction to retrace any concept, or explore the source and examples on GitHub.