Testing
Fake a tier behind the public ModelAdapter interface, keep the runtime real — and let the framework's own behavior suites carry the rest.
A keel app is deterministic everywhere except the model. The framework's testing story follows from that split: keel verifies its own contracts with per-module behavior suites, and your app verifies its behavior by scripting the one nondeterministic piece — the tier — while every gate, budget, and state rule runs for real.
Fake a tier, keep the runtime real
Testing a keel app needs no testing package. A tier is anything that
implements the public ModelAdapter interface, so a deterministic fake is a
few lines of your own code — and the whole arrangement fits in one test:
import assert from "node:assert/strict";
import { createRuntime, defineAgent, defineChain, PartRegistry, type ModelAdapter } from "@keel-dev/core";
const fakeTier: ModelAdapter = {
id: "fake",
windowTokens: 32_000,
maxOutputTokens: 4_000,
generate: () => ({ events: [{ type: "text", text: "Found it." }], finish: "stop" }),
};
const agent = defineAgent({
id: "assistant",
chain: defineChain({ id: "assistant-chain", tiers: [fakeTier] }),
prompt: "Answer honestly.",
reads: [], tools: [], emits: [],
});
const runtime = createRuntime({
registry: new PartRegistry(),
budgets: { "turn.wallClock": { limitMs: 30_000 }, "transport.envelopeBytes": { limit: 16_384 } },
});
const outcome = await runtime.runTurn({
agent,
ambient: { userId: "u-1", threadId: "t-1", scope: {} },
envelope: { threadId: "t-1", clientTurnId: "test-1", message: "hello", attachments: [] },
budget: { steps: 4 },
});
assert.equal(outcome.kind, "completed");
assert.ok(runtime.events.some((e) => e.kind === "gate" && e.gate === "close" && e.verdict === "pass"));Because the fake sits behind the same interface as a live provider, the test
exercises the real runtime — real gates, real budgets, real chain policy —
with only the text generation scripted. A fake that returns a call event
drives the real tool executor; a fake that throws a TierError proves your
chain's retry and fallback policy; a fake with stream exercises the live
drain. For time-budget tests, inject your own clock: Clock is a one-method
interface ({ now(): number }), so a manually advanced implementation passed
as RuntimeConfig.clock makes a wall-clock test run in microseconds.
Assert behavior, not just settlement: check the outcome's kind, then the
evidence — outcome.payload, the records from await runtime.store.records(),
and the bus events in runtime.events. For a search that finds nothing, check
that the answer says so; for a state write, check the stored block. Completion
alone proves neither.
Fakes stay out of production
A fake tier belongs in tests — never in a deployed chain, where scripted content answering real users would be a synthetic success. The Model chapter carries the same warning from the production side.
Run the suites
From the repository root:
pnpm --filter @keel-dev/core test
pnpm --filter example-assistant testBun runs these tests. The first checks the framework's own contracts; the
second checks the reference application's behavior. A scaffolded app runs its
own tests with npm test.
How the framework tests itself
The framework's tests live in packages/keel/test/, one folder per module —
state, wire, context, model, agent, turn, delegation — and each
suite pins the behavior its module promises, using the same technique this
chapter recommends: scripted ModelAdapter fakes driving the real runtime.
| Suite | What it pins |
|---|---|
state/ | The Store contract — block roundtrips, per-thread keying, listMessages window semantics — and block open/reduce/repair behavior. |
wire/ | Part registration, fixtures, marker promotion, and the reserved outcome part. |
context/ | Assembly order, section budgets, and render clipping. |
model/ | Chain retry and fallback classes, capacity validation, and provider request shapes. |
agent/ | Build-time declaration checks — undeclared reads, tool/tier compatibility, byproduct bindings. |
turn/ | The four gates, budgets, deliverables, streaming, tools, and exhaustion — each refusal driven end to end through runTurn. |
delegation/ | Caller/child ownership: grants, forwarding, and failure feedback. |
test/turn/gates.test.ts is a good first read: an oversized context settles
failed with fit-overflow and zero steps used; an open marker settles
marker-open; narration-only output settles empty. Every check is a real
turn with a scripted tier — no live provider, no flakiness.
Add a regression test
When a bug appears, reproduce it with a fake tier or an injected failure,
assert the observable result that matters to the caller, and keep the test
next to your app (test/, as in examples/assistant/test/ and the scaffolded
starter's test/app.test.ts). The assistant's workspace-search.test.ts
shows the shape: run a turn, assert the root outcome, find the child's record,
and check the wire events for what did — and did not — reach the client.
Mocks test the runtime; evals test the prompts
A scripted model verifies how the runtime responds to known model actions. Whether your prompts reliably produce those actions is a separate question — cover it with live-model evaluation.
That completes the tour. Head back to the Introduction to retrace any concept, or explore the source and examples on GitHub.
