Custom adapters
Implement a model tier and test chain behavior without provider calls.
A tier implements the exported ModelAdapter interface. You can use that
interface for an additional provider or a deterministic local test double —
a test double in Keel is just a ModelAdapter you write.
For an unlisted model using a supported provider API, use
customModel
with a shipped provider class. Implement an adapter when you need another API
protocol or a local test double.
Build a minimal local tier
// src/agents/assistant/assistant.mock.ts
import { defineChain, type ModelAdapter } from "@keel-dev/core";
export const greetingTier: ModelAdapter = {
id: "local-greeting",
windowTokens: 8_000,
maxOutputTokens: 256,
generate: () => ({
events: [{ type: "text", text: "Hello from the local adapter." }],
finish: "stop",
}),
};
export const localChain = defineChain({
id: "local-chain",
tiers: [greetingTier],
});This adapter serves a text-only agent with tools: []. Its fixed response
does not change when you edit the prompt. Do not declare tool support just
to bypass the capability check.
Understand the input
GenerateContext field | Content |
|---|---|
system | Assembled opening sections, excluding the separate user section, plus any delivery instructions. |
brief | The current request. |
tools | Available tool names, descriptions, and input schemas. |
transcript | This turn's assistant steps, tool results, and recovery nudges. |
step | The runtime's step index for this request. |
forceTool | Optional requested tool constraint during delivery recovery. |
maxOutputTokens | Output reserve supplied by the chain for this request, already clamped to any turn-level output grant. |
Adapters receive rendered context, not the raw state store. A live adapter
must translate these fields into the provider's message and tool formats,
and honor maxOutputTokens as its output ceiling. Direct callers may omit it;
use the adapter's configured ceiling in that case.
Return normalized events
ModelEvent has exactly two variants. Text uses { type: "text", text }.
A tool call uses { type: "call", tool, input, callId? } with
finish: "tool-calls" when the step requests tool execution. There is no
other event kind. Return usage when you can report it accurately.
An adapter supporting tools must also interpret the transcript on later steps. Repeating the same tool call forever is not a complete implementation. The runtime's budgets remain the boundary on that loop.
Stream live output
stream is the optional second half of the adapter interface:
stream?(ctx: GenerateContext): AsyncIterable<StreamEvent>;When a tier implements it, the runtime prefers it over generate: each
{ type: "text" } delta reaches the wire as it arrives instead of after the
full step. The runtime's drain holds back open and partially formed markers,
so a marker split across chunks still promotes whole.
StreamEvent is ModelEvent plus a terminal frame:
type StreamEvent =
| ModelEvent
| { type: "finish"; finish: StepResult["finish"]; usage?: StepResult["usage"] };Yield content events in order, then exactly one finish. A complete
streaming tier:
// src/agents/assistant/assistant.stream.ts
import {
consumeStream,
type GenerateContext,
type ModelAdapter,
type StreamEvent,
} from "@keel-dev/core";
async function* doStream(ctx: GenerateContext): AsyncIterable<StreamEvent> {
yield { type: "text", text: "Answering: " };
yield { type: "text", text: ctx.brief };
yield {
type: "finish",
finish: "stop",
usage: { inputTokens: 24, outputTokens: 8, estimated: true },
};
}
export const streamingTier: ModelAdapter = {
id: "local-streaming",
windowTokens: 8_000,
maxOutputTokens: 256,
stream: doStream,
generate: (ctx) => consumeStream(doStream(ctx)),
};consumeStream is exported for exactly this pattern: implement generate
by collecting your own stream, so both paths share one implementation. It
forwards text deltas through an optional callback and folds the events and
finish frame into a normalized StepResult. The shipped provider adapters
are built the same way.
Stream failure policy
Where a stream throws decides what the chain may do with it:
- Before any content event: the throw follows the normal
retry and fallback classes. Classify it
with
tierErrorlike ageneratefailure; nothing reached the wire, so a retry is safe. - After a content event: the step is folded into
finish: "error-post"and passed through. The chain never retries it — deltas already reached the wire, and a retry would duplicate output. The runtime cuts the turn with the typedstream-errorreason.
Inject a failure
To exercise fallback deterministically, replace generate with a classified
error and put a working local tier after it:
import { tierError } from "@keel-dev/core";
const unavailable = {
...greetingTier,
id: "local-unavailable",
generate: () => {
throw tierError("unavailable", "deliberate local outage");
},
};A transient error exercises the same-tier retry; unavailable exercises
immediate fallback. Test exhaustion with only failing tiers.
Inspect in Studio
Connect the local chain to your application's text-only assistant and run it in Studio. Verify the greeting and selected tier, then run the failure variants and inspect the chain events and final outcome.
Keep deterministic tiers in local and test configurations. See the testing guide for broader application checks.
Next: Context.
