keel
Model

Custom adapters

Implement a model tier and test chain behavior without provider calls.

A tier implements the exported ModelAdapter interface. You can use that interface for an additional provider or a deterministic local test double — a test double in Keel is just a ModelAdapter you write.

For an unlisted model using a supported provider API, use customModel with a shipped provider class. Implement an adapter when you need another API protocol or a local test double.

Build a minimal local tier

// src/agents/assistant/assistant.mock.ts
import { defineChain, type ModelAdapter } from "@keel-dev/core";

export const greetingTier: ModelAdapter = {
  id: "local-greeting",
  windowTokens: 8_000,
  maxOutputTokens: 256,
  generate: () => ({
    events: [{ type: "text", text: "Hello from the local adapter." }],
    finish: "stop",
  }),
};

export const localChain = defineChain({
  id: "local-chain",
  tiers: [greetingTier],
});

This adapter serves a text-only agent with tools: []. Its fixed response does not change when you edit the prompt. Do not declare tool support just to bypass the capability check.

Understand the input

GenerateContext fieldContent
systemAssembled opening sections, excluding the separate user section, plus any delivery instructions.
briefThe current request.
toolsAvailable tool names, descriptions, and input schemas.
transcriptThis turn's assistant steps, tool results, and recovery nudges.
stepThe runtime's step index for this request.
forceToolOptional requested tool constraint during delivery recovery.
maxOutputTokensOutput reserve supplied by the chain for this request, already clamped to any turn-level output grant.

Adapters receive rendered context, not the raw state store. A live adapter must translate these fields into the provider's message and tool formats, and honor maxOutputTokens as its output ceiling. Direct callers may omit it; use the adapter's configured ceiling in that case.

Return normalized events

ModelEvent has exactly two variants. Text uses { type: "text", text }. A tool call uses { type: "call", tool, input, callId? } with finish: "tool-calls" when the step requests tool execution. There is no other event kind. Return usage when you can report it accurately.

An adapter supporting tools must also interpret the transcript on later steps. Repeating the same tool call forever is not a complete implementation. The runtime's budgets remain the boundary on that loop.

Stream live output

stream is the optional second half of the adapter interface:

stream?(ctx: GenerateContext): AsyncIterable<StreamEvent>;

When a tier implements it, the runtime prefers it over generate: each { type: "text" } delta reaches the wire as it arrives instead of after the full step. The runtime's drain holds back open and partially formed markers, so a marker split across chunks still promotes whole.

StreamEvent is ModelEvent plus a terminal frame:

type StreamEvent =
  | ModelEvent
  | { type: "finish"; finish: StepResult["finish"]; usage?: StepResult["usage"] };

Yield content events in order, then exactly one finish. A complete streaming tier:

// src/agents/assistant/assistant.stream.ts
import {
  consumeStream,
  type GenerateContext,
  type ModelAdapter,
  type StreamEvent,
} from "@keel-dev/core";

async function* doStream(ctx: GenerateContext): AsyncIterable<StreamEvent> {
  yield { type: "text", text: "Answering: " };
  yield { type: "text", text: ctx.brief };
  yield {
    type: "finish",
    finish: "stop",
    usage: { inputTokens: 24, outputTokens: 8, estimated: true },
  };
}

export const streamingTier: ModelAdapter = {
  id: "local-streaming",
  windowTokens: 8_000,
  maxOutputTokens: 256,
  stream: doStream,
  generate: (ctx) => consumeStream(doStream(ctx)),
};

consumeStream is exported for exactly this pattern: implement generate by collecting your own stream, so both paths share one implementation. It forwards text deltas through an optional callback and folds the events and finish frame into a normalized StepResult. The shipped provider adapters are built the same way.

Stream failure policy

Where a stream throws decides what the chain may do with it:

  • Before any content event: the throw follows the normal retry and fallback classes. Classify it with tierError like a generate failure; nothing reached the wire, so a retry is safe.
  • After a content event: the step is folded into finish: "error-post" and passed through. The chain never retries it — deltas already reached the wire, and a retry would duplicate output. The runtime cuts the turn with the typed stream-error reason.

Inject a failure

To exercise fallback deterministically, replace generate with a classified error and put a working local tier after it:

import { tierError } from "@keel-dev/core";

const unavailable = {
  ...greetingTier,
  id: "local-unavailable",
  generate: () => {
    throw tierError("unavailable", "deliberate local outage");
  },
};

A transient error exercises the same-tier retry; unavailable exercises immediate fallback. Test exhaustion with only failing tiers.

Inspect in Studio

Connect the local chain to your application's text-only assistant and run it in Studio. Verify the greeting and selected tier, then run the failure variants and inspect the chain events and final outcome.

Keep deterministic tiers in local and test configurations. See the testing guide for broader application checks.

Next: Context.