keel

Tools

Give the model capabilities that validate their own input, declare who may call them, and fail on your terms — never silently.

Here is an incident shape worth remembering. A model calls a filing tool with priority: "urgent" — a value the API doesn't accept. The glue code helpfully coerces it to a string, the request files as garbage, and the turn reads as a success. Its cousin: a prompt says "only workspace members may search the workspace", and one sampled reply searches anyway.

The lesson from both is the same: a capability is only as safe as its border. In keel, a tool is a declaration that carries its whole border with it:

  • typed input — a zod schema is enforced at the executor; a call that doesn't parse is rejected with the reason, never coerced,
  • declared access — who may call it is a rule on the tool, checked by the framework, never a sentence in the prompt,
  • an explicit error policy — retries and what a thrown execution becomes are the caller's declared terms,
  • declared liveness — a long-running tool states how long it may go silent before the watchdog cuts it.

The typed border

modelcall search {"query"…}input schemazod parseparsedexecuteyour coderesult→ transcript · feeds Staterejected — the reason returns as the call's resulta throwing execute: retries on the caller's terms, thenonError — feedback (default) or fail the whole turn
Every tool call crosses a typed border. Invalid input bounces back with the reason — the model fixes its own call; tool code only ever sees parsed input.

When the model's arguments fail the schema, the rejection — with the exact field-level reasons — returns as the call's result, and the model corrects its own call in the same turn. Your execute only ever runs on input that parsed. A delegation call is guarded the same way: invalid input never opens a child turn.

Shared capabilities live in src/shared/tools/<name>.tool.ts (agent-private ones can live in the agent's own folder). The reference app's smallest tool:

// src/shared/tools/now.tool.ts
import { z } from "zod";
import { defineTool } from "@keel-dev/core";

export const nowTool = defineTool({
  id: "now",
  description: "Current date and time",
  access: "public",
  input: z.object({}),
  execute: () => ({ result: { now: new Date().toISOString() } }),
});

The timestamp is produced by tool code at execution — the model reports time rather than inventing it. Keep contracts this explicit when you connect real services.

Typing the input in your code

The runtime hands execute the parsed value, but the signature types it as unknown. When your code needs typed fields, parse with the same schema inside the callback: const { query } = inputSchema.parse(raw).

Access: declared once, enforced twice

access is required — defineTool throws without it. Either the tool is explicitly "public", or it names a predicate over the ambient context { userId, threadId, scope } — the shared core plus the scope your app declared with defineAmbient and keel validated at the admit gate:

type AccessRule<S> =
  | "public"
  | { id: string; allows: (ambient: AmbientContext<S>) => boolean | string };

The predicate's return value carries the verdict:

allows(ambient) returnsWhat happens
truethe call is allowed
falserefused as refused by access rule "<id>"
a stringrefused, with that string as the reason

The reference app's workspace search names its rule and its reason in one line — the scope field it reads (workspaceId) is typed by the app's own ambient declaration:

// src/shared/tools/workspace-search.tool.ts
access: {
  id: "workspace-members",
  allows: (a) => Boolean(a.scope.workspaceId) || "requires a workspace",
},

A rule is a pure function of the ambient context — no clock, no store, no network. That makes it matrix-printable: evaluate every rule over your app's ambient fixtures in a test and you have the access matrix, checked in CI instead of reverse-engineered during an incident.

The rule is enforced in two places, not one:

tool + accessplan · tenancy · sourcewindow filterassemble()renderedexecutor checkrunTurnfiltered — never in contextrefused — tool-refused defect
Access is enforced twice: a barred tool's definition never reaches the window, and the executor refuses it anyway if the model hallucinates the call.

A barred tool's definition never enters the model's window — it cannot be "talked into" using a tool it cannot see. And if the model hallucinates the call anyway, the executor refuses it with a typed defect. The turn record keeps both lists (toolsAvailable, toolsFiltered), so support reads the resolved set instead of the logs.

Ambient context is yours to earn

Access rules evaluate the ambient context your application supplies. The framework validates the scope's shape at the admit gate and enforces the rule; authenticating the request and building trusted scope values remains your job.

Failure on the caller's terms

A thrown execute is not an exception that escapes — it walks a ladder you declared:

SettingBehavior
retries: 0default: no retry after a thrown execution
retries: 1re-run a thrown execution once before giving up
onError: "feedback"default: the final failure reason returns to the model as the call's result
onError: "fail"the defect settles the caller's whole turn as failed

Use "feedback" when the model can degrade gracefully ("the lookup failed, here is what I know without it"). Use "fail" for tools whose failure invalidates the turn — a state commit, a payment.

Retries repeat side effects

A retry re-runs your execute. Before declaring retries on a tool with external effects, make sure repeating the operation is safe.

Long work declares its heartbeat

execute receives a third argument — the executor's control plane:

execute: (input, ambient, ctx) => { /* ctx.signal, ctx.progress */ }
  • ctx.progress(data) emits a progress part to the wire and resets the liveness silence window — one call serves the UI and the watchdog,
  • ctx.signal is an AbortSignal that fires when the watchdog trips, the turn's time budget runs out, or the user cancels.

A tool that takes real time declares duration: "long" and a liveness ceiling — defineTool refuses a long tool without one, and the ceiling is enforced whenever it is present:

// src/shared/tools/workspace-search.tool.ts
duration: "long",
liveness: { maxSilenceMs: livenessMs.workspaceSearch }, // 20s, from budgets.ts
input: z.object({ query: z.string().default("") }),
execute: (_input, _ambient, ctx) => {
  ctx.progress({ step: "scanning workspace" });
  ctx.progress({ step: "ranking results" });
  return { result: { hits: [/* … */] } };
},

The watchdog is real wall clock, not a suggestion: if the tool goes silent — no ctx.progress call — past maxSilenceMs, the executor aborts ctx.signal and the turn settles cut by "watchdog". A hung tool no longer hangs the turn. Independently of liveness, every tool call is capped by the turn's remaining time budget — a tool without a ceiling still cannot out-sleep the wall clock.

Abortion is cooperative: pass ctx.signal into your IO (fetch(url, { signal: ctx.signal })) so the underlying work actually stops when the executor gives up on it. The ceiling itself lives in budgets.ts with every other number in the app.

Where tools meet State and other agents

A tool run returns { result, commit? }. The successful result can land in a block two ways — both are byproducts of tool code, never model actions:

  • feeds: { block, event } routes the whole result through a declared reducer. The binding is build-gated: defineAgent throws unless the agent declares that block in writes and the block declares that reducer.
  • a returned commit: { block?, event, payload } picks the payload explicitly. It is validated at runtime, and a refusal is never silent: naming an undeclared block, an unknown reducer, or a writer already taken this turn appends STATE COMMIT REFUSED (…): reason to the tool's own result — the model sees it and reacts. Under onError: "fail", the refused commit fails the whole turn instead.

A feeds block with a ttlMs also turns the tool into a cache: while the last write is fresh, the tool is not executed — the executor answers with { servedFromState: "<block>", value } instead. Note the shape: a cache hit does not mimic the tool's own result, it says where the value came from. State covers blocks, reducers, and the single writer.

A specialist agent is also exposed as a tool:

asTool(researcherAgent(keys), { budget: grants.researcher })

The call opens a child turn under the caller's budget — Agents covers the wiring and the deliverable contract.

Next: The turn — the four gates every invocation crosses.