Tools
Give the model capabilities that validate their own input, declare who may call them, and fail on your terms — never silently.
Here is an incident shape worth remembering. A model calls a filing tool
with priority: "urgent" — a value the API doesn't accept. The glue code
helpfully coerces it to a string, the request files as garbage, and the turn
reads as a success. Its cousin: a prompt says "only workspace members may
search the workspace", and one sampled reply searches anyway.
The lesson from both is the same: a capability is only as safe as its border. In keel, a tool is a declaration that carries its whole border with it:
- typed input — a zod schema is enforced at the executor; a call that doesn't parse is rejected with the reason, never coerced,
- declared access — who may call it is a rule on the tool, checked by the framework, never a sentence in the prompt,
- an explicit error policy — retries and what a thrown execution becomes are the caller's declared terms,
- declared liveness — a long-running tool states how long it may go silent before the watchdog cuts it.
The typed border
When the model's arguments fail the schema, the rejection — with the exact
field-level reasons — returns as the call's result, and the model corrects
its own call in the same turn. Your execute only ever runs on input that
parsed. A delegation call is guarded the same way: invalid input never opens
a child turn.
Shared capabilities live in src/shared/tools/<name>.tool.ts (agent-private
ones can live in the agent's own folder). The reference app's smallest tool:
// src/shared/tools/now.tool.ts
import { z } from "zod";
import { defineTool } from "@keel-dev/core";
export const nowTool = defineTool({
id: "now",
description: "Current date and time",
access: "public",
input: z.object({}),
execute: () => ({ result: { now: new Date().toISOString() } }),
});The timestamp is produced by tool code at execution — the model reports time rather than inventing it. Keep contracts this explicit when you connect real services.
Typing the input in your code
The runtime hands execute the parsed value, but the signature types it
as unknown. When your code needs typed fields, parse with the same
schema inside the callback: const { query } = inputSchema.parse(raw).
Access: declared once, enforced twice
access is required — defineTool throws without it. Either the tool is
explicitly "public", or it names a predicate over the ambient context
{ userId, threadId, scope } — the shared core plus the scope your app
declared with defineAmbient and keel validated at the admit gate:
type AccessRule<S> =
| "public"
| { id: string; allows: (ambient: AmbientContext<S>) => boolean | string };The predicate's return value carries the verdict:
allows(ambient) returns | What happens |
|---|---|
true | the call is allowed |
false | refused as refused by access rule "<id>" |
| a string | refused, with that string as the reason |
The reference app's workspace search names its rule and its reason in one
line — the scope field it reads (workspaceId) is typed by the app's own
ambient declaration:
// src/shared/tools/workspace-search.tool.ts
access: {
id: "workspace-members",
allows: (a) => Boolean(a.scope.workspaceId) || "requires a workspace",
},A rule is a pure function of the ambient context — no clock, no store, no network. That makes it matrix-printable: evaluate every rule over your app's ambient fixtures in a test and you have the access matrix, checked in CI instead of reverse-engineered during an incident.
The rule is enforced in two places, not one:
A barred tool's definition never enters the model's window — it cannot be
"talked into" using a tool it cannot see. And if the model hallucinates the
call anyway, the executor refuses it with a typed defect. The turn record
keeps both lists (toolsAvailable, toolsFiltered), so support reads the
resolved set instead of the logs.
Ambient context is yours to earn
Access rules evaluate the ambient context your application supplies. The framework validates the scope's shape at the admit gate and enforces the rule; authenticating the request and building trusted scope values remains your job.
Failure on the caller's terms
A thrown execute is not an exception that escapes — it walks a ladder you
declared:
| Setting | Behavior |
|---|---|
retries: 0 | default: no retry after a thrown execution |
retries: 1 | re-run a thrown execution once before giving up |
onError: "feedback" | default: the final failure reason returns to the model as the call's result |
onError: "fail" | the defect settles the caller's whole turn as failed |
Use "feedback" when the model can degrade gracefully ("the lookup failed,
here is what I know without it"). Use "fail" for tools whose failure
invalidates the turn — a state commit, a payment.
Retries repeat side effects
A retry re-runs your execute. Before declaring retries on a tool with
external effects, make sure repeating the operation is safe.
Long work declares its heartbeat
execute receives a third argument — the executor's control plane:
execute: (input, ambient, ctx) => { /* ctx.signal, ctx.progress */ }ctx.progress(data)emits a progress part to the wire and resets the liveness silence window — one call serves the UI and the watchdog,ctx.signalis anAbortSignalthat fires when the watchdog trips, the turn's time budget runs out, or the user cancels.
A tool that takes real time declares duration: "long" and a liveness
ceiling — defineTool refuses a long tool without one, and the ceiling is
enforced whenever it is present:
// src/shared/tools/workspace-search.tool.ts
duration: "long",
liveness: { maxSilenceMs: livenessMs.workspaceSearch }, // 20s, from budgets.ts
input: z.object({ query: z.string().default("") }),
execute: (_input, _ambient, ctx) => {
ctx.progress({ step: "scanning workspace" });
ctx.progress({ step: "ranking results" });
return { result: { hits: [/* … */] } };
},The watchdog is real wall clock, not a suggestion: if the tool goes silent
— no ctx.progress call — past maxSilenceMs, the executor aborts
ctx.signal and the turn settles cut by "watchdog". A hung tool no
longer hangs the turn. Independently of liveness, every tool call is capped
by the turn's remaining time budget — a tool without a ceiling still cannot
out-sleep the wall clock.
Abortion is cooperative: pass ctx.signal into your IO
(fetch(url, { signal: ctx.signal })) so the underlying work actually
stops when the executor gives up on it. The ceiling itself lives in
budgets.ts with every other number in the app.
Where tools meet State and other agents
A tool run returns { result, commit? }. The successful result can land
in a block two ways — both are byproducts of tool code, never model
actions:
feeds: { block, event }routes the whole result through a declared reducer. The binding is build-gated:defineAgentthrows unless the agent declares that block inwritesand the block declares that reducer.- a returned
commit: { block?, event, payload }picks the payload explicitly. It is validated at runtime, and a refusal is never silent: naming an undeclared block, an unknown reducer, or a writer already taken this turn appendsSTATE COMMIT REFUSED (…): reasonto the tool's own result — the model sees it and reacts. UnderonError: "fail", the refused commit fails the whole turn instead.
A feeds block with a ttlMs also turns the tool into a cache: while the
last write is fresh, the tool is not executed — the executor answers
with { servedFromState: "<block>", value } instead. Note the shape: a
cache hit does not mimic the tool's own result, it says where the value
came from. State covers blocks, reducers, and the single
writer.
A specialist agent is also exposed as a tool:
asTool(researcherAgent(keys), { budget: grants.researcher })The call opens a child turn under the caller's budget — Agents covers the wiring and the deliverable contract.
Next: The turn — the four gates every invocation crosses.
