keel
Harness

Completion and deliverables

Require a usable result before a turn is allowed to complete.

Missing result
"I finished the research."

findings submitted: false
Validated result
submit_findings({
  items: ["source A"]
})
Missing → bounded nudge → recheck → delivered or failed
The caller defines the required result. The close gate checks whether that result exists.

A provider finishing generation is not enough to complete a turn. The harness checks how generation ended, whether structured output is complete, and whether the caller's required result exists.

Understand the baseline check

Without a deliverable contract, completion requires a clean stop and nonempty text or a registered content part. Progress and ephemeral parts do not count. Parts forwarded from a child do not count as the caller's own substance.

This is a structural check, not a judge of answer quality. Plain text such as “I finished” may pass the baseline. Use a deliverable contract when the task requires a specific result rather than a statement that work happened.

Declare a required result

import { asTool, definePart } from "@keel-dev/core";
import { z } from "zod";

const findingsPart = definePart(registry, {
  name: "findings",
  version: 1,
  schema: z.object({ items: z.array(z.string()).min(1) }),
  persist: "content",
  merge: "replace-by-id",
  fixture: { items: ["one validated example"] },
});

// researcher must declare this same part in its emits.
const research = asTool(researcher, {
  budget: { steps: 6 },
  deliverable: { part: findingsPart, nudges: 1 },
});

The registry and researcher come from your app. asTool checks that the child declares the part, that it persists as content, and that its derived submission-tool name does not conflict with an existing tool. One more check waits at run time: the contract's part must be registered in the runtime's PartRegistry, or runTurn throws a KeelError before the turn opens.

Validate, recover, and check again

The runtime exposes submit_findings using the part's schema. Invalid data returns validation feedback. The contract itself is "the validated part exists on this turn's wire": a schema-valid submission fulfills it, and so does a promoted <findings> marker of the contract's part name — the submit tool is the derived, forceable door, not the only one. A re-submission replaces the previous one, with a warn event on the bus.

A deliverable is also substance in its own right: a turn that submits its contracted part and says nothing else still completes — the one intended exception to the empty close check, which otherwise refuses a turn with no prose and no content part.

If the child stops without delivering, the harness grants the configured nudge rounds — model calls beyond budget.steps, bounded by nudges instead. During those rounds only the submission tool is accepted by the executor; every other call is refused with that reason, and tiers supporting forced tool selection receive the constraint too.

If the required part is still missing, the child fails with deliverable-missing. The plain deliverable: findingsPart form grants one nudge by default. A validated result still must pass the other close checks.

Keep the result with its owner

A completed contracted turn carries a branded Delivered<T>. A direct call to runTurn with a typed deliverable contract returns DeliveredTurnOutcome<T>, so narrowing to completed exposes the validated data without a nullable deliverable.

For delegation, the caller receives { ok: true, deliverable: ... } privately and resumes its own work. The child's completion does not complete the root. Explicit forwarding changes visibility, not ownership or the completion rule.

Recognize other endings

ConditionOutcome
Length limit with unsuccessful or disabled continuationtruncated
Open marker after a normal stopfailed with marker-open
No baseline substancefailed with empty
User stopstopped
Silent tool, time ceiling, or failed streamcut — by is watchdog, timeout, or stream-error

A successful schema check does not establish factual correctness. Encode structural requirements in the schema, and cover task quality with behavior tests and live-model evaluation — see Testing.

Inspect in Studio

Compare the researcher's WIRE and TURN RECORDS for a valid submission, an invalid submission followed by correction, and a missing deliverable. Then inspect the caller's separate continuation. Its final outcome should reflect its own work, not merely the existence of a completed child.

Next: Tracing and inspection.