Files
deepseek-harness/docs/core-data-structures/workflow.md
T
Tianyi Cui 2accf85714 workflow: simplify to the trust premise; settle result on cancellation
Two review responses that belong together — the same review argued the
engine was defending the wrong threat while a benign-input bug wedged
the product.

1) Drop hostile-value containment; state the trust premise.

Scripts are model-written — the same trust level as the model's bash
access — yet successive pre-push review rounds had ratcheted in defenses
that only matter against an adversarial author: trap-free proxy
rejection, accessor-never-invoked descriptor walks, realm-side
pre-rendering of thrown values, realm-built promises/arrays/error clones
with structural fatal recognition. That same author keeps a documented,
accepted, unkillable event-loop spin, so containing its error VALUES is
cost without a threat model — and the planned hardened engine
(worker/isolated-vm) gets value isolation by serialization and deletes
all of this machinery anyway.

What stays, because benign scripts hit it constantly: result never
rejects; dropped hook promises cannot become unhandled rejections; the
value boundary rejects LOUD everything JSON cannot carry (now a plain
recursive walk — getters are read ordinarily and their result is what
crosses; a throwing read fails loud); a "__proto__" key still copies as
a data property; the fatal-vs-null combinator discipline (now host
instanceof — unforgeable from the realm and simpler than clone-shape
recognition). What changes for scripts (documented in the engine
README): hooks hand back host values and host errors — in-script
`instanceof Error` on a hook failure is false (branch on e.name/e.code)
— and args are host-cloned once so a script cannot mutate the caller's
object. realm.ts drops 289 → 173 lines; the hostile-value test tables go
with it. The premise now leads the engine module doc, the README, and
the RFC's engine section, with the removed machinery recorded under
What was rejected.

2) result settles within the dispose grace of a cancellation.

Review finding (verified through the real registry + tool + engine): a
script parked on a promise no hook owns — `await new Promise(() => {})`,
`await Promise.race([])`, a returned never-settling thenable — could not
be settled by cancel(): hooks reject and children abort, but nothing
touches a promise the engine does not own, so `result` stayed pending
FOREVER (the previous cut even pinned that as intended). The tool awaits
run.result BEFORE its disposing finally, the registry awaits the tool,
the loop awaits the registry — one such script wedged the whole agent
turn past any abort, unrecoverable in-process; the mock engine in the
tool's abort test settles result on cancel, which is exactly the
behavior the real engine lacked, so no existing test could see it.

The seam contract now says it out loud: once a run is cancelled, result
SETTLES within the implementation's bounded grace even if the script
never does. The vm engine arms an abandon channel in cancel(); drive()
races the script against it, force-settling 'cancelled' at the grace
(the abandoned settlement stays contained; a post-slice synchronous spin
remains the documented limitation). dispose()'s outer race now exists
for child quiescence only, and `workflow/end` again fires exactly once
per started run. The old 'result stays pending' pin is FLIPPED to the
new contract (the pinned behavior was the bug); new regressions cover
cancel-then-settle on a parked script, a never-settling returned
thenable, and the full composition through the REAL registry + tool +
vm engine (tool-workflow gains workflow-vm/subagent devDeps for it).
agentsStarted JSDoc clarified while touching the vocabulary (accepted
calls, including ones still queued at cancellation).
2026-07-06 00:48:49 +08:00

5.0 KiB

Workflow

The workflow seam — an agent running a model-written orchestration SCRIPT that fans out subagents. Like subagent it is one optional capability, not part of the agent-loop spine, so its vocabulary lives here rather than in core.md. Unlike the subagent registry it takes the bash shape: ONE engine implementation per context provides ctx.workflows; there is no named-provider registry (a second engine is a plugin swap, not a co-resident).

Interface: dsh-workflow (ctx.workflows + the vocabulary below). The implementation is dsh-workflow-vm (an in-process node:vm engine); the model-facing consumer is dsh-tool-workflow. The proposal and rationale: the dynamic-workflows RFC.

Source: packages/workflow/workflow/src/types.ts

The start request

What a caller asks for when starting a run. The tool layer builds this from the model's { script, args } plus the calling agent; the engine validates the script's meta block BEFORE the body runs. parent is REQUIRED — every child the script spawns is attributed to it (cwd, lineage, and depth flow through the subagent seam). args must be plain host-realm JSON data; the engine exposes it to the script as the args global.

interface WorkflowStartRequest {
  script: string
  args?: unknown
  parent: Agent
  signal?: AbortSignal
}

The script's identity: WorkflowMeta

The validated export const meta block (Claude Code dynamic-workflows format — a PURE object literal heading the script). phases is progress vocabulary only: phase() calls match titles for observers; no execution structure is implied.

interface WorkflowMeta {
  name: string
  description: string
  whenToUse?: string
  phases?: WorkflowPhase[]
}

The terminal result: WorkflowResult

The outcome of one run, resolved by WorkflowRun.result. value is the script's materialized return value — plain host-realm JSON data (null when the script returned nothing) — meaningful only for completed. stopReason is a CLOSED union (engine-owned; consumers may exhaust it): completed | cancelled | error. A non-completed reason carries the failure in error, and the consumer maps it to an isError tool result rather than reporting partial output as success.

interface WorkflowResult {
  value: unknown
  stopReason: WorkflowStopReason
  error?: string
  agentsStarted: number
}

A live run: WorkflowRun

The handle the consumer holds while a script executes. The consumer awaits result, may cancel mid-flight, and MUST dispose on every path. result does NOT reject — a script failure resolves with stopReason: 'error' — and once the run is cancelled it SETTLES within the engine's bounded grace even if the script itself never settles (the engine abandons the script and reports cancelled), so a consumer awaiting result is never wedged past a cancellation. dispose() = cancel + that bounded settle + child quiescence (the engine documents what abandonment leaves behind); it never hangs on a stuck script.

interface WorkflowRun {
  readonly id: WorkflowRunId
  readonly meta: WorkflowMeta
  readonly result: Promise<WorkflowResult>
  cancel(reason?: string): void
  dispose(): Promise<void>
}

Failure discipline: WorkflowError.fatal

Hook misuse inside a script — bad arguments, unknown/deferred agent() options, a schema outside the structured-output subset, a tripped cap, a seam start failure, cancellation — throws a WorkflowError with fatal: true. The parallel()/pipeline() combinators RE-THROW fatal errors instead of mapping the item to null: a typo'd option must kill the script loudly, never dissolve into something that reads as an ordinary child failure. The per-item null is reserved for child-run failures (a non-completed stop reason) and ordinary in-stage script errors.

Events

The workflow/* events (workflow/start, workflow/phase, workflow/log, workflow/agent-start, workflow/agent-end, workflow/end — see the events catalog) are observe-only emits carrying DATA SNAPSHOTS: every payload starts with WorkflowRunInfo (id + meta), never the live WorkflowRun, so a subscriber cannot gain cancel/dispose, and workflow/end deliberately omits the result value (a listener observing outcomes must not receive a mutable alias of the caller's result). Every emit is per-listener contained — a throwing subscriber is logged, never propagated, and cannot starve the listeners registered after it — and every listener receives its own payload clone, so mutating it corrupts neither the engine nor other listeners; the containment mirrors subagent/start/subagent/end.