Files
deepseek-harness/docs/core-data-structures/tools.md
T
Yichen Jiang 740106a51a Merge remote-tracking branch 'origin/master' into codex/rfc-subagent-background-tasks
# Conflicts:
#	docs/config-catalog.md
#	docs/cordis-catalog/services.md
#	docs/core-data-structures/bash.md
#	docs/core-data-structures/core.md
#	docs/event-producer-consumer.md
#	docs/module-graph.md
#	docs/tool-catalog.md
#	examples/acp-agent/tests/snapshots/both-mode-turn/session.jsonl
#	examples/acp-agent/tests/snapshots/code-mode-turn/session.jsonl
#	examples/acp-agent/tests/snapshots/text-turn/session.jsonl
#	packages/README.md
#	packages/bash/README.md
#	packages/bash/bash-local/src/index.ts
#	packages/bash/bash/README.md
#	packages/bash/bash/package.json
#	packages/bash/bash/src/index.ts
#	packages/bash/bash/src/types.ts
#	packages/bash/bash/tests/service.spec.ts
#	packages/bash/bash/tsconfig.json
#	packages/bash/tool-bash/README.md
#	packages/bash/tool-bash/src/index.ts
#	packages/bash/tool-bash/tests/tools.spec.ts
#	packages/bash/tool-bash/tsconfig.json
#	packages/cordis/tool-cordis/src/api-catalog.ts
#	packages/core/agent-core/README.md
#	packages/core/agent-core/package.json
#	packages/core/agent-core/src/index.ts
#	packages/core/agent-core/tests/agent-core.spec.ts
#	packages/core/tools/tests/gen-tool-catalog.spec.ts
#	packages/hooks/hook-protocol/tests/runner.spec.ts
#	packages/ui/acp-agent/tests/acp-agent.spec.ts
#	packages/ui/stdio-agent/tests/stdio-agent.spec.ts
#	pnpm-lock.yaml
#	scripts/doc-budgets.manifest.json
#	scripts/gen-tool-catalog.ts
2026-07-11 23:04:27 +08:00

12 KiB

Tools

The tool pipeline of dsh-tools. core.md introduces ToolDefinition as the one pipeline-authoring type promoted to the spine and ToolSchema as the model-facing wire shape. This page owns the full ToolDefinition, the typed schema DSL that builds it, the waterfall execution shapes, and the UI-presentation vocabulary.

Source: packages/core/tools/src/index.ts · packages/core/tools/src/schema.ts · packages/core/tools/src/presentation.ts

ToolDefinition — a registered tool

A ToolSchema (the model-facing fields) plus the execute function and optional UI presenters. The registry holds these; the loop dispatches calls through them. The registry's schemas() builds the model-facing ToolSchema[] by an explicit allowlist — execute/presentCall/presentResult must never leak into a model request.

interface ToolDefinition extends ToolSchema {
  execute(args: unknown, exec: ToolExecution): Promise<ToolExecuteReturn>
  /**
   * Cooperative tool-call timeout budget in milliseconds. Omit for no deadline.
   * Enforced by `@deepseek-ai/dsh-timeout-policy` (a `tools/execute` wrapper); it
   * is NEVER sent to the model — `schemas()` whitelists only name/description/
   * parameters. Declaring it asserts this tool forwards `exec.signal` to a
   * cooperative implementation that can reach quiescence when the signal aborts.
   */
  timeoutMs?: number
  /**
   * Optional: how to present the PENDING state of one call in a UI, derived from
   * the call's `args` (parsed arguments, `unknown` — the tool validates/narrows
   * its own input). Returns a {@link ToolCallView} (a `card`-tagged render intent),
   * or `undefined` (or omit the method) to fall back to a generic presentation
   * (title = tool name, raw args as input). Pure and side-effect-free: a UI may
   * call it during live streaming AND a session-log replay, so it must depend
   * only on `args`.
   */
  presentCall?(args: unknown): ToolCallView | undefined
  /**
   * Optional: how to present the COMPLETED state, given the same `args` and the
   * `result` (`execute`'s content + whether it errored). Returns a
   * {@link ToolResultView}, or `undefined` (or omit the method) to keep the
   * pending title and render the raw result content. Pure and side-effect-free
   * for the same replay reason.
   */
  presentResult?(args: unknown, result: ToolResult): ToolResultView | undefined
}

execute receives args: unknown — a raw ToolDefinition validates its own input. First-party tools don't write that by hand; they use defineTool, which validates and narrows for them.

The typed schema DSL

Plugin authors write per-property specs with a boolean required: true, and a type-level helper maps the spec to the execute argument type — zero casts. The DSL is machinery that types ToolDefinition; it is intentionally a sub-page detail, not core.

Source: packages/core/tools/src/schema.ts

interface SchemaProp {
  type: SchemaType
  /** Per-property required flag (NOT the JSON Schema top-level required array). */
  required?: true
  /** Human-readable description, surfaced in the JSON Schema as well. */
  description?: string
  /** Enum of allowed values (strings only). */
  enum?: string[]
  /** Default value. */
  default?: unknown
  /** Nested properties for type: 'object'. */
  properties?: SchemaSpec
  /** Items schema for type: 'array'. */
  items?: SchemaProp
}
type SchemaSpec = Record<string, SchemaProp>

SchemaType is the primitive union 'string' | 'number' | 'boolean' | 'object' | 'array'. InferArgs<S> maps a SchemaSpec to the TS argument type — required: true props become required keys, everything else genuinely optional:

type InferArgs<S extends SchemaSpec> = Simplify<
  & { [K in RequiredKeys<S>]: InferPropValue<S[K]> }
  & { [K in Exclude<keyof S, RequiredKeys<S>>]?: InferPropValue<S[K]> }
>

defineTool({ name, description, parameters, execute, … }) ties it together: parameters is a SchemaSpec, execute(args, exec) gets args: InferArgs<typeof parameters>, and the helper converts the spec to JSON Schema (schemaSpecToJsonSchema) for the wire and validates model-generated args (validateArgs) before the typed body runs. A mismatch throws ToolArgsError (code: 'INVALID_ARGS'), which the registry turns into an isError result so the model can self-correct. Why a custom DSL and not schemastery: tool parameters need JSON Schema (the LLM wire format), not validation/transformation — the lightweight DSL gives the best authoring DX with the smallest surface.

Execution: the tools/pre-execute / tools/post-execute pipeline shapes

ctx.tools.execute() runs each call through a two-waterfall pipeline — tools/pre-execute (the allow/deny/ask gate) → core dispatch → tools/post-execute (inspect/replace the result, attach context) — the seams where sandbox, permission, hook, and plan-mode plugins gate or transform a call. The pending call is a ToolExecution; the outcome is a ToolExecutionResult.

interface ToolExecution {
  callId: CallId
  name: string
  /** Parsed JSON arguments (unknown — tools validate their own input). */
  arguments: unknown
  /** The agent on whose behalf the call runs (set by the agent loop). */
  agent?: Agent
  signal?: AbortSignal
}
interface ToolExecutionResult {
  callId: CallId
  content: ContentBlock[]
  isError: boolean
  /**
   * Set when the call failed with a {@link HarnessError}: machine-routable
   * `{ name, code }` for retry/sandbox plugins and replay. The model-facing
   * text in `content` is always present; this is extra structure for code.
   */
  error?: ToolErrorInfo
  /**
   * Extra model-facing context a `tools/post-execute` listener attached for the
   * NEXT request (Claude Code's PostToolUse `additionalContext`). It is NOT part
   * of this call's `content` — `content`/`feedback` shape the tool RESULT, but
   * `additionalContext` is a SEPARATE `context/message`. A step can carry
   * multiple tool calls, so the loop BUFFERS every call's `additionalContext`
   * and appends them only AFTER all `tool/result`s for the step, keeping
   * tool-call/result adjacency intact. Carried on the result purely to ferry it
   * from `execute()` up to the loop's per-step buffer.
   */
  additionalContext?: HookContext
  /**
   * The tool-private presentation payload from a successful `execute` (the object
   * return form). Threaded onto the `tool/result` session event and back into
   * {@link ToolResult} for `presentResult`. Opaque (`unknown`); absent when the
   * tool attached none or the call failed.
   */
  meta?: unknown
}

Each interception waterfall returns a typed Decision (the idiom shared with the agent/* seams). tools/pre-execute listeners receive (exec, next) and return a PreToolDecision; tools/post-execute listeners receive (exec, result, next) and return a PostToolDecision:

type PreToolDecision =
  | { kind: 'allow' }
  | { kind: 'deny'; reason: string }
  | { kind: 'ask'; reason?: string }
type PostToolDecision =
  | { kind: 'accept'; content?: ContentBlock[]; additionalContext?: HookContext }
  | { kind: 'block'; feedback: ContentBlock[]; additionalContext?: HookContext }

Call next() to delegate to the default (allow / accept-unchanged), or return a decision to short-circuit. A pre-execute deny (or ask, which degrades to deny until the permission system lands) skips dispatch and yields an isError result; input rewrite is deliberately NOT offered on PreToolDecision (it would desync the pre-execution audit/history/UI from what ran — its own proposed RFC). A post-execute accept may replace the model-facing content (clean, because tool/result is logged after execute() returns); a block turns the call into an isError whose content is the corrective feedback. Core dispatch sits between the waterfalls as plain code; the tool body keeps its own try/catch so a thrown tool still reaches post-execute as an isError. An unregistered tool routes through the same catch as a tool-thrown error, so both failure classes get a structured { name, code } (ToolNotFoundErrorUNKNOWN_TOOL) — the loop records a failed tool call instead of failing the whole turn.

The structured-output schema subset

The vocabulary a caller uses to demand a machine-readable result from a subagent (SubagentStartRequest.outputSchema, subagent.md) or a workflow agent() call. It is deliberately NOT full JSON Schema: the schema travels verbatim to the model as a forced tool's parameters, and the produced value is validated client-side by validateStructuredValue — so every accepted keyword must be one the validator actually enforces, and assertSupportedOutputSchema rejects anything else loud (OutputSchemaError, listing every violation). Both walkers reason over own enumerable properties only (JSON carries nothing else) and reject non-plain objects (Date, Map) that would serialize lossily.

type StructuredScalar = string | number | boolean | null
type StructuredSchemaType = 'object' | 'array' | 'string' | 'number' | 'integer' | 'boolean' | 'null'
interface StructuredSchemaNode {
  type: StructuredSchemaType
  properties?: Record<string, StructuredSchemaNode>
  required?: string[]
  additionalProperties?: boolean
  items?: StructuredSchemaNode
  enum?: StructuredScalar[]
  const?: StructuredScalar
  description?: string
  title?: string
  default?: unknown
  examples?: unknown
}

A schema is an object-rooted node (enum/const are scalar-only; description/title/default/examples are annotations, allowed and ignored but still required to be JSON data — they ride the wire):

type StructuredOutputSchema = StructuredSchemaNode & { type: 'object' }

Tool-presentation UI vocabulary

How a tool wants its call shown in a UI (an editor tool-call card, a CLI log line), provider-neutral so a tool describes itself without depending on any client protocol. presentCall/presentResult return a card-tagged render intent — a discriminated union a UI bridge switches on:

  • ToolCallView (pending): { card: 'generic', title, kind?, rawInput?, content?, locations? } (the default card; locations is { path, line? }[] files the call reads/modifies, for editor follow-along), { card: 'terminal', title, description?, cwd? } (a shell command → a terminal card), or { card: 'diff', title, diffs, locations? } (a file create/modify → an inline diff card; diffs is { path, oldText, newText }[], oldText: null for a new file).
  • ToolResultView (completed): { card: 'generic', title?, content? }, { card: 'terminal', title?, output?, exitCode?, signal? } (the captured run output + exit; a capable UI shows an exit-status pill, an incapable one gets a fenced ```console fallback the bridge derives from output), or { card: 'diff', title?, diffs } (a completed file mutation → the change to show, typically the applied hunks with context lines computed from the before/after content, or a whole-file diff when there is no before-image — e.g. a file create. A tool_call_update's content REPLACES the call's content, so a mutation tool returns this even when it duplicates the call-time snippet, to keep the result from clobbering the diff with result text).

ToolCallKind ('read' | 'edit' | 'delete' | 'move' | 'search' | 'execute' | 'fetch' | 'other') picks an icon on a generic card. FileLocation ({ path, line? }) and FileDiff ({ path, oldText, newText }) are the shared file-card vocabulary. The design is pinned in the render-intent-union RFC; the ACP bridge maps a diff card to a { type: 'diff' } content block, a terminal card to the _meta terminal convention, and relativizes a file card's title against the session cwd.

The full presentation field docs live in packages/core/tools/src/presentation.ts. The bash schema and executor are on bash.md; generic background controls are on tasks.md.