Files
deepseek-harness/packages/web/tool-web

@deepseek-ai/dsh-tool-web

The model-facing web tool suite — web_search and web_fetch — over the web capability seam (ctx.web). It owns model-facing concerns only: tool names, JSON schemas, snake_case argument names, prompt sections, the result-count bound, result formatting, HTML→markdown presentation, and presentCall. All web access goes through ctx.web; this package never imports a concrete provider. Neither tool exposes a model-facing timeout — each tool's cooperative tool-call budget is declared here via config (fetchTimeoutMs/searchTimeoutMs, attached as ToolDefinition.timeoutMs) and enforced by @deepseek-ai/dsh-timeout-policy (a tools/execute wrapper); each tool just forwards exec.signal to the seam.

Each tool is registered independently; a product that wants only one disables the other via config ({ search: false } / { fetch: false }).

Tools

Tool Args Behavior
web_search query (string) Discovery. Returns an optional answer plus source URLs. max_results is not model-facing — the tool sets the bound (the searchMaxResults config, default 8) and passes it to the seam.
web_fetch url (string) Retrieves a specific URL. HTML bodies are rendered to markdown-ish text; text bodies pass through. A non-2xx status is reported, not an error. The tool-call timeout is deployment policy (dsh-timeout-policy), not a model argument.

Config

Key Default Meaning
search true Register web_search.
fetch true Register web_fetch.
searchMaxResults 8 Upper bound on sources returned by one web_search call (the seam truncates a longer provider list and flags it).
fetchTimeoutMs 30000 Cooperative tool-call timeout budget (ms) for web_fetch.
searchTimeoutMs 30000 Cooperative tool-call timeout budget (ms) for web_search.

fetchTimeoutMs/searchTimeoutMs declare each tool's cooperative timeout budget (attached as ToolDefinition.timeoutMs), enforced by @deepseek-ai/dsh-timeout-policy; the model-facing schema exposes no timeout argument.

- id: tool-web
  name: '@deepseek-ai/dsh-tool-web'

Stable registration

Tool registration follows product enablement, not backend availability. A tool stays visible even when its selected provider is missing, misconfigured, ambiguous, or temporarily unavailable; the seam resolves the provider at execution time and execution fails with a structured WebError (e.g. WEB_PROVIDER_UNAVAILABLE, WEB_PROVIDER_AMBIGUOUS), which ToolRegistry.execute() turns into an error tool result the model can read and hooks/UI can route on. This keeps the model schema stable without making plugin load order, credential state, or HMR timing part of the model-facing contract. To remove a web tool entirely, disable it here in config.

The tool never calls a provider's status() and never enumerates providers — its only execution path is ctx.web.search() / ctx.web.fetch(), and provider unavailability reaches it as the structured WebError codes selection throws at execution time. Provider selection stays entirely inside the seam, with one owner.

Model Experience

Context surface What the model sees Token effect
System prompt Each config-enabled tool adds one short section: search guidance says to discover current sources and follow with fetch; fetch guidance says to retrieve a specific HTTP(S) URL and cite it. A scoped tool restriction does not remove these independently registered sections. Fixed guidance cost per request for each config-enabled tool, even when a restriction hides its schema.
Tool schemas According to config and scoped restrictions, the model sees web_search(query), web_fetch(url), or both. Result-count and timeout budgets are deployment settings, not model arguments. Fixed schema cost per request; config disablement removes both schema and guidance, while a scoped restriction removes only the schema.
Tool-call history and results Search returns an optional answer and bounded source entries; fetch returns status plus decoded text or markdown-shaped HTML, or a structured error. Queries and URLs remain in call history. Data-dependent results are resent until compaction. Search sources are capped by searchMaxResults; fetch providers cap body size, and timeout policy can replace a late result with a short error.

Known Limitations and Deferred Work

  • htmlToMarkdown is a minimal regex converter, not an HTML parser — it strips script/style/noscript, keeps headings/bullets/links, and decodes about a dozen named entities; tables, images, and nested formatting are lost.
  • The model-facing surface is minimal by design, with promotions deferredmax_results stays a config bound (not a model argument), and web_fetch takes only url (no format/prompt/LLM-summarization mode); both are named later steps in the seam RFC.
  • No web-specific permission policy — both tools execute without requesting ctx.approval; a deployment that needs confirmation must add a tools/pre-execute policy, and the package does not define persistent URL/domain grants.