Files
deepseek-harness/docs/rfc/012-optional-code-mode.md
T
Tianyi Cui a1eea4d36e docs: note multi-language Code Mode backends in RFC 012
Clarify that the CodeRuntime seam can host backends differing by
language/runtime, not just trust level — e.g. an AssemblyScript/WASM
backend (naturally sandboxed) and a Python backend over CPython or a
more controllable/embeddable interpreter. Note the execution contract
is language-agnostic while SDK codegen/prompt presentation is per-
language, and add these backends to the deferred follow-up list.
2026-06-15 08:37:45 +08:00

32 KiB

RFC 012: Optional Code Mode — model writes TypeScript against an SDK of all tools

Status: proposed

Problem

Today the agent loop advertises every registered tool to the model as a native JSON-schema function definition. ToolRegistry feeds its schemas into ctx.systemPrompt, the loop puts them on GenerateOptions.tools, and the adapter serializes them to the provider's function-calling wire format. The model then invokes one tool-call block per step, the loop dispatches each call through ctx.tools.execute() sequentially (parallel tool execution is an explicit open TODO in dsh-tools and docs/architecture.md), and every intermediate tool-result re-enters the model's context on the next request.

For multi-step tool work this is token-heavy and serial. The model cannot compose tools — loop over a result set, branch on an intermediate value, fan out, post-process — without a full model round-trip per call, and each of those round-trips drags the entire intermediate result back into context whether the model needs it or not.

Cloudflare's Code Mode (shipped as the @cloudflare/codemode npm package) proposes an alternative grounded in a simple observation: LLMs are better at writing code than at emitting tool calls, because they have seen millions of lines of real code and comparatively few contrived tool-calling traces. Instead of one tool call per step, the model writes a TypeScript program against a generated SDK that wraps all the tools, and that program is executed. The model curates what comes back — only what it console.logs and/or returns — instead of every intermediate result. The SDK functions are async, so the model can express fan-out (Promise.all) naturally in code; this RFC initially serializes those dispatches (§ Concurrency) until the tool contract grows concurrency-safety metadata, so the early win is composition and fewer round-trips, not parallelism.

This RFC proposes an optional Code Mode for the DeepSeek Harness, covering all tools uniformly — built-in and future MCP — with no per-tool work, implemented Cordis-style with zero core-package changes. It fully specifies the code-execution seam and the SDK-generation pipeline, but ships only a minimal node:vm reference stub for execution; the hardened, sandboxed execution substrate is deferred to a follow-up RFC (see Risks). This RFC does not change the agent loop, and it leaves native tool-calling exactly as it is — Code Mode is a plugin you load, not a replacement.

Proposal

The design follows the codebase's capability-seam pattern (ADR 0009, the bash template) as a three-package split, plus one consumer plugin. Nothing in dsh-session, dsh-agent, dsh-agent-loop, dsh-llm, dsh-tools, or dsh-system-prompt changes.

Prior art. @cloudflare/codemode validates this shape directly and several of its decisions are adopted below. Its Executor interface is deliberately tiny — execute(code, fns) → { result, error?, logs? } — with a production DynamicWorkerExecutor (isolated Workers) and a six-line NodeVMExecutor example as two implementations behind it: exactly the interface/implementation split ADR 0009 prescribes. It generates TypeScript type definitions from tools for the model's context and runs the generated JavaScript in a sandbox, capturing console output alongside the return value. It normalizes model output into an async arrow function via AST parsing (acorn) and sanitizes tool names into valid JS identifiers (my-toolmy_tool, deletedelete_). It blocks outbound network by default. The transferable lessons — minimal executor contract, host-side type derivation, capture-output-and-return-value, name sanitization, AST-normalize the code, isolate by default — are folded into the design below. What does not transfer is the substrate: Cloudflare's isolation is Workers-specific; our equivalent hardened substrate is the deferred follow-up.

Prompt-budget tradeoff (Code Mode is not unconditionally cheaper). Deriving the SDK types host-side costs no extra discovery round-trip, but the generated .d.ts is injected into the system prompt (§3a), so the type definitions themselves do consume context — and for an all-tools SDK that cost scales with every registered tool and can be comparable to, or larger than, the native JSON schemas it replaces. Code Mode's saving is on the output/result side (the model curates what comes back; intermediate results never re-enter context) and on round-trips (compose many calls in one program), not on the input-side tool description. The net win is workload-dependent: it pays off for multi-call, large-intermediate-result workflows and can cost more for a single call against a large tool surface. The .d.ts section is a prefix-stable prompt prefix, so prompt caching amortizes its per-turn cost across a session; the RFC notes that caching is what keeps the injected SDK affordable, and that a deployment with a very large tool surface should weigh the SDK size against native schemas rather than assume Code Mode is strictly cheaper.

1. Interface package packages/code-runtime/ — a new package @deepseek-ai/dsh-code-runtime owning ctx.codeRuntime, depending only on cordis. It defines an abstract CodeRuntime extends Service plus the execution vocabulary. The runtime knows nothing about ctx.tools: it is handed a set of named async functions (the resolved SDK bindings), runs the program, and captures output. The result shape mirrors Cloudflare's proven-minimal contract so an error is a field on a resolved result, not a throw the runtime is expected to make:

  • CodeRunRequest = { code: string; sdk: SdkBinding[]; signal?: AbortSignal }
  • CodeRunResult = { result: unknown; logs: string[]; error?: string }
  • a readonly safe: boolean on the CodeRuntime service — false for an unsandboxed stub, true only for a real isolating substrate; consumers gate on it (§2).
  • SdkBinding = { namespace: string; fns: Record<string, (args: unknown) => Promise<unknown>> }

Per the "explicit > implicit at seams" convention, the request spells out every field the runtime acts on; defaulting (e.g. an output cap, a timeout derived from signal) is the implementation's explicit job, not a hidden ?? default inside run(). The split into interface + implementation is justified under ADR 0009 because there is genuinely more than one planned implementation — the node:vm stub and the hardened substrate (a real isolate, or the generated program run as a sandboxed process through the existing ctx.bash seam) that is scheduled follow-up work, not speculative optionality. ADR 0009 warns against splitting preemptively when only one implementation is conceivable; here a second is not just conceivable but required before any untrusted use, so the seam earns its keep.

Backends can differ by language/runtime, not only by trust level. The two implementations above (unsafe stub vs. hardened substrate) differ along the trust axis while staying TypeScript/JS, but nothing in the CodeRuntime contract — a program string plus a set of named async SDK bindings in, and a { result, logs, error? } out — is bound to one source language. The same seam can host backends that differ along the language axis, executing a program written in something other than TypeScript. Two illustrative directions:

  • An AssemblyScript backend. AssemblyScript is a strict TypeScript subset that compiles to WebAssembly, so a program stays familiar to a TS-fluent model while the WASM boundary supplies exactly the sandboxing the hardened substrate is meant to provide — memory isolation and no ambient host authority come from the runtime rather than from after-the-fact hardening of node:vm. This is an appealing route to a safe = true backend.
  • A Python backend. Python is arguably the model's most native language — it has seen far more real Python than any tool-calling trace — which is the same "LLMs write better code than tool calls" argument that motivates Code Mode, taken one step further. A Python backend is itself a sub-seam over different Python runtimes: CPython (in-process or a sandboxed subprocess via ctx.bash) for maximum fidelity and ecosystem access, or a more controllable / embeddable interpreter — Pyodide (CPython on WASM), RustPython, or a restricted embedded interpreter — when isolation, deterministic resource limits, or a clean capability boundary matter more than running arbitrary native extensions.

These are illustrations of the seam's reach, not commitments — the MVP ships only the TypeScript path. The honest caveat is that the execution contract is language-agnostic but the presentation is not: the SDK-generation pipeline below (§3a and the jsonSchemaToTs codegen, which emits a TypeScript .d.ts) is TypeScript-specific, so a non-TS backend pairs the shared CodeRuntime contract with its own language-appropriate SDK generator and system-prompt section (a .pyi stub and Python usage instructions for the Python backend, AssemblyScript-flavored types for that one). The runtime seam is reused as-is; only the codegen/prompt half is per-language.

2. Implementation package packages/code-runtime-vm/ — a new package @deepseek-ai/dsh-code-runtime-vm, the node:vm reference stub. It type-erases the model's TypeScript via the compiler's transpileModule (or sucrase) — the types exist only to guide the model; the runtime is plain JS — then wraps the body in an async IIFE for top-level await (Cloudflare's NodeVMExecutor does literally new AsyncFunction("codemode", "return await (${code})()")), runs it in a vm.Context whose globals are a capturing console and the SDK namespace objects, awaits the IIFE, and captures the return value, the buffered logs, and any thrown error (as error: string). It applies an output cap (truncate captured logs) and a timeout tied to request.signal. These caps limit blast radius; they are not a security boundary. node:vm is not isolation: withholding require/process does not contain anything (code escapes via constructor/prototype reflection), and per AGENTS.md the harness must never hand model output the ambient environment.

The unsafe-runtime guard is enforceable, not a README warning. Because a README caveat is not a control — and AGENTS.md's "never hand model output ambient authority" is a hard rule, not advice — the design makes the danger refuse to run by construction. Two layers:

  • The runtime declares its trust level. CodeRuntime carries a readonly safe: boolean (a node:vm-class stub returns safe = false; a real isolate/sandboxed-process substrate returns safe = true). The code-runtime-vm constructor additionally requires an explicit opt-in — new VmCodeRuntime({ unsafe: true }) — and throws if that flag is absent, so merely depending on the package and wiring it cannot silently produce a live unsafe runtime; the operator must type the word unsafe.
  • The consumer refuses to expose run_code over an unsafe runtime by default. When code-mode initializes, if ctx.codeRuntime.safe === false it does not register run_code unless the plugin itself is configured with an explicit acknowledgement (e.g. code-mode config allowUnsafeRuntime: true). Absent that, it logs a typed error and registers nothing — so a real model never reaches an unsandboxed runtime by a single config slip. The refusal path is tested: with the acknowledgement unset and an unsafe runtime, run_code is absent (and the wire tool list is unchanged from native); with both opt-ins set, it registers and runs. This keeps the unsafe reference backend usable for tests and trusted local demos while making production misuse take two deliberate, greppable flags rather than one mistake.

code-runtime-vm is therefore documented as reference / test-only / unsafe-for-untrusted-input, acceptable in the MVP only because the code runs at harness trust and both opt-in flags must be set. Signal handling is best-effort: it aborts in-flight sub-dispatches but cannot reliably interrupt a hot synchronous loop (while(true){}) in node:vm — another reason the hardened substrate is deferred, not optional-forever.

3. Consumer plugin packages/code-mode/ — a new package @deepseek-ai/dsh-code-mode, the plugin that wires everything together. It declares inject = ['tools', 'systemPrompt', 'codeRuntime'] — Cordis throws on access to a service that is not injected, and keeps the plugin inactive until all three exist (the same pattern as tool-bash's inject = ['tools', 'bash']), which also gives correct load-ordering relative to code-runtime/code-runtime-vm. The plugin contributes four things, all through existing seams:

3a. Tool presentation — a lazy system-prompt section (the injection seam already exists). dsh-system-prompt already provides the Cordis-idiomatic way for any plugin to inject prompt snippets: ctx.systemPrompt.section({ name, order, text }), fiber-scoped and auto-disposed via ctx.effect(), where text may be a lazy () => string re-evaluated at each assembly. No new mechanism is needed or invented. Code Mode registers a lazy section (high order so it lands last) whose thunk reads ctx.tools.schemas() at assembly time and regenerates the SDK .d.ts plus usage instructions from the currently-registered tool set. Because the thunk reads the live registry, coverage of every tool — built-in, MCP, future — is automatic.

3b. Wire tool-list enforcement — an agent/request listener (the authoritative seam). The goal "exactly one tool reaches the wire" must be enforced where the wire request is finalized. The loop calls ctx.systemPrompt.assemble() first, then builds GenerateOptions (seeding tools from assembly.tools), then runs the agent/request waterfall, then calls ctx.llm.stream(). A system-prompt/assemble listener can only influence the seed; agent/request is the last seam before the model call, so it is authoritative. The plugin registers an agent/request listener that does const final = await next(); return { ...final, tools: [runCodeSchema] } — overriding the value returned by next(), not the inbound argument, so it dominates the cooperative request listeners it wraps. It registers with prepend: true to sit at the outer edge of the waterfall chain. One honest caveat, stated in the RFC body: ctx.llm.stream() itself runs a further llm/stream waterfall before the adapter, so the guarantee is "authoritative within the agent request pipeline," not an absolute wire invariant; if a hard invariant is ever required, a defensive llm/stream assertion with a spy adapter covers it in tests.

3c. The single tool — run_code. Registered normally in ctx.tools with one parameter { code: string (required) }. Because it is an ordinary tool, the unchanged loop dispatches it through the normal path — this is the crux of "zero loop changes." Its execute(args, exec):

  1. Builds the SDK bindings. For each real tool, an async invoke(callArgs) that checks exec.signal?.aborted (throwing if set) before and after calling ctx.tools.execute({ callId: <deterministic sub-id>, name, arguments: callArgs, agent: exec.agent, signal: exec.signal }), then maps the resulting ContentBlock[] to a simplified { output, isError } (text blocks for the MVP), and emits an observability event. The explicit abort check matters because ctx.tools.execute() catches thrown tool errors and converts them to isError results — without the check, an aborted sub-call would look like ordinary error data and the program would keep running instead of stopping. Sub-dispatch still flows through the tools/execute waterfall, so permission/sandbox/hook plugins apply to code-mode calls exactly as to native ones.
  2. Calls ctx.codeRuntime.run({ code: args.code, sdk: bindings, signal: exec.signal }).
  3. Surfaces the outcome. A successful run returns [{ type: 'text', text: <console logs + return value> }]. A runtime-error result cannot be reported by returning content, because a normal ToolDefinition.execute() returns only Promise<ContentBlock[]> and ToolRegistry.execute() hardcodes isError: false on any successful return — isError: true arises only from the registry's catch path. So on an error result the tool throws a CodeRunError extends HarnessError (HarnessError is exported from dsh-llm; the registry catch turns any throw into isError: true with the message as text, and a HarnessError additionally carries structured { name, code }). An alternative — registering run_code handling as a tools/execute listener that returns a full ToolExecutionResult and can set isError directly — is noted; the throw is simpler and preferred.

3d. Result discipline — what the model receives. The model gets back only the captured console output and/or the program's return value (the model chooses which to surface). Intermediate sub-call results are never returned to the model. This is the core context-saving benefit: the agent curates its own output, exactly as a script's stdout curates a pipeline's intermediate state.

Sub-call CallIds. Real tool calls dispatched from inside run_code need ids, but CallId is normally provider-issued (a branded string for correlating a call with its result — only brand-wrapped via CallId(), with no generator and no documented session-global-uniqueness guarantee). The plugin mints deterministic sub-ids scoped to the parent: `${exec.callId}:code:${n}` with a per-run counter n. These are unique within one run_code run (assuming the parent callId is unique, which the provider guarantees per turn); the code/dispatch event additionally carries the session log's seq so the UI and persistence can order and disambiguate globally without relying on the id alone. ToolExecution.agent is optional; the normal loop always supplies it (and with it exec.agent.session, the log code/dispatch appends to). A run_code execution arriving without exec.agent still runs (sub-calls propagate agent: undefined, exactly as the loop's own contract allows) but skips session-log observability — with no session to append to, those direct runs are simply not logged.

Observability without context cost. Each sub-dispatch emits a session event declared by the dsh-code-mode plugin itself via SessionEventMap declaration merging (the map is merge-extensible precisely so plugins can add events without touching dsh-session). Shape: code/dispatch with { parentCallId, subCallId, name, arguments (or redacted), isError, summary }, ordered by the session log's own seq. deriveMessages() does not translate it into a model message — an unknown event type falls through its default, per the merge-extensible-union convention — so the UI and persistence (RFC 009) can render every sub-call while the model's context only ever receives the single run_code tool-result. Because the event lives in the plugin, this adds no core change.

SDK codegen. A pure jsonSchemaToTs(schema) in code-mode maps the JSON-schema subset the defineTool DSL produces (object/string/number/boolean/array, properties, required[], enum → string-literal union, nested objects, array items) to a TS type literal. It is total: any unsupported construct ($ref, oneOf/anyOf, integer, null, additionalProperties, or any raw MCP shape it does not recognize) degrades to unknown without throwing — it never crashes codegen. Typing is best-effort, not a guarantee, because MCP tools accept arbitrary JSON Schema and ToolSchema.parameters is typed only as Record<string, unknown>. Because ToolSchema.name is an arbitrary string (not necessarily a valid TS identifier), the SDK is generated as a namespace with quoted access (e.g. tools["some-mcp-tool"](args)) plus safe camelCase aliases where the name is a clean identifier; alias collisions and TS reserved words fall back to quoted-only access (no duplicate alias emitted). This mirrors Cloudflare's sanitizeToolName. run_code itself is filtered out of the SDK. The MVP surfaces text content only; image and other block types in sub-results are deferred (noted as a limitation).

Concurrency — serialized by default (the binding must enforce it). The SDK functions are async, so a model writing await Promise.all([tools.a(...), tools.b(...)]) would start both immediately, and each would call ctx.tools.execute right away — i.e. the binding shape makes concurrent dispatch the default, not an opt-in. Because the tool contract carries no concurrency-safety metadata today (parallel tool execution and a concurrency-safety hint are an open TODO in both dsh-tools and docs/architecture.md: "phase 1 executes tool calls sequentially"), concurrent dispatch through a not-yet-hardened tool may race. So a prose "may serialize" is not sufficient. Decision: the MVP SDK bindings enforce serialization — each run_code invocation owns a per-run dispatch queue, and every invoke() chains onto it (tail = tail.then(() => ctx.tools.execute(...))), so even Promise.all over SDK calls executes them one at a time in submission order. This is a hard acceptance criterion, with a test that issues Promise.all([...]) from a program and asserts the underlying ctx.tools.execute calls did not overlap (e.g. a probe tool records enter/exit and the test asserts no interleaving). The .d.ts may describe the model-visible functions as async (they are), but the implementation guarantees serial execution. Lifting serialization is deferred: only once a tool can declare itself read-only / concurrency-safe does the binding allow those specific tools to overlap. The same per-run queue is where the before/after abort checks (§3c) live, so an aborted run drains no further queued dispatches.

Tool visibility tiers (design intentionally skipped). A natural extension is to mark each tool with a visibility tier: some tools "direct-call eligible" (still offered as native wire tools alongside run_code), some "code-mode only" (reachable solely from within a run_code program, never on the wire), and the default "both." This would let a deployment keep a few high-frequency or approval-gated tools as direct calls while routing the long tail through Code Mode, or hide composition-only primitives from the native surface entirely. This RFC notes the possibility but intentionally skips the detailed design — the per-tool metadata, how it interacts with the agent/request enforcement in 3b, and the presentation split in 3a are left to a follow-up. The MVP is the simple two-state model: Code Mode on (everything via run_code) or off (everything native).

Optionality / toggle. Loading the code-mode plugin enables Code Mode for that context; not loading it leaves today's native tool-calling untouched. The two are mutually exclusive within one ctx, because Code Mode rewrites the wire tool list down to [run_code]. Per-agent selection via ctx forks, and the visibility tiers above, are future work; the MVP toggle is plugin presence.

Alternatives

Result elision / summarization over native tool-calling (the narrower route). The Problem has two halves — context bloat (every intermediate tool-result re-enters context) and serial composition (one tool call per round-trip). The context-bloat half can be addressed without any code-execution runtime: keep provider tool-calling exactly as it is, and add a plugin on the agent/request waterfall (or a compaction pass akin to RFC 009's session work) that elides or summarizes older tool-result blocks before they re-enter the model's context — drop them past a window, replace large payloads with a digest, or keep only the blocks the model still references. This is strictly less invasive than Code Mode: no new runtime seam, no model-written programs, no new safety surface. It is the right tool if context growth is the only pain.

It is insufficient for the composition / round-trip half, which is the decisive reason this RFC does not stop there. Elision still pays one model round-trip per tool call: a loop over N items is N turns, a branch on an intermediate value is a turn to fetch then a turn to act, and post-processing (filter, join, reduce) either happens in the model's head over full payloads or not at all. Code Mode collapses all of that into one program — the loop, the branch, the join run in the runtime, and only the curated result returns. Elision also cannot express fan-out or data-dependent control flow; it only shrinks what comes back. So the two are complementary, not competing: elision could even layer under Code Mode for the residual native-tool paths. The RFC chooses Code Mode because the round-trip/composition cost is the larger structural limit, and accepts the new code-execution surface as the price — which is exactly why the execution substrate is gated behind the enforceable safety guard (§2) and the hardened backend is a hard prerequisite for untrusted use.

Why not change the loop to dispatch native tool calls in parallel instead? That is the other obvious answer to the round-trip cost, and it remains valid future work (it is the open dsh-tools/architecture.md TODO). But it is a core-loop change requiring the same concurrency-safety metadata Code Mode defers, and it still does not give the model composition (branch/loop/post-process between calls) — only parallelism of independent calls the model already decided to make in one step. Code Mode delivers composition with zero core change; parallel native dispatch and Code Mode can coexist later.

Plan

  1. Scaffold the interface package packages/code-runtime/ per the cookbook: abstract CodeRuntime extends Service (super(ctx, 'codeRuntime')) with a readonly safe: boolean, the declare module 'cordis' ctx key, the CodeRunRequest/CodeRunResult/SdkBinding vocabulary, method contracts documented in JSDoc (what run captures, abort semantics, that an error is a result field not a throw, what safe means). HMR-safety test (dispose the contributing fiber, assert ctx.codeRuntime is gone).
  2. Scaffold the implementation package packages/code-runtime-vm/: the node:vm stub — safe = false, a constructor that throws unless given { unsafe: true }, transpile/type-erase, async-IIFE wrap, capturing console, SDK globals, return-value/logs/error capture, output cap, signal-tied timeout. Tests for output capture, return value, error-as-field, abort, the constructor refusal without unsafe, and a README documenting the "not a sandbox, trusted-only" caveat prominently.
  3. Scaffold the consumer plugin packages/code-mode/: jsonSchemaToTs codegen with namespace/quoted-access + alias handling (unit tests, including non-identifier MCP names and unsupported-shape → unknown); the registered lazy ctx.systemPrompt.section() carrying the SDK .d.ts; the agent/request listener (prepend: true) collapsing request.tools to [run_code] after await next(); the unsafe-runtime gate (refuse to register run_code when ctx.codeRuntime.safe === false unless allowUnsafeRuntime is set); the run_code tool with the dispatch bridge (per-run serialization queue, deterministic sub-call ids, before/after abort checks, CodeRunError on error results); and the code/dispatch event declared here via SessionEventMap merge. Declare inject = ['tools', 'systemPrompt', 'codeRuntime'].
  4. Tests: HMR-safety (dispose removes the tool, the section, and the listener); a waterfall test that the wire tool list is exactly [run_code] (spy adapter, asserting via agent/request and optionally llm/stream); an integration test that a program calling two tools returns only its printed/returned output (verify the world, not the self-report); a serialization test that Promise.all([...]) over SDK calls does not overlap the underlying ctx.tools.execute invocations (a probe tool records enter/exit; assert no interleaving); deriveMessages() ignores code/dispatch; abort mid-program stops further dispatches; CodeRunError surfaces as isError: true; and the unsafe-runtime refusal test (§3, the VM-guard): with the unsafe flag unset, a non-mock agent's run_code is refused; with it set, the program runs.
  5. Wire an example: examples/coding-agent-code-mode (or a config flag on the existing example) loading the trio. Running it against the node:vm stub requires both opt-ins (VmCodeRuntime({ unsafe: true }) and code-mode's allowUnsafeRuntime); the example sets them explicitly and comments why, or uses a mock model — a real model never reaches the unsandboxed stub without those deliberate flags. Add a yarn demo:* entry.
  6. Docs: update docs/architecture.md (a ctx.codeRuntime row in the service map, a Code Mode note under the tool pipeline / capability seams sections); add a cookbook note on writing a CodeRuntime backend; and file the follow-up RFC for the hardened execution substrate (the isolate/sandboxed-process design, the additional-language backends sketched in §1 — AssemblyScript/WASM, Python — with their per-language SDK generators, plus the tool-visibility-tier design skipped here). Append the | 012 | … | proposed | row to the RFC index.

Risks

node:vm is not a sandbox. This is the single biggest caveat. Withholding require/process is not a boundary; the MVP runs at harness trust only; the hardened substrate is a hard prerequisite before any untrusted use and is the explicit subject of a follow-up RFC. The guard is enforceable, not just documented: the runtime exposes safe: boolean, the VM stub throws unless constructed with { unsafe: true }, and code-mode refuses to register run_code over an unsafe runtime unless separately acknowledged (allowUnsafeRuntime) — production misuse requires two deliberate, greppable flags, and the refusal path is tested.

Wrong seam would leak tools. If the wire tool list were enforced only in system-prompt/assemble, a later agent/request listener could re-add tools. Mitigation: enforce request.tools = [run_code] in the agent/request waterfall (the authoritative seam, run last before llm.stream()) with prepend: true, and assert exactly one wire tool in tests. The residual llm/stream caveat is documented, not hidden.

Concurrency before the contract supports it. The binding shape makes concurrent dispatch the default, and the tool contract has no concurrency-safety metadata yet, so unguarded Promise.all over SDK calls could race a not-yet-hardened tool. Mitigation: the MVP bindings enforce a per-run serialization queue (every invoke chains onto the previous), with a test asserting Promise.all from a program does not overlap the underlying ctx.tools.execute calls. Per-tool parallelism is unlocked only once a tool can declare itself concurrency-safe.

Two presentation modes to keep coherent. A tool added later must work in both native and Code Mode. Mitigation: both the codegen thunk and the agent/request listener read ctx.tools.schemas(), so coverage is automatic; a test asserts every registered schema produces valid .d.ts, including non-identifier MCP names via quoted access.

Type-erased runtime is not type-checked. The model can write code that type-checks against the advisory .d.ts but throws at runtime, and MCP-schema typing is best-effort. Mitigation: errors are captured as CodeRunResult.error and surfaced so the model can self-correct; the .d.ts is explicitly advisory.

Lost observability of sub-calls. Routing everything through one run_code result hides the individual calls from the model — and could hide them from operators too. Mitigation: the plugin-declared code/dispatch event keeps every sub-call in the session log and UI without polluting model context.

Abort granularity. node:vm cannot reliably interrupt hot synchronous code, and ctx.tools.execute() converts thrown aborts into isError data. Mitigation: the SDK bindings check signal.aborted and throw before/after each dispatch so an aborted sub-call stops the program; the vm stub wraps the run in a signal-tied timeout; the hardened substrate addresses the hot-loop case.

Unsafe example wiring. A demo running a real model through the node:vm stub would hand model output ambient authority. Mitigation: examples are mock-model or explicitly marked unsafe; code-runtime-vm is labeled reference/test-only.

Non-text sub-results dropped in the MVP. Image and other block types from sub-calls are not surfaced into the program yet. Mitigation: noted as a known limitation; block-type handling deferred.