Files
deepseek-harness/examples/coding-agent
Tianyi Cui e98c1c5d42 Add examples/coding-agent and the docs cookbook
The first real agent wiring: DeepSeek V4 + the bash tool suite + stdio
chat + JSONL persistence, runnable via yarn demo:coding (reads the
gitignored repo-root .env through process.loadEnvFile).

- examples/coding-agent: cordis.yml wiring both real plugin families
  (llm-deepseek with !!js env secrets; bash-local + tool-bash), a
  bash-only coding system prompt, a max-steps-guard plugin (bounds
  runaway turns via the agent/turn-continuation waterfall — abort()
  from step-end is a no-op by then), and a stdio UI with dimmed
  reasoning and exit-on-idle for piped stdin.
- e2e (yarn test:e2e, key-gated): full-loop.e2e.ts runs a real model
  against the real bash tool; coding-task.e2e.ts is the swebench-style
  smoke — the model fixes a buggy add.js in a temp dir and the test
  re-runs node add.test.js itself rather than trusting the agent.
- docs/cookbook: adding-a-package (the verified checklist),
  adding-a-tool (execute() contract, background pattern, seams),
  adding-an-llm-adapter (protocol obligations, mock-server testing,
  e2e policy). AGENTS.md layout/commands/secrets sections updated;
  architecture.md points at both examples and the cookbook.
- vitest.e2e.config.ts: serialize test files + retry twice — parallel
  e2e files trip the shared internal key's concurrency quota.
- fix: the !js YAML tag spelling in docs/JSDoc is actually !!js
  (js-yaml resolves custom tags under tag:yaml.org,2002:js).
2026-06-13 18:30:50 +08:00
..

coding-agent

The first REAL agent wiring: DeepSeek V4 + the bash tool suite + stdio chat

  • JSONL persistence, loaded from cordis.yml. Where echo-agent proves the skeleton with mocks, this example is a usable coding assistant.

Run it

# repo root .env (gitignored) or exported env:
#   DEEPSEEK_API_KEY=sk-…
#   DEEPSEEK_BASE_URL=https://…   # optional; defaults to the public API
yarn demo:coding

Type a coding task. The agent's only tools are bash (+ bash_output / bash_kill for background tasks): file reads, writes, searches, and test runs all happen through shell commands, each in a fresh bash -c (the system prompt tells the model to pass workdir instead of cd). Reasoning streams dimmed; tool calls/results render inline.

> fix the failing test in /path/to/project
[main turn 1] (reasoning…)
  [tool call] bash({"command": "node --test", "workdir": "/path/to/project"})
  [tool result] … [exit code: 1]
  …

What each plugin demonstrates

Entry Demonstrates
llm-deepseek real LlmAdapter via config (!!js process.env.… secrets); swap one line to @deepseek-ai/dsh-llm-pi-ai for the library-backed twin
bash (dsh-bash-local) + tool-bash the executor seam + tool schemas as separate plugins
agent-loop agent created from config with a coding system prompt
src/session-jsonl.ts write-behind persistence on session/event + session/flush (copied from echo-agent)
src/stdio-chat.ts UI as a plugin; copied from echo-agent with reasoning-dimming and an exit-on-idle close handler for piped stdin. Example-local on purpose — extract a shared UI package when a third example needs it

End-to-end tests (yarn test:e2e, key-gated)

  • tests/full-loop.e2e.ts — the canary: real model runs echo e2e-ok through the real bash tool; asserts tool/call/tool/result session events and the final answer.
  • tests/coding-task.e2e.ts — the swebench-style smoke: a temp dir holds add.js (with a - b where a + b belongs) and a failing add.test.js; the agent must fix the bug and verify. The test re-runs node add.test.js ITSELF and inspects the files — agent claims are not trusted.

Both self-skip without DEEPSEEK_API_KEY.