Files
deepseek-harness/examples/coding-agent
Tianyi Cui d6a2ab30c8 feat(types): brand bash ids + stop brand erosion; extract Branded to dsh-brand
Type-only change (brands are zero-cost casts; no runtime/wire impact). Closes
the two gaps in the "brand ids that cross package boundaries" policy and fixes
the dependency direction so a capability package never pulls in an unrelated one.

- Extract the `Branded<B>` primitive into a new standalone type-only package
  `@deepseek-ai/dsh-brand` (packages/util/brand) with no harness-package deps.
  dsh-llm keeps its owned CallId but imports Branded from dsh-brand; dsh-session,
  dsh-agent, and dsh-bash all import Branded from there. dsh-bash depends on
  dsh-brand ALONE — never on dsh-llm or dsh-session (the architectural fix: a
  generic execution backend must not couple to the LLM or session vocabulary).
- Mint BashTaskId + OwnerToken in dsh-bash and thread them through BashTask.id,
  the get/ownerOf/list/readOutput/kill seam, the bash-local generation site, and
  the dsh-tool-bash validate/access surface. OwnerToken is a DISTINCT brand from
  SessionId so the seam stays decoupled; dsh-tool-bash is the single boundary
  that casts SessionId -> OwnerToken.
- Brand at the SOURCE, not via mid-pipeline casts: agent-loop's Config types
  agents[].id as AgentId and resumeSessionId as SessionId, so the brand enters
  at the config boundary and the inner create()/resume casts disappear (only the
  genuinely-new per-run session-id string is cast).
- Stop brand erosion: propagate CallId/SessionId/AgentId to the registry/store
  Map keys and public params/exports (SessionStore, AgentRegistry + factory
  options, the ACP session-id surface + ToolPresenter CallId map, the
  persistence coordinator, invariants pendingCalls, the pi-ai tool-call maps).
- Docs: document BashTaskId/OwnerToken in bash.md (type-equiv re-pasted), point
  the Branded type-equiv at dsh-brand, fix stale param types in the session/
  agent/bash READMEs, regenerate the cordis catalog + module graph.

Implements docs/rfc/proposed/architecture/2026-06-20-branded-ids.md
2026-06-21 07:19:59 +08:00
..

coding-agent

The first REAL agent wiring: DeepSeek V4 + the bash tool suite + stdio chat

  • JSONL persistence, loaded from cordis.yml. Where echo-agent proves the skeleton with mocks, this example is a usable coding assistant.

Run it

# repo root .env (gitignored) or exported env:
#   DEEPSEEK_API_KEY=sk-…
#   DEEPSEEK_BASE_URL=https://…   # optional; defaults to the public API
pnpm run demo:coding

Type a coding task. The agent's only tools are bash (+ bash_output / bash_kill for background tasks): file reads, writes, searches, and test runs all happen through shell commands, each in a fresh bash -c (the system prompt tells the model to pass workdir instead of cd). Reasoning streams dimmed; tool calls/results render inline.

> fix the failing test in /path/to/project
[main turn 1] (reasoning…)
  [tool call] bash({"command": "node --test", "workdir": "/path/to/project"})
  [tool result] … [exit code: 1]
  …

Resuming a prior session

Each run starts a fresh session by default (its event log lands under ./.sessions/). To continue a previous conversation, set RESUME_SESSION_ID to that session's id — the main agent then rehydrates the persisted log instead of starting fresh, so the model sees the earlier turns as history:

RESUME_SESSION_ID=<prior-session-id> pnpm run demo:coding

The id is wired through cordis.yml (resumeSessionId: !!js process.env.RESUME_SESSION_ID); unset, the agent starts a new session. A missing/unreadable id is non-fatal — it logs a warning and starts no main agent.

What each plugin demonstrates

Entry Demonstrates
llm-deepseek real LlmAdapter via config (!!js process.env.… secrets); swap one line to @deepseek-ai/dsh-llm-pi-ai for the library-backed twin
bash (dsh-bash-local) + tool-bash the executor seam + tool schemas as separate plugins
agent-loop agent created from config with a coding system prompt
session-persistence (dsh-session-persistence-jsonl) durable JSONL persistence (root: ./.sessions): append-only event log per session, crash-safe atomic writes — the shared backend, no per-example file
src/stdio-chat.ts UI as a plugin; copied from echo-agent with reasoning-dimming and an exit-on-idle close handler for piped stdin. Example-local on purpose — extract a shared UI package when a third example needs it

End-to-end tests (pnpm run test:e2e, key-gated)

  • tests/full-loop.e2e.ts — the canary: real model runs echo e2e-ok through the real bash tool; asserts tool/call/tool/result session events and the final answer.
  • tests/coding-task.e2e.ts — the swebench-style smoke: a temp dir holds add.js (with a - b where a + b belongs) and a failing add.test.js; the agent must fix the bug and verify. The test re-runs node add.test.js ITSELF and inspects the files — agent claims are not trusted.
  • tests/resume.e2e.ts — durable continuity across processes: run 1 tells the real model a secret code and persists the turn to a temp JSONL root, then the whole context is disposed; run 2 is a fresh context over the same root that RESUMES the session id and asks the model to recall the code. The recall can only come from the rehydrated log.

Both self-skip without DEEPSEEK_API_KEY.