Files
deepseek-harness/examples/acp-agent
Tianyi Cui 9a5a3835c8 ci+fix: run snapshot tests in CI and load .env only when recording
Holistic-review fixes for integration gaps the per-commit reviews missed:

- CI now runs `pnpm run test:snapshot` (a step after the coverage gate). It was
  wired into pre-push but not .github/workflows/ci.yml, so the RFC/AGENTS claim
  that snapshot replay runs in the default PR gate was only half-true — CI is
  the real gate.
- vitest.snapshot.config.ts loads the repo .env ONLY when DSH_SNAPSHOT=record.
  Loading it unconditionally contradicted the replay safety story (replay must
  never reach the network), and runScenario forwards process.env to the child.
  Non-ENOENT load errors now surface instead of being swallowed.
- start.ts: the graceful-shutdown comment said "RECORD runs" but the path
  applies to both snapshot modes (replay also closes stdin → dispose → exit).
- docs/development.md: list the new pre-push snapshot job and the CI snapshot
  gate.
2026-06-19 04:28:09 +08:00
..

acp-agent example

The DeepSeek Harness coding agent exposed as an Agent Client Protocol (ACP) server over JSON-RPC stdio — drive it from Zed or any other ACP client.

pnpm run demo:acp          # needs DEEPSEEK_API_KEY (repo-root .env or env)

This boots @deepseek-ai/dsh-acp over the shared provider/tool core (../base.yml), with agent-loop configured with no pre-created agents (ACP session/new creates them on demand) and JSONL session persistence (so session/load works).

stdout is the protocol

This example loads no stdout loggerstdout carries the JSON-RPC frames, and any other write corrupts them. Do not add @cordisjs/plugin-logger-console or a stdio UI here. Use a stderr exporter if you need logs.

Zed configuration

Add to your Zed settings.json under agent_servers:

{
  "agent_servers": {
    "DeepSeek Harness": {
      "command": "pnpm",
      "args": ["run", "demo:acp"],
      "env": { "DEEPSEEK_API_KEY": "sk-…" }
    }
  }
}

The editor sets each session's cwd to the project it opens; the agent's bash tools run there (see the per-session cwd note in packages/acp), so the server does not need to be launched in the workspace.

Snapshot tests (record-once / replay-deterministic)

This example is the home of the harness's snapshot tests — they boot this server as a real subprocess, drive it with a deterministic input script, and diff its normalized output against committed golden files. The model is made deterministic by src/llm-replay.ts, a function/namespace plugin that installs an llm/stream waterfall listener and short-circuits it, serving model streams reconstructed from a recorded session JSONL fixture (<scenario>/session.jsonl) — so replay needs no API key. The fixture IS the persisted session log: its assistant/chunk events carry every StreamChunk, so grouping them by (turn, step) reconstructs each stream() call (one model call per loop step). Recording is therefore "run the real agent once and harvest the .jsonl". The two failure modes not expressible as logged chunks — a pure throw before any chunk, and cancel/hang — use an optional <scenario>/replay.override.json sidecar (a ReplayEntry[] that replaces the derived script). See docs/rfc/implemented/2026-06-19-acp-snapshot-tests.md for the full design.

MVP limitations

The bridge supports N concurrent sessions per connection, each in its own workspace cwd (RFC 011). Remaining limits: text-only prompts, additionalDirectories rejected (a session operates in its single cwd), and the tool-permission gate is deferred (TODO(rfc010-permission-gate) — tools run with the executor's full authority). See packages/acp/README.md for the full contract.