The first real LlmAdapter implementations, shipped as a deliberate pair:
same models and wire protocol, completely different internals, so the
StreamChunk protocol is verified across independent implementations.
- dsh-llm-deepseek: hand-rolled fetch + SSE parser + chunk-translation
state machine against the official chat-completions format (thinking
mode via top-level thinking/reasoning_effort; the empty-string
reasoning_content first chunk; usage attached to the finish chunk or
trailing; reasoning_content passback on tool-call turns; disjoint
cache-token accounting).
- dsh-llm-pi-ai: the same endpoint through @earendil-works/pi-ai,
mapping its event vocabulary (parsed tool arguments, in-stream error
events, folded reasoning tokens) onto the same chunks.
The agent loop now honors the in-band error path: an adapter that ends
its stream with finish {kind:error|aborted} (the only option for
adapters that can't throw mid-stream, like pi-ai) is translated into a
step error, so the turn ends error/aborted with a logged error event
instead of a normal completed assistant message. This makes the
StreamChunk error contract real for both adapters; docs/architecture.md
and the StreamChunk doc are updated accordingly.
New yarn test:e2e (vitest.e2e.config.ts, *.e2e.ts) runs key-gated
real-API matrices for both adapters across V4 Flash/Pro and all
thinking/effort levels; it self-skips without DEEPSEEK_API_KEY. Unit
suites run against local node:http mock SSE servers at 100% per-file
coverage.
Packages
Harness packages, all under the @deepseek-ai/dsh-* scope. Each package is a
Cordis service (microkernel plugin-style): it exports a default Service class
that gets registered via ctx.plugin(), declares its ctx key and events through
declaration merging, and exposes extension points through ctx.effect(),
ctx.on(), and ctx.waterfall().
Dependency graph
dsh-llm (no harness deps — pure vocabulary)
dsh-session ← dsh-llm
dsh-system-prompt ← dsh-llm
dsh-agent ← dsh-llm, dsh-session
dsh-tools ← dsh-llm, dsh-system-prompt, dsh-agent
dsh-agent-loop ← dsh-llm, dsh-session, dsh-system-prompt, dsh-tools, dsh-agent
The rule: plugins depend on interfaces, never on the concrete loop.
dsh-agent-loop is swappable — UI/hook/tool plugins keep working against the
dsh-agent vocabulary if the loop is replaced.
What goes where
| Package | Role | ctx key |
|---|---|---|
llm/ |
Abstract LLM service + content-block vocabulary + chunk assembler | ctx.llm |
session/ |
Event-sourced session log + in-memory store | ctx.sessions |
system-prompt/ |
Prompt-section + tool-schema assembly registry | ctx.systemPrompt |
tools/ |
Tool registry + tools/execute waterfall |
ctx.tools |
agent/ |
Agent interface, registry, agent/* event vocabulary |
ctx.agents |
agent-loop/ |
THE concrete plugin: LoopAgent + the loop driver |
ctx.agentLoop |
Each package has its own README.md with purpose, service API, events,
extension points, and deliberate non-goals (TODOs).
Conventions (applied across all harness packages)
- Registrations are effects: every contribution (adapter, tool, section,
agent, event listener) goes through
ctx.effect()/ctx.on(), so disposal and HMR clean up automatically. Everyregister()returns the disposer. - Declaration merging for events and ctx: services declare their events in
declare module 'cordis' { interface Events { ... } }and their ctx key ininterface Context. - Waterfall semantics:
ctx.waterfalllisteners receive(...args, next)and MUST callnext()to delegate; returning without it short-circuits (the veto mechanism). - Extensible unions:
ContentBlockMap,MessageSourceMap,FinishReasonMap,TurnTriggerMap,TurnEndReasonMap, andSessionEventMapuse the merge-extensible-map pattern so plugins can add variants via declaration merging. - ESM everywhere; imports use package names across package boundaries,
.tsextensions within a package. - Tests: vitest, colocated under
packages/<name>/tests/*.spec.ts. Every registry needs an HMR-safety test. Err on the side of more tests.