The first real LlmAdapter implementations, shipped as a deliberate pair:
same models and wire protocol, completely different internals, so the
StreamChunk protocol is verified across independent implementations.
- dsh-llm-deepseek: hand-rolled fetch + SSE parser + chunk-translation
state machine against the official chat-completions format (thinking
mode via top-level thinking/reasoning_effort; the empty-string
reasoning_content first chunk; usage attached to the finish chunk or
trailing; reasoning_content passback on tool-call turns; disjoint
cache-token accounting).
- dsh-llm-pi-ai: the same endpoint through @earendil-works/pi-ai,
mapping its event vocabulary (parsed tool arguments, in-stream error
events, folded reasoning tokens) onto the same chunks.
The agent loop now honors the in-band error path: an adapter that ends
its stream with finish {kind:error|aborted} (the only option for
adapters that can't throw mid-stream, like pi-ai) is translated into a
step error, so the turn ends error/aborted with a logged error event
instead of a normal completed assistant message. This makes the
StreamChunk error contract real for both adapters; docs/architecture.md
and the StreamChunk doc are updated accordingly.
New yarn test:e2e (vitest.e2e.config.ts, *.e2e.ts) runs key-gated
real-API matrices for both adapters across V4 Flash/Pro and all
thinking/effort levels; it self-skips without DEEPSEEK_API_KEY. Unit
suites run against local node:http mock SSE servers at 100% per-file
coverage.
@deepseek-ai/dsh-llm-deepseek
DeepSeek chat-completions adapter for the harness LLM seam: hand-rolled
fetch + SSE translation from the official wire format (source of truth:
the API docs — guides/thinking_mode, guides/tool_calls,
api/create-chat-completion) into the StreamChunk protocol.
A second, independent implementation of the same seam exists in
@deepseek-ai/dsh-llm-pi-ai (library-backed). Same Config shape — pick one
per context (registering both for the same model names throws by design).
Config
- id: llm-deepseek
name: '@deepseek-ai/dsh-llm-deepseek'
config:
apiKey: !!js process.env.DEEPSEEK_API_KEY # or rely on the env fallback
baseURL: !!js process.env.DEEPSEEK_BASE_URL # default: https://api.deepseek.com
models: [deepseek-v4-flash, deepseek-v4-pro] # one adapter, registered for each name
thinking: enabled # optional; provider default is enabled
reasoningEffort: high # optional; high | max — omitted ⇒ not sent
models lists every model name this one adapter instance serves: the adapter
registers itself for each (the harness model name IS the wire model string),
so a generate/stream call routes to it whenever options.model is any of
them. Registering a second adapter for a name already taken throws
LlmError('DUPLICATE_ADAPTER') (the LLM service enforces one adapter per
model, all-or-nothing).
reasoningEffort is omitted by default — when unset, the reasoning_effort
wire field is not sent and the server applies its own default for the model.
The only accepted values are high and max (DeepSeek's official effort
levels). It is meaningful only with thinking enabled (the provider default).
thinking/reasoningEffort are adapter-level request defaults serialized as
the official top-level thinking: {type} / reasoning_effort wire fields.
They live in adapter config (not GenerateOptions) to keep the core
vocabulary provider-neutral.
Wire-format notes (verified live + against the official docs)
- Streaming only (
stream_options.include_usagealways on).usagemay arrive attached to the finish chunk or as a trailing usage-only chunk — the translator defers both to[DONE], sousagealways precedesfinishand nothing followsfinish. - The first thinking-mode chunk carries
reasoning_content: ""— handled (no spurious reasoning block). - Reasoning passback rule: on assistant turns that carried tool calls,
reasoning_contentis serialized back in history (required by the API in thinking mode); on tool-call-free turns it is dropped (ignored anyway — saves tokens). stricton tool schemas passes through (officially Beta; the public API wants the/betabase URL for it, the internal endpoint accepts it directly).- Cache accounting:
cacheReadTokens←prompt_cache_hit_tokens/prompt_tokens_details.cached_tokens; DeepSeek reports no cache-write metric.
Limitations (MVP, documented deliberately)
prefillthrowsLlmError('UNSUPPORTED')— DeepSeek's chat-prefix completion is a Beta feature on the/betabase URL; future work.imageblocks are skipped (no vision support on these models).tool_choiceis not mapped (not part of the core vocabulary).
Errors
Non-2xx responses throw LlmError with stable codes: AUTH (401/403),
RATE_LIMIT (429), INVALID_REQUEST (400), SERVER (5xx), HTTP_<status>
otherwise. Protocol violations throw STREAM_CLOSED (no [DONE]) or
MALFORMED_RESPONSE (bad JSON payload). Unknown wire finish_reasons
(e.g. content_filter, insufficient_system_resource) become
finish {kind: 'error', code: <REASON>} chunks.
Testing
Unit suites run against a local node:http mock SSE server (no network).
Real-API coverage lives in tests/adapter.e2e.ts (yarn test:e2e,
key-gated): V4 Flash + V4 Pro across thinking enabled/disabled and both
official effort levels, including the thinking+tools round trip with
reasoning passback.