Codex diff review, round 1, two (A) findings: - The agent/request fallback resolved the RAW seed object — on later steps the session's cached header fold — so a delegating listener (await next(), mutate, return) could rewrite the fold in place and the change would compare as already-baseline: no delta logged, the persisted log unable to reconstruct the request (the dev invariant would fire on the divergence, but the log would still lie). One structuredClone'd, deep-frozen seed now serves both the listener chain and the fallback — in-place shaping after delegation throws — and Session.requestHeader() freezes its fold on update, so the leak class is unrepresentable from either side. Pinned by a loop-level delegating-mutator test. - Doc sweep for the old contract: agent README's event row (mutate GenerateOptions / tool filtering → frozen config seed, replacement out, logged header), compact-basic's module JSDoc (summarize routed through agent/request → direct one-shot at llm/stream), and architecture.md's event-domain line (request mutation → call-config shaping).
@deepseek-ai/dsh-compact-basic
The basic compaction backend: a BasicCompactService implementing the @deepseek-ai/dsh-compact seam with a chars-per-token heuristic (the charsPerToken config, default 4), token-budget retention, and summarization routed through the agent request pipeline.
This is the implementation tier of the compaction capability — see the interface package for the seam and the capability-seam RFC for the design.
What it owns
The abstract contract states only WHAT compaction does; this backend owns every HOW decision:
- Token estimation —
estimateContentTokens(): chars divided by thecharsPerTokenconfig (default 4) with per-block structural overhead (text/reasoning=ceil(len/charsPerToken) + 4,tool-callfrom name + arguments,tool-resultrecursive, unknown blocks via JSON length). - Retention policy —
compactIfNeeded()walks the surface nodes tail→head summing per-node token estimates, and retains the smallest tail-run of WHOLE units (a closed step, or a single no-step node such as a pre-stepuser/messageor inter-stepsteering/message) whose total reachesretainTokens; everything older is compacted. Retention is turn-agnostic — turn boundaries play no role, so a single runaway turn that alone exceeds the window compacts its OWN early closed steps rather than being retained verbatim (the failure mode that motivated dropping turn-protection: a tool-heavy turn must stay compactable or the harness dies exactly when compaction is needed). The only structural guard is tool-pairing balance: the compacted region's edges are balanced cuts on the surface (no unanswered tool-call crosses either edge), so it never splits a step'sassistant/messagetool-calls from theirtool/results. When the only compactable content left is an un-splittable open tail step, it declines (returnsnull) and retries once an older step closes. Single-unit overflow is out of scope, by design: if one retained unit (a single closed step, or a large pasteduser/message) ALONE exceeds the budget, compaction cannot help and the call may go out over-budget — bounding an individual unit's size is a separate concern.compactRegion()enforces tool-pairing balance strictly, throwing on a boundary that would split a step.dsh-sessionexportsisToolPairingBalancedfor the check. - Dynamic convergence — no static summary-length config pretends to bound what the model will write. If framing/estimator/system overhead leaves the compacted surface above threshold,
compactIfNeeded()re-compacts the head checkpoint up tocompactionRetriesextra times; if it still cannot get below threshold, it throws. A summary whose estimated stored size is not smaller than the shadowed content fails closed before it mutates the surface. - Summarization —
summarize(): aGenerateOptionsrequest assembled viaBlockAssemblerwith a fixed system prompt that asks for a structured checkpoint (Primary Request and Intent · Key Technical Concepts · Files and Code · Errors and Fixes · Pending Tasks · Current Work · Next Step · Critical Context), every section mandatory, exact paths/commands/identifiers preserved. The request is a direct one-shotctx.llm.stream()call — NOT a loop step, so it does not runagent/request(that seam shapes the loop's conversation requests); the model comes fromsummarizationModelfalling back to the agent's own, and per-call routing happens atllm/streamlike any other direct call.maxTokensis the provider-side generation cap; only text blocks from the model's reply are kept before the checkpoint is stored (reasoning is dropped so private chain-of-thought never leaks into the durable summary, and a straytool-callis dropped so the synthesizeduser/messagesummary cannot land an orphaned call with no matchingtool-result). The compacted region is flattened to a plain-text transcript first: text and reasoning contribute their text, and every non-text block (tool-call, tool-result, plugin-added types) contributes a type-tagged placeholder ([tool-call: name(args)],[tool-result: …], …) so the summarizer is told what existed rather than silently dropping it. - Checkpoint framing — the raw summary is not landed directly.
compactRegion()wraps it in a checkpoint preamble (so a resuming model reads it as a checkpoint, not a fresh user request, and builds on the captured context rather than restating it) plus<compacted-summary>…</compacted-summary>tags. Because region compaction can be invoked manually, a surface may hold several checkpoints, so the framing does not claim everything after it is recent or verbatim. The tags make a prior checkpoint detectable in the transcript on the next compaction cycle: the summarization prompt then instructs the model to merge it in place (preserve still-true facts, drop stale ones) rather than re-summarize it verbatim — a cheap incremental merge that needs no extra log/event machinery. The unframed summary stays on thecompact/summaryprovenance event. - Surface mutation —
compactRegion()appends thecompact/start→compact/summary→compact/endlog records and the singleuser/messagereplace node carrying the framed summary (see the interface README). - Auto-compaction — an
agent/pre-steplistener delegates tocompactIfNeeded()before every step (not just a turn's first — a tool-heavy turn grows the surface mid-turn, so a runaway turn still compacts, and per-step firing is the only moment to rescue it before overflow).agent/pre-stepis a serial (awaited, in-order) surface-mutation checkpoint that fires afterturn/startand BEFORE the step opens (step/start) and its request history is derived, so compaction mutates the surface — with its log-onlycompact/*records landing cleanly outside any step — and the loop derives once from the result: no double-derive, and the listener cannot see (or need to rewrite) an already-assembledmessagesarray. The listener owns no threshold logic of its own (the single token-pressure check lives incompactIfNeeded()); because Cordisserialbails early on non-void return values, the listener returnsvoidand does not use the dispatcher's bail channel as a veto surface. - Failure handling — the
compact/start … compact/endbracket is a log-recorded lock: it makes a crash mid-summarization a detectable orphan (acompact/startwith nocompact/end), records provenance, and prevents a concurrent compaction. Two failure paths: a crash (the loop dies mid-summarization) leaves a danglingcompact/startthat is inert —compact/*events are log-only, the surface replacement never landed, so the full history derives fine and generic turn-repair closes the turn; a recoverable failure (summarization throws but the loop survives) appendscompact/endwith itserrorfield set, leaving the surface untouched so the call proceeds with full history. Core session repair stays compaction-agnostic by design — it never learns aboutcompact/*.
estimateContentTokens() and summarize() are overridable hooks: a tokenizer-based or template-based backend can subclass BasicCompactService and override just those, reusing the retention walk and surface plumbing. summarize() returns the summary blocks together with the call envelope it actually used ({ summary, model, maxTokens? }) — the caller logs that envelope on the compact/summary provenance event, so an overriding backend reports its own envelope honestly.
Config (BasicCompactConfig)
Every knob is required except auto — there is no concrete data yet to justify default thresholds/budgets, so a consumer states each value explicitly rather than inherit a guessed default. auto alone defaults to true.
| Key | Required | Meaning |
|---|---|---|
contextWindow |
yes | Context window size in tokens. |
thresholdRatio |
yes | Compact when estimated usage exceeds this fraction of the window. |
retainTokens |
yes | Tokens of recent context to keep intact. |
summarizationModel |
yes | Model for summarization ('' → use the agent's model). |
maxTokens |
yes | Provider generation cap for the summarization call; may include reasoning tokens. |
compactionRetries |
yes | Extra compaction attempts after the first if the compacted surface remains over threshold. |
auto |
no (default true) |
Register the agent/pre-step auto-compaction listener. Set false for manual-only. |
charsPerToken |
no (default 4) |
Token-estimator text density (estimated tokens = chars / charsPerToken; may be fractional). The default suits English text; CJK-heavy deployments should set ~1-2 or the estimate undershoots several-fold and compaction fires too late. |
Usage
import type { Context } from 'cordis'
import { BasicCompactService } from '@deepseek-ai/dsh-compact-basic'
export const name = 'compact-basic'
export const inject = ['llm']
export function apply(ctx: Context): void {
ctx.plugin(BasicCompactService, {
contextWindow: 128000,
thresholdRatio: 0.8,
retainTokens: 20480,
summarizationModel: '',
maxTokens: 8192,
compactionRetries: 1,
})
}
Loading the plugin registers ctx.compact. With auto: true (the default) it compacts automatically under token pressure; a consumer (a future /compact tool) can also call ctx.compact.compactIfNeeded(...) or ctx.compact.compactRegion(...) directly.