Two review responses that belong together — the same review argued the
engine was defending the wrong threat while a benign-input bug wedged
the product.
1) Drop hostile-value containment; state the trust premise.
Scripts are model-written — the same trust level as the model's bash
access — yet successive pre-push review rounds had ratcheted in defenses
that only matter against an adversarial author: trap-free proxy
rejection, accessor-never-invoked descriptor walks, realm-side
pre-rendering of thrown values, realm-built promises/arrays/error clones
with structural fatal recognition. That same author keeps a documented,
accepted, unkillable event-loop spin, so containing its error VALUES is
cost without a threat model — and the planned hardened engine
(worker/isolated-vm) gets value isolation by serialization and deletes
all of this machinery anyway.
What stays, because benign scripts hit it constantly: result never
rejects; dropped hook promises cannot become unhandled rejections; the
value boundary rejects LOUD everything JSON cannot carry (now a plain
recursive walk — getters are read ordinarily and their result is what
crosses; a throwing read fails loud); a "__proto__" key still copies as
a data property; the fatal-vs-null combinator discipline (now host
instanceof — unforgeable from the realm and simpler than clone-shape
recognition). What changes for scripts (documented in the engine
README): hooks hand back host values and host errors — in-script
`instanceof Error` on a hook failure is false (branch on e.name/e.code)
— and args are host-cloned once so a script cannot mutate the caller's
object. realm.ts drops 289 → 173 lines; the hostile-value test tables go
with it. The premise now leads the engine module doc, the README, and
the RFC's engine section, with the removed machinery recorded under
What was rejected.
2) result settles within the dispose grace of a cancellation.
Review finding (verified through the real registry + tool + engine): a
script parked on a promise no hook owns — `await new Promise(() => {})`,
`await Promise.race([])`, a returned never-settling thenable — could not
be settled by cancel(): hooks reject and children abort, but nothing
touches a promise the engine does not own, so `result` stayed pending
FOREVER (the previous cut even pinned that as intended). The tool awaits
run.result BEFORE its disposing finally, the registry awaits the tool,
the loop awaits the registry — one such script wedged the whole agent
turn past any abort, unrecoverable in-process; the mock engine in the
tool's abort test settles result on cancel, which is exactly the
behavior the real engine lacked, so no existing test could see it.
The seam contract now says it out loud: once a run is cancelled, result
SETTLES within the implementation's bounded grace even if the script
never does. The vm engine arms an abandon channel in cancel(); drive()
races the script against it, force-settling 'cancelled' at the grace
(the abandoned settlement stays contained; a post-slice synchronous spin
remains the documented limitation). dispose()'s outer race now exists
for child quiescence only, and `workflow/end` again fires exactly once
per started run. The old 'result stays pending' pin is FLIPPED to the
new contract (the pinned behavior was the bug); new regressions cover
cancel-then-settle on a parked script, a never-settling returned
thenable, and the full composition through the REAL registry + tool +
vm engine (tool-workflow gains workflow-vm/subagent devDeps for it).
agentsStarted JSDoc clarified while touching the vocabulary (accepted
calls, including ones still queued at cancellation).
RFCs
One kind of design doc lives here. An RFC records a decision or proposal that shapes this codebase — the why and what we gave up, the parts code and docs can't carry.
Layout and naming
Every RFC has two axes, both encoded in its path — {lifecycle}/{class}/yyyy-mm-dd-topic-title.md:
- Lifecycle (the top-level folder) is the RFC's status, and an RFC moves between folders as that status changes:
proposed/— proposals reviewed before implementation; not yet built (or only partly).implemented/— the decision shipped. The file records what was decided and what was rejected, and is kept current with what actually shipped: when the code later moves a file, renames a package, or changes a key/default, the RFC is updated in the same change to match (facts only — paths, names, structure — not the decision itself). See implemented/AGENTS.md.rejected/— the proposal was considered and declined. Kept for the record so the rejection isn't re-litigated.
- Class (the nested folder) is the kind of decision — see Classification below.
The date in the filename is when the topic was first proposed (per git history). Cross-references between RFCs use relative markdown links ([topic](../../implemented/architecture/2026-…-….md)) — never bare prose or numbers — so they are mechanically checkable and survive moves between folders.
Classification
Each RFC is filed under exactly one class — the kind of decision it records. The class is encoded in the path (the folder is the label, so a file's location declares its class) and the set is closed: scripts/rfc-index.ts owns the canonical set, scripts/verify-rfc-classification.ts rejects any folder outside it, and the index tables below are generated from the tree (pnpm run gen-rfc-index rewrites the marker-delimited regions from each RFC's path, H1 title, and filename date; the gate fails when they are stale). Adding a new class means amending that const and this section, not just dropping a new folder. See the classification RFC for why the taxonomy is path-encoded and gated, and the index-generation RFC for why the tables are generated while this prose stays curated.
| Class | What it covers |
|---|---|
feature |
A new user- or model-facing capability. |
bug-fix |
Corrects a defect or closes a gap a postmortem surfaced. |
simplification |
Removes code, behavior, or surface area without adding a capability. |
architecture |
A structural decision about the shipped source — how packages relate, what the runtime vocabulary is. |
process |
Tooling, policy, or workflow around the code — gates, the package manager, vendoring — not runtime behavior. |
testing |
Test infrastructure and strategy. |
The architecture / process line: architecture is about the source we ship; process is the surrounding tooling and workflow. (refactor is deliberately absent — it overlaps simplification, whose discriminator, "does observable behavior change?", already covers it.)
When to write one
Write an RFC when a decision is durable (it shapes the codebase beyond a single function or package), contested (there was a real alternative a reasonable engineer might have chosen), and surprising (a future reader would otherwise ask "why on earth is it done this way?"). A proposal for substantial future work starts in proposed/; a decision already made starts in implemented/. Pick the class folder that matches the decision (see Classification).
Do NOT write one for a mechanical or local choice (a variable name, a one-file refactor), for anything already enforced and explained by a gate or a convention in AGENTS.md, or for a still-provisional decision tagged TODO(...) in the code — record those as TODOs and promote to an RFC only once they settle. An RFC is never edited into a different decision: supersede it with a new one and cross-link. (Editing an implemented/ RFC to track where its already-made decision now lives — a moved file, a renamed package — is not a different decision and is required, not forbidden; see implemented/AGENTS.md.)
Proposed
Feature
| Title | First proposed |
|---|---|
| Agent Client Protocol (ACP) support — drive the coding agent from external editors | 2026-06-14 |
| Multiplex concurrent ACP sessions over one connection | 2026-06-14 |
| Optional Code Mode — model writes TypeScript against an SDK of all tools | 2026-06-15 |
| Pre-tool input rewrite — a consistent design | 2026-06-30 |
Simplification
| Title | First proposed |
|---|---|
| Unify the agent id and the session id | 2026-06-20 |
Prune dead core-spine surface — SurfaceManager.invalidate(), the loop-internal exports, ToolExecutionResult.callId |
2026-07-04 |
Architecture
| Title | First proposed |
|---|---|
| Runtime schemas for the event vocabulary (Zod vs the merge-extensible-map pattern) | 2026-06-16 |
| Extract a generic long-running tool runtime | 2026-06-20 |
Process
| Title | First proposed |
|---|---|
| API extractor reports | 2026-06-11 |
| Architectural conformance — dependency rules and the adapter kit | 2026-06-11 |
| Supply chain checks and vendor drift verification | 2026-06-11 |
| Discover package inventories instead of maintaining static lists | 2026-06-20 |
Testing
| Title | First proposed |
|---|---|
| Deterministic tests, the replay invariant fixture, and race stress | 2026-06-11 |
| Mutation testing as the coverage counterweight | 2026-06-11 |
Implemented
Feature
Simplification
Architecture
Process
| Title | First proposed |
|---|---|
| Doc-sync enforcement | 2026-06-11 |
| Mechanical quality gates over prose guidelines | 2026-06-11 |
| tsdown for JS bundling instead of dumble | 2026-06-11 |
| Vendor Cordis as source, not npm dependencies | 2026-06-11 |
| pnpm as the package manager instead of Yarn 4 | 2026-06-16 |
| TSC-first build and one tsconfig | 2026-06-17 |
| Markdown cross-link validity linting | 2026-06-18 |
Core-data-structures catalog and the ts type-equiv drift gate |
2026-06-20 |
| Generated cordis events + services catalog | 2026-06-20 |
| Classify RFCs by kind via path-encoded subdirectories | 2026-06-20 |
| Bilingual documentation via paired sibling files and a pairing gate | 2026-07-02 |
| Generated tool-schema catalog (boot-and-harvest) | 2026-07-02 |
| JSDoc completeness gate for the cordis surface | 2026-07-04 |
| Documentation tiers, budgets, and the ceiling gate | 2026-07-04 |
| Generate the RFC index tables | 2026-07-04 |
| Generated persistence log event catalog | 2026-07-04 |
Testing
| Title | First proposed |
|---|---|
| Property-based testing for protocol-shaped code | 2026-06-11 |
| ACP snapshot tests — record-once / replay-deterministic | 2026-06-19 |
| Real-API e2e in CI against the external DeepSeek API | 2026-06-19 |
Use session.jsonl as the only snapshot session-log artifact |
2026-06-20 |
| Persist the seed boundary so fork-child replay routes correctly | 2026-06-22 |
| Record fork and mixed spawn+fork snapshot scenarios | 2026-06-22 |
| Per-session snapshot replay for nested agents | 2026-06-22 |
| Hook snapshot matrix — end-to-end goldens for both bridges | 2026-07-04 |
| Single-source the acp-agent replay config | 2026-07-04 |
Rejected
Simplification
| Title | First proposed |
|---|---|
| Persist assembled assistant messages, not stream chunks | 2026-06-20 |
| Drop ACP session/load until resume has a product shape | 2026-06-20 |
Drop ACP terminal _meta rendering |
2026-06-20 |
| Drop bash full-output spill files | 2026-06-20 |
| Drop durable step boundary events | 2026-06-20 |
| Drop unused session lineage metadata | 2026-06-20 |
| Fold the persistence interface into dsh-session | 2026-06-20 |
| Collapse tool-owned UI presentation | 2026-06-20 |
| Retire mid-turn steering | 2026-06-20 |
| Return the ACP bridge to one live session per connection | 2026-06-20 |
| Truncate interrupted final turns on load | 2026-06-20 |
| Prune the unimplemented subagent seam vocabulary | 2026-07-04 |
Architecture
| Title | First proposed |
|---|---|
| Deep-readonly public surfaces | 2026-06-11 |
| Make the shared example base providerless | 2026-06-20 |