Codex convergence findings on the delegate-and-fold fix (code path verified
correct, prose only):
- The hook-bridges RFC claimed a downstream `block` "carries the bridge context
too" for BOTH seams. True for `tools/post-execute` (PostToolDecision.block has
an additionalContext field) but false for `agent/prompt-submit`
(PromptDecision.block is `{kind,reason}` with no context field). The code is
already correct — a blocked prompt drops the context, which is right since the
prompt never reaches the model. Reworded the RFC to state the per-seam
difference accurately.
- Two test comments narrated "Before the fix…", which the current-state-only
doc rule forbids. Reworded to describe the behavior, not its history.
- Documented on concatContext (both bridges) why the merged block carries a
single source: a HookContext holds one MessageSource and the seam cannot
represent mixed provenance; rendering distinguishes only by source.kind, so a
downstream plugin's text stays framed as plugin context.
Address review on the hook-bridges PR — two composability/compatibility bugs
in both the CC and Codex bridges:
1. A hook that only attaches additionalContext (no block/deny) returned
`allow`/`accept` WITHOUT calling next(), short-circuiting every later
agent/prompt-submit / tools/post-execute listener. A policy/sandbox plugin
registered after the bridge never saw the prompt. Now the context-only path
delegates via next() and folds its context onto the downstream decision
(concatContext): a downstream block/deny still wins and carries the bridge
context; a downstream allow/accept keeps its own content rewrite and gains
the context. Only a real hook deny/block short-circuits.
2. CLAUDE_PROJECT_DIR was empty in the default ACP wiring (no projectDir
configured), breaking common unmodified hooks that reference
$CLAUDE_PROJECT_DIR. It now defaults per-run to the agent's session
workspace (the same cwd the hook runs in); an explicit config.projectDir
still wins.
Regression tests per bridge: a later listener blocks a prompt a context-only
hook allowed; both contexts survive when the downstream also adds one; the
default CLAUDE_PROJECT_DIR reaches the hook. Each proven red on the pre-fix
code.
Address review on the bridges:
- Hook cwd (blocking): the bridges never passed a workdir to runHook, so hooks
ran in the executor default (the ACP server launch dir), not the session
cwd — a hook doing `pwd`/relative reads/marker writes operated in the wrong
tree. Both bridges now thread the agent's session `header.cwd` (the
session/new.cwd) as the hook workdir for agent-scoped points. Regression per
bridge: server cwd ≠ session cwd, a `pwd` hook proves it ran in the session
workspace (proven red without the workdir).
- Example config honesty (blocking): `configPath: ./hooks.json` is read ONCE at
load against the PROCESS cwd, not per-session — the comment/README now say so
explicitly (a project-local per-session hooks.json is not discovered;
TODO(per-session-hook-config)). The hooks-run-in-session-cwd fix above is the
distinct, separately-documented half.
- Session-start timing (blocking): agent/session-start is a synchronous emit and
the hook runs on a detached .then, so injected context is BEST-EFFORT — not
guaranteed before the first request. Downgrade the contract in code comments +
README + RFC (TODO(session-start-gating)) rather than implying "first request
sees it", and add a no-wait regression that asserts the safe properties
without pre-waiting for the inject.
- systemMessage (non-blocking): the merge collects merged.systemMessages but no
bridge surfaced it. Warn per hook (like updatedInput) and document it as
deferred in both READMEs + the RFC; tests assert the warn + non-surfacing.
Round-2 Codex review of the round-1 fixes:
- (A) The Codex plain-stdout→additionalContext fold (F1) was not gated on exit
code, so a NON-clean hook's stdout still injected: a SessionStart `echo stale;
exit 2` (an emit — cannot block) wrongly injected "stale", and a
UserPromptSubmit `exit 1` (non-blocking error → falls through to context) did
too. Gate the fold on `output.exitCode === 0`, matching the codec's own
structured-stdout rule. Guard tests for both paths, proven red without the gate.
- (B) The Codex "SessionStart no-context no-op" absence test was unsound (a
completed turn doesn't prove the detached hook finished). It now touches a
marker and waitFor()s it before asserting no context.
- (B) Both HMR tests used a no-op `true` hook, so a leaked listener would still
pass. They now use a BLOCKING (exit 2) UserPromptSubmit hook and assert the
post-dispose turn is NOT blocked and logs no hook/invoked — a leaked listener
fails loudly.
The bridge tests that drive observe-only emit listeners (session-start,
subagent/start, subagent/end) fire their hook on a detached `.then` the test
cannot await. They waited a fixed 50-80ms, which flaked under the full
test:coverage run's heavy parallel load (transform ~400s): the sleep expired
before the async hook completed, so the injected context / marker file / warn
call had not landed. Replace each fixed sleep with a `waitFor(predicate)` poll
that retries until the observable effect appears (5s deadline) — "async state is
not synchronous state": wait for the signal that actually fires, not a guessed
duration. No behavior change; the same assertions, made robust to scheduling.
Round-1 Codex review findings on the bridges:
- Stop force-continue (both bridges): a blocking Stop hook with EMPTY stderr
yielded decision 'deny' + reason undefined, and the `&& reason !== undefined`
guard let the turn STOP — the opposite of a blocking Stop hook. Force-continue
on any deny; fall back to a generic steering line when there is no reason.
- Codex payload tool_name: hardcoded "Bash" disagreed with the exec.name matcher
subject, so a real Codex `matcher:"Bash"` never fired against the harness's
lowercase `bash` tool. Use exec.name in both payload builders (matches the
matcher subject and the sibling CC bridge). Doc/RFC updated.
- Codex plain-stdout context: SessionStart/UserPromptSubmit are documented to
treat a clean hook's PLAIN (non-JSON) stdout as additionalContext, but nothing
folded it. runPoint now folds plain stdout into context for those two events,
gated on the codec's JSON gate so structured stdout is never dumped as prose.
- continue:false is deferred, not honored: the seams have no hard-halt primitive
yet. TODO(hook-continue-false) at both bridges + an RFC deferred note; the two
tests now assert the LOG records the halt request AND that the run is NOT
actually halted (no longer misleading).
- README concurrency wording: hooks run SERIALLY (deliberate — adjacent
invoked/result log pairs, order-independent fold), not concurrently. Fixed the
CC README claim + an RFC note.
Regression guards proven red on the unfixed code, then reverted. The mismatched-
hookEventName discard (also flagged) is fixed in dsh-hook-protocol and merged down.
The two bridge plugins that run a user's existing Claude Code / Codex hook
config on the harness's typed interception seams, built on the shared
dsh-hook-protocol library. A bridge is a faithfulness adapter, not a power
tool: anything it does a native cordis plugin does more powerfully — the
bridge exists only to run UNMODIFIED external hooks.
- dsh-hooks-claude: CC dialect. Seven hook points (SessionStart,
UserPromptSubmit, PreToolUse, PostToolUse, Stop, SubagentStart,
SubagentStop), CC per-event stdin payloads, env + ${CLAUDE_PLUGIN_ROOT}/
${CLAUDE_PROJECT_DIR} substitution, literal-or-regex matcher.
- dsh-hooks-codex: Codex dialect — a deliberate subset. Five hook points,
always-regex matcher, snake_case payloads (turn_id/model, no trailing
newline), no env/substitution, block-only decisions.
Both map the neutral merged outcome onto the seam's typed Decision and stamp
an explicit {kind:'plugin'} source on injected context (so it is never
mislabeled as a user prompt). Config parse-failure is contained; only command
hooks run. updatedInput is logged+warned (input rewrite deferred); the Stop
loop-guard is deferred (TODO).
Tests: per-file 100% — config-parse unit branches + per-seam mappings
end-to-end through the REAL loop + REAL bash + REAL shell scripts (scripted
mock model only) + a real-Loader export-shape guard. A keyless ACP snapshot
scenario (hook-prompt-block) proves a UserPromptSubmit hook blocks a prompt
end-to-end (rejected turn -> ACP cancelled, hook/* events in the log); a
with-key e2e (hooks.e2e.ts) proves a PreToolUse hook blocks real bash
(verified on disk). The snapshot normalizer now scrubs hook/result.durationMs.
RFC: docs/rfc/implemented/feature/2026-06-30-hook-bridges.md