Merge refreshed docs/i18n-batch-cds-postmortem into docs/i18n-batch-rfc
# Conflicts: # .agents/notes/implemented/architecture/2026-06-11-content-block-vocabulary.i18n.yaml # .agents/notes/implemented/architecture/2026-06-11-content-block-vocabulary.zh.md # .agents/notes/implemented/architecture/2026-06-11-custom-schema-dsl.i18n.yaml # .agents/notes/implemented/architecture/2026-06-11-custom-schema-dsl.zh.md # .agents/notes/implemented/architecture/2026-06-11-dev-invariants-over-deep-readonly.i18n.yaml # .agents/notes/implemented/architecture/2026-06-11-dev-invariants-over-deep-readonly.zh.md # .agents/notes/implemented/architecture/2026-06-11-event-sourced-sessions.i18n.yaml # .agents/notes/implemented/architecture/2026-06-11-event-sourced-sessions.zh.md # .agents/notes/implemented/architecture/2026-06-11-microkernel-event-taxonomy.i18n.yaml # .agents/notes/implemented/architecture/2026-06-11-microkernel-event-taxonomy.zh.md # .agents/notes/implemented/architecture/2026-06-11-runtime-arg-validation.i18n.yaml # .agents/notes/implemented/architecture/2026-06-11-runtime-arg-validation.zh.md # .agents/notes/implemented/architecture/2026-06-11-structured-error-taxonomy.i18n.yaml # .agents/notes/implemented/architecture/2026-06-11-structured-error-taxonomy.zh.md # .agents/notes/implemented/architecture/2026-06-11-tool-schemas-in-prompt-assembly.i18n.yaml # .agents/notes/implemented/architecture/2026-06-11-tool-schemas-in-prompt-assembly.zh.md # .agents/notes/implemented/architecture/2026-06-13-capability-seams.i18n.yaml # .agents/notes/implemented/architecture/2026-06-13-capability-seams.zh.md # .agents/notes/implemented/architecture/2026-06-13-twin-llm-adapters.i18n.yaml # .agents/notes/implemented/architecture/2026-06-13-twin-llm-adapters.zh.md # .agents/notes/implemented/architecture/2026-06-14-session-persistence.i18n.yaml # .agents/notes/implemented/architecture/2026-06-14-session-persistence.zh.md # .agents/notes/implemented/architecture/2026-06-15-turn-enclosure-invariant.i18n.yaml # .agents/notes/implemented/architecture/2026-06-15-turn-enclosure-invariant.zh.md # .agents/notes/implemented/architecture/2026-06-17-filesystem-capability-seam.i18n.yaml # .agents/notes/implemented/architecture/2026-06-17-filesystem-capability-seam.zh.md # .agents/notes/implemented/architecture/2026-06-18-agent-lifecycle-and-ownership-seams.i18n.yaml # .agents/notes/implemented/architecture/2026-06-18-agent-lifecycle-and-ownership-seams.zh.md # .agents/notes/implemented/architecture/2026-06-18-session-surface.i18n.yaml # .agents/notes/implemented/architecture/2026-06-18-session-surface.zh.md # .agents/notes/implemented/architecture/2026-06-18-shared-persistence-write-coordinator.i18n.yaml # .agents/notes/implemented/architecture/2026-06-18-shared-persistence-write-coordinator.zh.md # .agents/notes/implemented/architecture/2026-06-20-branded-ids.i18n.yaml # .agents/notes/implemented/architecture/2026-06-20-branded-ids.zh.md # .agents/notes/implemented/architecture/2026-06-20-extract-example-app-packages.i18n.yaml # .agents/notes/implemented/architecture/2026-06-20-extract-example-app-packages.zh.md # .agents/notes/implemented/architecture/2026-06-20-package-hierarchy.i18n.yaml # .agents/notes/implemented/architecture/2026-06-20-package-hierarchy.md # .agents/notes/implemented/architecture/2026-06-20-package-hierarchy.zh.md # .agents/notes/implemented/architecture/2026-06-21-mandatory-app-attribution-headers.i18n.yaml # .agents/notes/implemented/architecture/2026-06-21-mandatory-app-attribution-headers.zh.md # .agents/notes/implemented/architecture/2026-06-24-web-capability-seam.i18n.yaml # .agents/notes/implemented/architecture/2026-06-24-web-capability-seam.zh.md # .agents/notes/implemented/architecture/2026-06-26-file-context-as-event-gate.i18n.yaml # .agents/notes/implemented/architecture/2026-06-26-file-context-as-event-gate.zh.md # .agents/notes/implemented/architecture/2026-06-30-bash-stdin-env-trusted-plugin-surface.i18n.yaml # .agents/notes/implemented/architecture/2026-06-30-bash-stdin-env-trusted-plugin-surface.zh.md # .agents/notes/implemented/architecture/2026-06-30-event-domain-semantics.i18n.yaml # .agents/notes/implemented/architecture/2026-06-30-event-domain-semantics.zh.md # .agents/notes/implemented/architecture/2026-07-02-fs-per-session-cwd.i18n.yaml # .agents/notes/implemented/architecture/2026-07-02-fs-per-session-cwd.zh.md # .agents/notes/implemented/architecture/2026-07-02-result-time-applied-hunk-diffs.i18n.yaml # .agents/notes/implemented/architecture/2026-07-02-result-time-applied-hunk-diffs.zh.md # .agents/notes/implemented/architecture/2026-07-02-tool-render-intent-union.i18n.yaml # .agents/notes/implemented/architecture/2026-07-02-tool-render-intent-union.zh.md # .agents/notes/implemented/architecture/2026-07-03-filesystem-directory-listing-seam.i18n.yaml # .agents/notes/implemented/architecture/2026-07-03-filesystem-directory-listing-seam.zh.md # .agents/notes/implemented/architecture/2026-07-05-prompt-variables-and-tool-guidance-ownership.i18n.yaml # .agents/notes/implemented/architecture/2026-07-05-prompt-variables-and-tool-guidance-ownership.zh.md # .agents/notes/implemented/architecture/2026-07-05-reconstructable-requests.i18n.yaml # .agents/notes/implemented/architecture/2026-07-05-reconstructable-requests.zh.md # .agents/notes/implemented/architecture/2026-07-05-subagent-provider-lifecycle-events.i18n.yaml # .agents/notes/implemented/architecture/2026-07-05-subagent-provider-lifecycle-events.zh.md # .agents/notes/implemented/architecture/2026-07-06-timeout-deadline-library.i18n.yaml # .agents/notes/implemented/architecture/2026-07-06-timeout-deadline-library.zh.md # .agents/notes/implemented/architecture/2026-07-07-tool-call-timeout-policy.i18n.yaml # .agents/notes/implemented/architecture/2026-07-07-tool-call-timeout-policy.zh.md # .agents/notes/implemented/architecture/2026-07-08-agent-scope-contexts.i18n.yaml # .agents/notes/implemented/architecture/2026-07-08-agent-scope-contexts.zh.md # .agents/notes/implemented/architecture/2026-07-12-agent-scope-runtime-design.i18n.yaml # .agents/notes/implemented/architecture/2026-07-12-agent-scope-runtime-design.zh.md # .agents/notes/implemented/feature/2026-06-14-acp-agent-client-protocol.i18n.yaml # .agents/notes/implemented/feature/2026-06-14-acp-agent-client-protocol.zh.md # .agents/notes/implemented/feature/2026-06-14-acp-multi-session.i18n.yaml # .agents/notes/implemented/feature/2026-06-14-acp-multi-session.zh.md # .agents/notes/implemented/feature/2026-06-15-code-mode.i18n.yaml # .agents/notes/implemented/feature/2026-06-15-code-mode.zh.md # .agents/notes/implemented/feature/2026-06-17-filesystem-tool-schemas.i18n.yaml # .agents/notes/implemented/feature/2026-06-17-filesystem-tool-schemas.zh.md # .agents/notes/implemented/feature/2026-06-18-acp-terminal-and-tool-rendering.i18n.yaml # .agents/notes/implemented/feature/2026-06-18-acp-terminal-and-tool-rendering.zh.md # .agents/notes/implemented/feature/2026-06-18-compaction-capability-seam.i18n.yaml # .agents/notes/implemented/feature/2026-06-18-compaction-capability-seam.zh.md # .agents/notes/implemented/feature/2026-06-21-subagent-capability-seam.i18n.yaml # .agents/notes/implemented/feature/2026-06-21-subagent-capability-seam.md # .agents/notes/implemented/feature/2026-06-21-subagent-capability-seam.zh.md # .agents/notes/implemented/feature/2026-06-22-acp-subagent-backend.i18n.yaml # .agents/notes/implemented/feature/2026-06-22-acp-subagent-backend.zh.md # .agents/notes/implemented/feature/2026-06-25-ask-user-question.i18n.yaml # .agents/notes/implemented/feature/2026-06-25-ask-user-question.zh.md # .agents/notes/implemented/feature/2026-06-29-todo-write-tool.i18n.yaml # .agents/notes/implemented/feature/2026-06-29-todo-write-tool.zh.md # .agents/notes/implemented/feature/2026-06-30-hook-bridges.i18n.yaml # .agents/notes/implemented/feature/2026-06-30-hook-bridges.zh.md # .agents/notes/implemented/feature/2026-06-30-hook-protocol-lib.i18n.yaml # .agents/notes/implemented/feature/2026-06-30-hook-protocol-lib.zh.md # .agents/notes/implemented/feature/2026-06-30-interception-seams.i18n.yaml # .agents/notes/implemented/feature/2026-06-30-interception-seams.zh.md # .agents/notes/implemented/feature/2026-06-30-session-store-fork-api.i18n.yaml # .agents/notes/implemented/feature/2026-06-30-session-store-fork-api.zh.md # .agents/notes/implemented/feature/2026-06-30-subagent-observe-enrich.i18n.yaml # .agents/notes/implemented/feature/2026-06-30-subagent-observe-enrich.zh.md # .agents/notes/implemented/feature/2026-07-05-dynamic-workflows.i18n.yaml # .agents/notes/implemented/feature/2026-07-05-dynamic-workflows.zh.md # .agents/notes/implemented/feature/2026-07-05-skill-system.i18n.yaml # .agents/notes/implemented/feature/2026-07-05-skill-system.zh.md # .agents/notes/implemented/feature/2026-07-06-approval-seam.i18n.yaml # .agents/notes/implemented/feature/2026-07-06-approval-seam.zh.md # .agents/notes/implemented/feature/2026-07-06-explicit-tool-order.i18n.yaml # .agents/notes/implemented/feature/2026-07-06-explicit-tool-order.zh.md # .agents/notes/implemented/feature/2026-07-06-sandbox.i18n.yaml # .agents/notes/implemented/feature/2026-07-06-sandbox.zh.md # .agents/notes/implemented/feature/2026-07-07-mcp-client-plugin.i18n.yaml # .agents/notes/implemented/feature/2026-07-07-mcp-client-plugin.zh.md # .agents/notes/implemented/feature/2026-07-07-session-prefix.i18n.yaml # .agents/notes/implemented/feature/2026-07-07-session-prefix.zh.md # .agents/notes/implemented/feature/2026-07-08-repeat-tool-guard.i18n.yaml # .agents/notes/implemented/feature/2026-07-08-repeat-tool-guard.zh.md # .agents/notes/implemented/feature/2026-07-08-self-referential-cordis-toolset.i18n.yaml # .agents/notes/implemented/feature/2026-07-08-self-referential-cordis-toolset.zh.md # .agents/notes/implemented/feature/2026-07-10-session-query-service.i18n.yaml # .agents/notes/implemented/feature/2026-07-10-session-query-service.zh.md # .agents/notes/implemented/feature/2026-07-12-subagent-persona-tool-filter-and-depth.i18n.yaml # .agents/notes/implemented/feature/2026-07-12-subagent-persona-tool-filter-and-depth.zh.md # .agents/notes/implemented/process/2026-06-11-doc-sync-enforcement.i18n.yaml # .agents/notes/implemented/process/2026-06-11-doc-sync-enforcement.zh.md # .agents/notes/implemented/process/2026-06-11-quality-gates.i18n.yaml # .agents/notes/implemented/process/2026-06-11-quality-gates.md # .agents/notes/implemented/process/2026-06-11-quality-gates.zh.md # .agents/notes/implemented/process/2026-06-11-tsdown-over-dumble.i18n.yaml # .agents/notes/implemented/process/2026-06-11-tsdown-over-dumble.zh.md # .agents/notes/implemented/process/2026-06-11-vendor-cordis-as-source.i18n.yaml # .agents/notes/implemented/process/2026-06-11-vendor-cordis-as-source.zh.md # .agents/notes/implemented/process/2026-06-16-pnpm-over-yarn.i18n.yaml # .agents/notes/implemented/process/2026-06-16-pnpm-over-yarn.zh.md # .agents/notes/implemented/process/2026-06-17-ts-build-config.i18n.yaml # .agents/notes/implemented/process/2026-06-17-ts-build-config.zh.md # .agents/notes/implemented/process/2026-06-18-markdown-cross-link-lint.i18n.yaml # .agents/notes/implemented/process/2026-06-18-markdown-cross-link-lint.zh.md # .agents/notes/implemented/process/2026-06-20-core-data-structures-catalog.i18n.yaml # .agents/notes/implemented/process/2026-06-20-core-data-structures-catalog.zh.md # .agents/notes/implemented/process/2026-06-20-generated-cordis-catalog.i18n.yaml # .agents/notes/implemented/process/2026-06-20-generated-cordis-catalog.zh.md # .agents/notes/implemented/process/2026-06-20-rfc-classification.i18n.yaml # .agents/notes/implemented/process/2026-06-20-rfc-classification.zh.md # .agents/notes/implemented/process/2026-07-02-tool-schema-catalog.i18n.yaml # .agents/notes/implemented/process/2026-07-02-tool-schema-catalog.zh.md # .agents/notes/implemented/process/2026-07-03-documentation-graph-atlas.i18n.yaml # .agents/notes/implemented/process/2026-07-03-documentation-graph-atlas.zh.md # .agents/notes/implemented/process/2026-07-04-cordis-jsdoc-completeness-gate.i18n.yaml # .agents/notes/implemented/process/2026-07-04-cordis-jsdoc-completeness-gate.zh.md # .agents/notes/implemented/process/2026-07-04-doc-tiers-and-budgets.i18n.yaml # .agents/notes/implemented/process/2026-07-04-doc-tiers-and-budgets.zh.md # .agents/notes/implemented/process/2026-07-04-generate-rfc-index-tables.i18n.yaml # .agents/notes/implemented/process/2026-07-04-generate-rfc-index-tables.zh.md # .agents/notes/implemented/process/2026-07-04-persistence-log-catalog.i18n.yaml # .agents/notes/implemented/process/2026-07-04-persistence-log-catalog.zh.md # .agents/notes/implemented/process/2026-07-05-uniform-rfc-format.i18n.yaml # .agents/notes/implemented/process/2026-07-05-uniform-rfc-format.zh.md # .agents/notes/implemented/process/2026-07-06-export-surface-jsdoc-gate.i18n.yaml # .agents/notes/implemented/process/2026-07-06-export-surface-jsdoc-gate.zh.md # .agents/notes/implemented/process/2026-07-06-generated-config-catalog.i18n.yaml # .agents/notes/implemented/process/2026-07-06-generated-config-catalog.zh.md # .agents/notes/implemented/process/2026-07-06-node-engine-floor.i18n.yaml # .agents/notes/implemented/process/2026-07-06-node-engine-floor.zh.md # .agents/notes/implemented/process/2026-07-06-parallel-github-ci-gates.i18n.yaml # .agents/notes/implemented/process/2026-07-06-parallel-github-ci-gates.zh.md # .agents/notes/implemented/process/2026-07-06-parallel-pre-push-gates.i18n.yaml # .agents/notes/implemented/process/2026-07-06-parallel-pre-push-gates.zh.md # .agents/notes/implemented/process/2026-07-10-readme-known-limitations-gate.i18n.yaml # .agents/notes/implemented/process/2026-07-10-readme-known-limitations-gate.zh.md # .agents/notes/implemented/process/2026-07-12-package-model-experience-contract.i18n.yaml # .agents/notes/implemented/process/2026-07-12-package-model-experience-contract.zh.md # .agents/notes/implemented/simplification/2026-06-19-drop-mutable-session-summary.i18n.yaml # .agents/notes/implemented/simplification/2026-06-19-drop-mutable-session-summary.zh.md # .agents/notes/implemented/simplification/2026-06-20-collapse-trace-only-session-events.i18n.yaml # .agents/notes/implemented/simplification/2026-06-20-collapse-trace-only-session-events.zh.md # .agents/notes/implemented/simplification/2026-06-20-drop-unconsumed-llm-adapter-change-event.i18n.yaml # .agents/notes/implemented/simplification/2026-06-20-drop-unconsumed-llm-adapter-change-event.zh.md # .agents/notes/implemented/simplification/2026-06-20-drop-unconsumed-llm-assembled-surfaces.i18n.yaml # .agents/notes/implemented/simplification/2026-06-20-drop-unconsumed-llm-assembled-surfaces.zh.md # .agents/notes/implemented/simplification/2026-06-20-prune-dead-seam-methods.i18n.yaml # .agents/notes/implemented/simplification/2026-06-20-prune-dead-seam-methods.md # .agents/notes/implemented/simplification/2026-06-20-prune-dead-seam-methods.zh.md # .agents/notes/implemented/simplification/2026-06-20-public-agent-stop-surface.i18n.yaml # .agents/notes/implemented/simplification/2026-06-20-public-agent-stop-surface.zh.md # .agents/notes/implemented/simplification/2026-06-20-remove-agent-boundary-mirror-events.i18n.yaml # .agents/notes/implemented/simplification/2026-06-20-remove-agent-boundary-mirror-events.zh.md # .agents/notes/implemented/simplification/2026-06-26-fsspec-style-fs-seam.i18n.yaml # .agents/notes/implemented/simplification/2026-06-26-fsspec-style-fs-seam.zh.md # .agents/notes/implemented/simplification/2026-07-02-remove-stream-chunk-mirror.i18n.yaml # .agents/notes/implemented/simplification/2026-07-02-remove-stream-chunk-mirror.zh.md # .agents/notes/implemented/simplification/2026-07-04-drop-image-content-block.i18n.yaml # .agents/notes/implemented/simplification/2026-07-04-drop-image-content-block.zh.md # .agents/notes/implemented/simplification/2026-07-04-drop-inert-request-knobs.i18n.yaml # .agents/notes/implemented/simplification/2026-07-04-drop-inert-request-knobs.zh.md # .agents/notes/implemented/simplification/2026-07-04-drop-unconsumed-web-observation-surface.i18n.yaml # .agents/notes/implemented/simplification/2026-07-04-drop-unconsumed-web-observation-surface.zh.md # .agents/notes/implemented/simplification/2026-07-04-fold-stdio-ui-helper.i18n.yaml # .agents/notes/implemented/simplification/2026-07-04-fold-stdio-ui-helper.md # .agents/notes/implemented/simplification/2026-07-04-fold-stdio-ui-helper.zh.md # .agents/notes/implemented/simplification/2026-07-04-prune-producerless-vocabulary-variants.i18n.yaml # .agents/notes/implemented/simplification/2026-07-04-prune-producerless-vocabulary-variants.zh.md # .agents/notes/implemented/simplification/2026-07-04-prune-write-only-fs-surface.i18n.yaml # .agents/notes/implemented/simplification/2026-07-04-prune-write-only-fs-surface.zh.md # .agents/notes/implemented/simplification/2026-07-04-remove-agent-steering-mirror.i18n.yaml # .agents/notes/implemented/simplification/2026-07-04-remove-agent-steering-mirror.zh.md # .agents/notes/implemented/simplification/2026-07-04-share-app-bin-boot-glue.i18n.yaml # .agents/notes/implemented/simplification/2026-07-04-share-app-bin-boot-glue.zh.md # .agents/notes/implemented/simplification/2026-07-04-tighten-hook-protocol-contract.i18n.yaml # .agents/notes/implemented/simplification/2026-07-04-tighten-hook-protocol-contract.zh.md # .agents/notes/implemented/simplification/2026-07-04-trim-acp-bridge-unreachable-surface.i18n.yaml # .agents/notes/implemented/simplification/2026-07-04-trim-acp-bridge-unreachable-surface.zh.md # .agents/notes/implemented/simplification/2026-07-12-drop-unconsumed-skill-provider-events.i18n.yaml # .agents/notes/implemented/simplification/2026-07-12-drop-unconsumed-skill-provider-events.zh.md # .agents/notes/implemented/simplification/2026-07-12-prune-unused-web-seam-fields.i18n.yaml # .agents/notes/implemented/simplification/2026-07-12-prune-unused-web-seam-fields.zh.md # .agents/notes/implemented/testing/2026-06-11-property-based-testing.i18n.yaml # .agents/notes/implemented/testing/2026-06-11-property-based-testing.zh.md # .agents/notes/implemented/testing/2026-06-19-acp-snapshot-tests.i18n.yaml # .agents/notes/implemented/testing/2026-06-19-acp-snapshot-tests.zh.md # .agents/notes/implemented/testing/2026-06-19-real-api-e2e-ci.i18n.yaml # .agents/notes/implemented/testing/2026-06-19-real-api-e2e-ci.zh.md # .agents/notes/implemented/testing/2026-06-20-remove-redundant-snapshot-log-goldens.i18n.yaml # .agents/notes/implemented/testing/2026-06-20-remove-redundant-snapshot-log-goldens.zh.md # .agents/notes/implemented/testing/2026-06-22-fork-child-replay-seed-boundary.i18n.yaml # .agents/notes/implemented/testing/2026-06-22-fork-child-replay-seed-boundary.zh.md # .agents/notes/implemented/testing/2026-06-22-fork-snapshot-scenarios.i18n.yaml # .agents/notes/implemented/testing/2026-06-22-fork-snapshot-scenarios.zh.md # .agents/notes/implemented/testing/2026-06-22-subagent-snapshot-replay.i18n.yaml # .agents/notes/implemented/testing/2026-06-22-subagent-snapshot-replay.zh.md # .agents/notes/implemented/testing/2026-07-04-hook-snapshot-matrix.i18n.yaml # .agents/notes/implemented/testing/2026-07-04-hook-snapshot-matrix.zh.md # .agents/notes/implemented/testing/2026-07-04-single-source-acp-replay-config.i18n.yaml # .agents/notes/implemented/testing/2026-07-04-single-source-acp-replay-config.zh.md # .agents/notes/implemented/testing/2026-07-06-pin-request-header-content-in-one-scenario.i18n.yaml # .agents/notes/implemented/testing/2026-07-06-pin-request-header-content-in-one-scenario.zh.md # .agents/notes/implemented/testing/2026-07-08-shared-acp-snapshot-package.i18n.yaml # .agents/notes/implemented/testing/2026-07-08-shared-acp-snapshot-package.zh.md # .agents/notes/proposed/architecture/2026-06-16-typed-event-schemas.i18n.yaml # .agents/notes/proposed/architecture/2026-06-16-typed-event-schemas.zh.md # .agents/notes/proposed/architecture/2026-06-20-generic-long-running-tool-runtime.i18n.yaml # .agents/notes/proposed/architecture/2026-06-20-generic-long-running-tool-runtime.zh.md # .agents/notes/proposed/feature/2026-06-30-pre-tool-input-rewrite.i18n.yaml # .agents/notes/proposed/feature/2026-06-30-pre-tool-input-rewrite.zh.md # .agents/notes/proposed/feature/2026-07-07-claude-code-and-codex-subagent-backends.i18n.yaml # .agents/notes/proposed/feature/2026-07-07-claude-code-and-codex-subagent-backends.zh.md # .agents/notes/proposed/feature/2026-07-08-interactive-side-sessions.i18n.yaml # .agents/notes/proposed/feature/2026-07-08-interactive-side-sessions.zh.md # .agents/notes/proposed/feature/2026-07-10-sqlite-session-query-provider.i18n.yaml # .agents/notes/proposed/feature/2026-07-10-sqlite-session-query-provider.zh.md # .agents/notes/proposed/feature/2026-07-13-stream-workflow-progress-through-tool-calls.i18n.yaml # .agents/notes/proposed/feature/2026-07-13-stream-workflow-progress-through-tool-calls.zh.md # .agents/notes/proposed/process/2026-06-11-api-extractor-reports.i18n.yaml # .agents/notes/proposed/process/2026-06-11-api-extractor-reports.md # .agents/notes/proposed/process/2026-06-11-api-extractor-reports.zh.md # .agents/notes/proposed/process/2026-06-11-architectural-conformance.i18n.yaml # .agents/notes/proposed/process/2026-06-11-architectural-conformance.zh.md # .agents/notes/proposed/process/2026-06-11-supply-chain-and-vendor-drift.i18n.yaml # .agents/notes/proposed/process/2026-06-11-supply-chain-and-vendor-drift.zh.md # .agents/notes/proposed/process/2026-06-20-discover-package-inventory.i18n.yaml # .agents/notes/proposed/process/2026-06-20-discover-package-inventory.zh.md # .agents/notes/proposed/simplification/2026-06-20-unify-agent-and-session-id.i18n.yaml # .agents/notes/proposed/simplification/2026-06-20-unify-agent-and-session-id.zh.md # .agents/notes/proposed/simplification/2026-07-04-prune-dead-core-spine-surface.i18n.yaml # .agents/notes/proposed/simplification/2026-07-04-prune-dead-core-spine-surface.zh.md # .agents/notes/proposed/simplification/2026-07-12-simplify-session-log-representation.i18n.yaml # .agents/notes/proposed/simplification/2026-07-12-simplify-session-log-representation.zh.md # .agents/notes/proposed/testing/2026-06-11-deterministic-and-stress-testing.i18n.yaml # .agents/notes/proposed/testing/2026-06-11-deterministic-and-stress-testing.zh.md # .agents/notes/proposed/testing/2026-06-11-mutation-testing.i18n.yaml # .agents/notes/proposed/testing/2026-06-11-mutation-testing.zh.md # .agents/notes/rejected/architecture/2026-06-11-immutable-public-surfaces.i18n.yaml # .agents/notes/rejected/architecture/2026-06-11-immutable-public-surfaces.zh.md # .agents/notes/rejected/architecture/2026-06-20-providerless-example-base.i18n.yaml # .agents/notes/rejected/architecture/2026-06-20-providerless-example-base.zh.md # .agents/notes/rejected/simplification/2026-06-20-assembled-assistant-messages-only.i18n.yaml # .agents/notes/rejected/simplification/2026-06-20-assembled-assistant-messages-only.zh.md # .agents/notes/rejected/simplification/2026-06-20-drop-acp-session-load.i18n.yaml # .agents/notes/rejected/simplification/2026-06-20-drop-acp-session-load.zh.md # .agents/notes/rejected/simplification/2026-06-20-drop-acp-terminal-meta.i18n.yaml # .agents/notes/rejected/simplification/2026-06-20-drop-acp-terminal-meta.zh.md # .agents/notes/rejected/simplification/2026-06-20-drop-bash-output-spill-files.i18n.yaml # .agents/notes/rejected/simplification/2026-06-20-drop-bash-output-spill-files.zh.md # .agents/notes/rejected/simplification/2026-06-20-drop-durable-step-boundaries.i18n.yaml # .agents/notes/rejected/simplification/2026-06-20-drop-durable-step-boundaries.zh.md # .agents/notes/rejected/simplification/2026-06-20-drop-unused-session-lineage.i18n.yaml # .agents/notes/rejected/simplification/2026-06-20-drop-unused-session-lineage.zh.md # .agents/notes/rejected/simplification/2026-06-20-fold-session-persistence-interface.i18n.yaml # .agents/notes/rejected/simplification/2026-06-20-fold-session-persistence-interface.zh.md # .agents/notes/rejected/simplification/2026-06-20-generic-tool-rendering.i18n.yaml # .agents/notes/rejected/simplification/2026-06-20-generic-tool-rendering.zh.md # .agents/notes/rejected/simplification/2026-06-20-retire-mid-turn-steering.i18n.yaml # .agents/notes/rejected/simplification/2026-06-20-retire-mid-turn-steering.zh.md # .agents/notes/rejected/simplification/2026-06-20-single-session-acp-bridge.i18n.yaml # .agents/notes/rejected/simplification/2026-06-20-single-session-acp-bridge.zh.md # .agents/notes/rejected/simplification/2026-06-20-truncate-interrupted-turns.i18n.yaml # .agents/notes/rejected/simplification/2026-06-20-truncate-interrupted-turns.zh.md # .agents/notes/rejected/simplification/2026-07-04-prune-unimplemented-subagent-vocabulary.i18n.yaml # .agents/notes/rejected/simplification/2026-07-04-prune-unimplemented-subagent-vocabulary.zh.md # .agents/notes/rejected/simplification/2026-07-12-collapse-workflow-to-foreground-core.i18n.yaml # .agents/notes/rejected/simplification/2026-07-12-collapse-workflow-to-foreground-core.zh.md # .agents/notes/rejected/simplification/2026-07-12-prune-unused-skill-registry-surface.i18n.yaml # .agents/notes/rejected/simplification/2026-07-12-prune-unused-skill-registry-surface.zh.md # docs/rfc/implemented/architecture/2026-06-18-agent-lifecycle-and-ownership-seams.md # docs/rfc/implemented/architecture/2026-06-18-session-surface.md # docs/rfc/implemented/architecture/2026-06-20-branded-ids.md # docs/rfc/implemented/architecture/2026-07-02-fs-per-session-cwd.md # docs/rfc/implemented/feature/2026-06-18-compaction-capability-seam.md # docs/rfc/implemented/feature/2026-07-07-session-prefix.md # docs/rfc/implemented/process/2026-06-20-rfc-classification.md # docs/rfc/implemented/process/2026-07-04-generate-rfc-index-tables.md # docs/rfc/implemented/process/2026-07-05-uniform-rfc-format.md # docs/rfc/implemented/process/2026-07-06-parallel-github-ci-gates.md # docs/rfc/implemented/process/2026-07-06-parallel-pre-push-gates.md # docs/rfc/implemented/process/2026-07-12-package-model-experience-contract.md # docs/rfc/implemented/simplification/2026-06-20-remove-agent-boundary-mirror-events.md # docs/rfc/implemented/simplification/2026-07-04-prune-producerless-vocabulary-variants.md # docs/rfc/implemented/testing/2026-06-20-remove-redundant-snapshot-log-goldens.md # docs/rfc/implemented/testing/2026-07-08-shared-acp-snapshot-package.md # docs/rfc/proposed/architecture/2026-06-20-generic-long-running-tool-runtime.md # docs/rfc/proposed/simplification/2026-06-20-unify-agent-and-session-id.md # docs/rfc/proposed/simplification/2026-07-12-simplify-session-log-representation.md # docs/rfc/rejected/simplification/2026-07-04-prune-unimplemented-subagent-vocabulary.md # scripts/translation-pairing.manifest.json
This commit is contained in:
2944 files changed
+205120
-33445
No files matched your search
@@ -0,0 +1,6 @@
|
||||
# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write
|
||||
2026-06-11-property-based-testing.md: 153584d3a2b77c8f2d103db02646f18a9d424b57
|
||||
2026-06-11-property-based-testing.zh.md: e11d11f7db9eee97ab81bc678afab5d0fea36bf9
|
||||
@@ -0,0 +1,29 @@
|
||||
# Agent Note: Property-based testing for protocol-shaped code
|
||||
|
||||
Status: implemented
|
||||
|
||||
English | [中文](2026-06-11-property-based-testing.zh.md)
|
||||
|
||||
> Merges the original proposal and the decision record for one topic. It found a real BlockAssembler duplicate-`block-end` bug on first run.
|
||||
|
||||
## Problem
|
||||
|
||||
Example-based tests pin the cases we thought of. The harness's core is protocol-shaped — chunk streams, event logs, schema conversion, inbox scheduling — where the input space is combinatorial and the interesting bugs live in interleavings nobody wrote an example for. The motivating evidence: a block-assembly ordering bug once survived 100% line coverage of the happy paths. Per-file 100% coverage proves every line ran, not that every interleaving is correct.
|
||||
|
||||
## Decision
|
||||
|
||||
Adopt `fast-check` (a root devDependency) with one `tests/properties.spec.ts` per protocol-shaped package, generators tuned for *realistic-but-adversarial* inputs (not uniform noise) and `numRuns` kept so the suite stays well under ~10s locally. Failures print a reproducible seed. (The original proposal also sketched a nightly CI job running 100× the iterations; that was not shipped — the property suite runs only in the normal `push`/`pull_request` CI, and a scheduled high-iteration job remains possible future work.)
|
||||
|
||||
- **dsh-llm / BlockAssembler:** arbitrary chunk streams (valid + malformed: duplicate indices, stragglers, missing block-start). Invariants: `blocks()` count ≤ distinct indices seen; re-assembly idempotent (`blocks()` is stable across repeated calls and `message().content` mirrors it); `blocks()` never throws and yields only valid content-block tags; `finish` reflects the last `finish` chunk, defaulting to `{kind:'stop'}` when none arrives.
|
||||
- **dsh-session:** arbitrary event logs. Invariants: `deriveMessages` deterministic; replay-from-seed identical; seq strictly monotonic; non-message events never affect derived history; derived content is decoupled from the log.
|
||||
- **dsh-tools:** arbitrary `SchemaSpec`. Invariants: JSON Schema `required` equals the `required:true` keys at every level; conversion total; **and the composition with [runtime arg validation](../architecture/2026-06-11-runtime-arg-validation.md)** — generated args satisfying a spec pass `validateArgs`, and targeted corruptions (dropped required key, non-object top level) are rejected. This closes the validator/`InferArgs` drift risk.
|
||||
- **dsh-agent-loop:** arbitrary send schedules against a never-exhausting adapter, driven through the `agent/status` settle signal (no wall-clock sleeps). Invariants: no message lost; turn numbers strictly increase; status transitions stay on the legal machine.
|
||||
|
||||
## Consequences
|
||||
|
||||
- Generator quality is the value lever — the generators bias toward small index pools and short strings so collisions and interleavings are common.
|
||||
- **It already paid off:** the BlockAssembler stream found a real bug — a duplicate `block-end` at the same index rewrote a completed block. Fixed (first close wins, matching the existing straggler rule) with a dedicated regression test.
|
||||
- A property flake from a timeout is a finding, not something to retry away. The loop properties are deterministic by construction (settle on `agent/status`), so a hang is a real defect.
|
||||
- Property tests supplement, not replace, the example tests that pin specific branches for the 100%-coverage gate.
|
||||
|
||||
<!-- agent-note-format: alternatives-not-recorded (pre-format Agent Note) -->
|
||||
@@ -0,0 +1,29 @@
|
||||
# RFC: 对协议形态代码进行基于属性的测试
|
||||
|
||||
Status: implemented
|
||||
|
||||
[English](2026-06-11-property-based-testing.md) | 中文
|
||||
|
||||
> 将原始提案与同一主题的决策记录合并为一篇。首次运行即发现了 BlockAssembler 重复 `block-end` 的真实 bug。
|
||||
|
||||
## 问题
|
||||
|
||||
基于示例的测试只能固定我们想到的用例。harness 的核心是协议形态的代码:分片流、事件日志、schema 转换、收件箱调度。这些场景的输入空间是组合式的,有趣的 bug 藏在没人写过示例的交错序列中。佐证:一个块组装的排序 bug 曾在 happy path 100% 行覆盖率下存活。逐文件 100% 覆盖率证明每一行都跑过了,但不能证明每种交错都是正确的。
|
||||
|
||||
## 决策
|
||||
|
||||
引入 `fast-check`(作为根 devDependency),在每个协议形态的包(package)中编写一个 `tests/properties.spec.ts`。生成器调优为*逼真但对抗性*的输入(而非均匀噪声),`numRuns` 控制在本地套件总耗时远低于约 10 秒。失败时打印可复现的 seed。(原始提案还草拟了一个夜间 CI job,以 100 倍迭代运行;该部分未交付。属性测试套件仅在常规的 `push`/`pull_request` CI 中运行,定时高迭代 job 仍属可能的后续工作。)
|
||||
|
||||
- **dsh-llm / BlockAssembler:** 任意分片流(合法 + 畸形:重复索引、滞后分片、缺少 block-start)。不变式:`blocks()` 计数 ≤ 已见到的不同索引数;重组幂等(`blocks()` 在重复调用间稳定,且 `message().content` 与之一致);`blocks()` 从不抛异常且仅产出合法的 content-block 标签;`finish` 反映最后一个 `finish` 分片,无此类分片时默认为 `{kind:'stop'}`。
|
||||
- **dsh-session:** 任意事件日志。不变式:`deriveMessages` 确定性;从 seed 回放结果一致;seq 严格单调递增;非消息事件不影响推导出的历史;推导出的内容与日志解耦。
|
||||
- **dsh-tools:** 任意 `SchemaSpec`。不变式:JSON Schema 的 `required` 等于每一层 `required:true` 的键集;转换是全函数;**并且与[运行时参数校验](../architecture/2026-06-11-runtime-arg-validation.md)组合验证**——满足 spec 的生成参数通过 `validateArgs`,而定向破坏(删除必填键、顶层非对象)被拒绝。这封堵了 validator 与 `InferArgs` 漂移的风险。
|
||||
- **dsh-agent-loop:** 任意发送调度,对接一个永不耗尽的适配器,通过 `agent/status` settle 信号驱动(无挂钟 sleep)。不变式:无消息丢失;轮次编号严格递增;状态转换保持在合法状态机上。
|
||||
|
||||
## 后果
|
||||
|
||||
- 生成器质量是价值杠杆——生成器偏向小索引池和短字符串,使碰撞与交错频繁发生。
|
||||
- **已经产出回报:** BlockAssembler 流测试发现了一个真实 bug——同一索引的重复 `block-end` 覆盖了已刷出的块,导致流式前缀与最终 `blocks()` 不一致。已修复(首次关闭生效,与既有的滞后分片规则一致),并附带一个专门的回归测试。
|
||||
- 属性测试因超时而 flake 是一个发现,不应通过重试消除。循环属性测试在设计上是确定性的(通过 `agent/status` settle),因此挂起即为真实缺陷。
|
||||
- 属性测试是对示例测试的补充而非替代;示例测试固定特定分支,服务于 100% 覆盖率门禁。
|
||||
|
||||
<!-- rfc-format: alternatives-not-recorded (pre-format RFC) -->
|
||||
@@ -0,0 +1,6 @@
|
||||
# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write
|
||||
2026-06-19-acp-snapshot-tests.md: c336b4864b73b8db29c0a8bb983d974348a9515a
|
||||
2026-06-19-acp-snapshot-tests.zh.md: bc9488562b0698b172ccffff74815a893acf722d
|
||||
@@ -0,0 +1,84 @@
|
||||
# Agent Note: ACP snapshot tests — record-once / replay-deterministic
|
||||
|
||||
Status: implemented
|
||||
|
||||
English | [中文](2026-06-19-acp-snapshot-tests.zh.md)
|
||||
|
||||
## Problem
|
||||
|
||||
Unit tests do not exercise the complete ACP subprocess transcript, while real-API tests are nondeterministic and key-gated. Editor-facing `session/update` output can therefore regress despite green unit coverage, as the [default-export postmortem](../../../../docs/postmortem/0001-acp-default-export-drops-inject.md) demonstrated.
|
||||
|
||||
The blocker for a full-transcript test is the model: the agent's output is driven by a non-deterministic LLM, and a key-gated test that hits the real API on every run is neither deterministic nor CI-runnable. We want the fidelity of a real run with the determinism of a fixture.
|
||||
|
||||
This Agent Note records the decision to add a third test tier — **snapshot tests** — and the design choices that make it deterministic, keyless-in-CI, and cheap to maintain.
|
||||
|
||||
## Decision
|
||||
|
||||
A snapshot test boots the real ACP example, drives its stdio protocol from a deterministic script, and compares normalized output with committed expected outputs. A session log recorded once from the real API supplies all later model streams. The fixture is the product's ordinary persisted JSONL.
|
||||
|
||||
### The fixture is the persisted session JSONL
|
||||
|
||||
Each scenario's `session.jsonl` is harvested from a real run. `assistant/chunk` events reproduce the model streams; tool, message, and boundary events capture the harness behavior. One ordinary session artifact therefore serves as both replay source and behavioral expected output.
|
||||
|
||||
When a scenario pins an alternative physical storage layout, its fixture is mechanically derived from a real unpacked counterpart. The scenario test requires every intended storage-row kind and exact event-for-event equality after decoding before the ordinary replay and log comparison proves that the assembled process consumes and reproduces that layout.
|
||||
|
||||
### Replay derives the model script from the log
|
||||
|
||||
`llm-replay` short-circuits the provider-agnostic `llm/stream` waterfall. `deriveReplayScript()` groups recorded chunks by `(turn, step)` and serves one group per model call. The loop makes one stream call per step, so the grouping is exact and includes error finish chunks without special handling.
|
||||
|
||||
### The in-memory replay entry honors the full LLM contract
|
||||
|
||||
`deriveReplayScript` produces a list of `ReplayEntry`, the in-memory unit the replay listener serves positionally:
|
||||
|
||||
```
|
||||
{ kind: 'chunks', chunks: StreamChunk[] }
|
||||
| { kind: 'throw', chunks: StreamChunk[], message: string, code: string }
|
||||
| { kind: 'hang' }
|
||||
```
|
||||
|
||||
Logs derive chunk entries. Pre-stream throws and hangs have no reconstructable chunk representation, so those scenarios provide `replay.override.json`. A throw entry may include prefix chunks for mid-stream failure. Explicit overrides avoid inferring adapter behavior from lossy turn-end reasons.
|
||||
|
||||
### Positional replay, one in-flight stream
|
||||
|
||||
Replay is positional and therefore permits only one in-flight model stream per scenario. Concurrent-session snapshots require request-keyed entries. Changed call order requires re-recording, and missing or exhausted fixtures fail loudly.
|
||||
|
||||
### Recording harvests the log; keyless replay needs a providerless config
|
||||
|
||||
Recording runs the scenario with the real `llm-deepseek` adapter and the JSONL persistence backend configured with `persistenceCompression: 'none'`, then copies the produced `.jsonl` into the scenario dir. The explicit raw mode keeps committed replay fixtures line-readable while ordinary deployments use the backend's compressed default. Per-event appends are durable, but the harness shuts the subprocess down gracefully (close stdin → `await ctx.dispose()`) before harvesting so the final events are flushed. `llm-replay` itself does no recording — it is replay-only.
|
||||
|
||||
Replay uses a `cordis.snapshot.yml` overlay that replaces the real adapter with `llm-replay` while retaining the live composition. Recording uses the ordinary config and a harness-supplied persistence root. Replay mode skips `.env` loading, so a stray API key cannot trigger a live call. See the [single-source config Agent Note](2026-07-04-single-source-acp-replay-config.md).
|
||||
|
||||
### Two surfaces: normalize, then compare
|
||||
|
||||
A snapshot run asserts **two** normalized surfaces, because the harness's external surfaces are distinct:
|
||||
|
||||
1. The **stdout transcript** — the framed `session/update` JSON-RPC the editor sees. Catches regressions in the ACP bridge's event→update translation (`streamSessionEventUpdate`). Compared against a committed `stdout.expected.jsonl`.
|
||||
2. The **re-persisted session JSONL**, normalized and compared with `session.jsonl`. The same fixture is both replay source and expected log. Prompt text is scrubbed; one scenario per header class pins readable prompt and tool content as described in the [header-pinning Agent Note](2026-07-06-pin-request-header-content-in-one-scenario.md). Override scenarios derive model behavior solely from their sidecar.
|
||||
|
||||
The surfaces are complementary: stdout covers bridge projection, while JSONL covers loop, tool, and boundary structure that the projection omits.
|
||||
|
||||
Normalization replaces session, cwd, protocol-id, timestamp, path, and process volatility while preserving deterministic sequence numbers. Scenarios constrain real bash use to stable commands. The stdout expected output remains wire-shaped JSONL and every raw line must parse as JSON. Vitest updates only the stdout expected output; normalized session equality never overwrites the replay fixture.
|
||||
|
||||
### Isolation: normalization now, sandbox later
|
||||
|
||||
Tool determinism comes from a generated cwd, scrubbed environment, fresh non-login shell, constrained commands, and normalization. The cwd defaults to the platform temp directory; a scenario can instead supply its parent when temp is an always-writable policy root and the behavior needs an independent project location. Concurrent replay runs own separate cwd, persistence, and fixed-length scenario-keyed spill roots, so one scenario's teardown cannot delete another's in-flight full-output recovery while real-path preview budgets remain stable. This tier does not claim OS confinement. A sandboxed executor can replace the local backend through the existing [capability seam](../architecture/2026-06-13-capability-seams.md) if a stronger tier is needed.
|
||||
|
||||
### The replay plugin is its own package
|
||||
|
||||
`@deepseek-ai/dsh-llm-replay` is a support package rather than example-local glue. It replaces the real adapter by short-circuiting `llm/stream` with streams reconstructed from JSONL, and its package placement keeps the replay logic under normal coverage gates.
|
||||
|
||||
### Two subcommands, replay in the default gate
|
||||
|
||||
`pnpm run test:snapshot` replays committed fixtures keylessly; `test:snapshot:record` uses the real API and rewrites the harvested session log and stdout expected output. Missing fixtures fail loud. Every scenario carries `input.json`, `stdout.expected.jsonl`, and `session.jsonl`; no-model cases use a header-only log. `replay.override.json` is required only for scenarios marked `overridden`, because its presence replaces derived replay. Fixture guards reject missing, mismatched, and orphaned files. Both commands accept scenario filters.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
- **A hand-authored `llm.json` of model chunks** — the earlier draft; reusing the real session log makes the fixture a genuine product of the system rather than a hand-built mock, and doubles it as a behavioral expected output.
|
||||
- **A byte-level HTTP-record library (Polly/nock/MSW)** — rejected: adapter-specific, awkward with streaming SSE, and lower-level than the thing under test.
|
||||
- **Synthesizing throw/cancel entries from `turn/end {kind:'error'|'aborted'}`** — rejected: it couples `llm-replay` to loop-internal turn-closing semantics, and the `turn/end` reason is lossy (it cannot distinguish a thrown 401 from a finish-error); the explicit `replay.override.json` sidecar is the cleaner seam.
|
||||
|
||||
## Consequences
|
||||
|
||||
The new tier adds reviewed per-scenario input, session, stdout, optional override, and optional workspace fixtures. Workspace seeds are copied into the generated cwd for both record and replay. In return the tier provides deterministic keyless transcript coverage through the real Loader and tool composition. The subprocess, input, workspace, normalization, and replay harness can support examples beyond ACP.
|
||||
|
||||
This Agent Note relates to but does not supersede the [proposed determinism Agent Note](../../proposed/testing/2026-06-11-deterministic-and-stress-testing.md): that proposal's "universal replay fixture" re-derives session *message history* after every test (an internal-consistency invariant), whereas snapshot tests pin the *external protocol output*. They are complementary — one guards the event-sourcing invariant, the other guards the editor-facing contract.
|
||||
@@ -0,0 +1,82 @@
|
||||
# RFC: ACP 快照测试——一次录制 / 确定性回放
|
||||
|
||||
Status: implemented
|
||||
|
||||
[English](2026-06-19-acp-snapshot-tests.md) | 中文
|
||||
|
||||
## 问题
|
||||
|
||||
单元测试无法覆盖完整的 ACP(Agent Client Protocol)子进程 transcript(文本记录),而真实 API 测试既不确定又需要密钥。因此,面向编辑器的 `session/update` 输出可能在单元覆盖率全绿的情况下发生回归,正如 [default-export 事后分析](../../../postmortem/0001-acp-default-export-drops-inject.md)所揭示的那样。
|
||||
|
||||
全 transcript 测试的阻塞因素在于模型:agent 的输出由非确定性的 LLM(大语言模型)驱动,而每次运行都命中真实 API 的密钥门控测试既不确定也无法在 CI 中运行。我们需要真实运行的保真度与 fixture(测试前置数据)的确定性兼得。
|
||||
|
||||
本 RFC 记录了新增第三层测试——**快照测试**——的决策,以及使其具备确定性、CI 中无需密钥、维护成本低的设计选择。
|
||||
|
||||
## 决策
|
||||
|
||||
快照测试启动真实的 ACP 示例,通过确定性脚本驱动其 stdio 协议,并将归一化后的输出与已提交的 golden 文件比对。一次从真实 API 录制的会话日志为后续所有模型流提供数据。fixture 就是产品正常持久化的 JSONL。
|
||||
|
||||
### fixture 即持久化的会话 JSONL
|
||||
|
||||
每个场景的 `session.jsonl` 从一次真实运行中采集。`assistant/chunk` 事件重现模型流;tool、message 和 boundary 事件捕获 harness 行为。一份普通的会话产物因此同时充当回放源和行为 golden。
|
||||
|
||||
### 回放从日志推导模型脚本
|
||||
|
||||
`llm-replay` 短路了提供方无关的 `llm/stream` waterfall(瀑布式事件)。`deriveReplayScript()` 按 `(turn, step)` 对已录制的 chunk 分组,每次模型调用服务一组。agent loop(智能体循环)每个 step 发起一次流调用,因此分组精确对应,错误结束 chunk 也无需特殊处理。
|
||||
|
||||
### 内存中的回放条目遵守完整的 LLM 契约
|
||||
|
||||
`deriveReplayScript` 产出一组 `ReplayEntry`,即回放监听器按位置服务的内存单元:
|
||||
|
||||
```
|
||||
{ kind: 'chunks', chunks: StreamChunk[] }
|
||||
| { kind: 'throw', chunks: StreamChunk[], message: string, code: string, status?: number }
|
||||
| { kind: 'hang' }
|
||||
```
|
||||
|
||||
日志推导出 chunk 条目。流开始前的抛出和挂起没有可重建的 chunk 表示,因此这些场景提供 `replay.override.json`。throw 条目可以包含前缀 chunk 以模拟流中途失败。显式覆盖避免了从有损的轮次结束原因推断适配器行为。
|
||||
|
||||
### 位置式回放,单个在途流
|
||||
|
||||
回放是位置式的,因此每个场景只允许一个在途模型流。并发会话快照需要按请求键索引的条目。调用顺序变更需要重新录制,fixture 缺失或耗尽时立即报错。
|
||||
|
||||
### 录制采集日志;无密钥回放需要无提供方的配置
|
||||
|
||||
录制使用真实的 `llm-deepseek` 适配器和 JSONL 持久化后端运行场景,然后将产出的 `.jsonl` 复制到场景目录。逐事件追加是持久的,但 harness 在采集前会优雅关闭子进程(关闭 stdin → `await ctx.dispose()`),确保最终事件已刷盘。`llm-replay` 本身不做录制,它只负责回放。
|
||||
|
||||
回放使用 `cordis.snapshot.yml` 覆盖配置,将真实适配器替换为 `llm-replay`,同时保留活跃的组合。录制使用普通配置和 harness 提供的持久化根目录。回放模式跳过 `.env` 加载,因此一个意外存在的 API key 不会触发真实调用。见[单源配置 RFC](2026-07-04-single-source-acp-replay-config.md)。
|
||||
|
||||
### 两个表面:归一化后比对
|
||||
|
||||
快照运行断言**两个**归一化后的表面,因为 harness 的外部表面是不同的:
|
||||
|
||||
1. **stdout transcript**——编辑器看到的带帧 `session/update` JSON-RPC。捕获 ACP bridge 事件→update 转换(`streamSessionEventUpdate`)中的回归。与已提交的 `stdout.golden.jsonl` 比对。
|
||||
2. **重新持久化的会话 JSONL**,归一化后与 `session.jsonl` 比对。同一份 fixture 既是回放源也是预期日志。提示词文本被擦除;每个 header 类别一个场景固定可读的 prompt 和 tool 内容,见 [header-pinning RFC](2026-07-06-pin-request-header-content-in-one-scenario.md)。覆盖场景的模型行为完全来自其伴随文件。
|
||||
|
||||
两个表面互补:stdout 覆盖 bridge 投影,JSONL 覆盖投影所省略的 loop、tool 和 boundary 结构。
|
||||
|
||||
归一化替换 session、cwd、protocol-id、时间戳、路径和进程相关的易变值,同时保留确定性序列号。场景将真实 bash 使用限制在稳定命令范围内。stdout golden 保持协议格式(wire format)的 JSONL,每一行原始数据必须可解析为 JSON。Vitest 只更新 stdout golden;归一化后的会话相等性检查从不覆盖回放 fixture。
|
||||
|
||||
### 隔离:当前靠归一化,后续可加沙箱
|
||||
|
||||
工具的确定性来自临时 cwd、擦除的环境变量、全新的非登录 shell、受限命令和归一化。它不声称具备操作系统级隔离。如果需要更强的隔离层级,可通过既有的[能力 seam](../architecture/2026-06-13-capability-seams.md) 将沙箱执行器替换本地后端。
|
||||
|
||||
### 回放插件是独立的包
|
||||
|
||||
`@deepseek-ai/dsh-llm-replay` 是一个支撑包(package),而非示例本地的胶水代码。它通过用从 JSONL 重建的流短路 `llm/stream` 来替换真实适配器,其包级放置使回放逻辑处于正常覆盖率门禁之下。
|
||||
|
||||
### 两个子命令,回放在默认门禁中
|
||||
|
||||
`pnpm run test:snapshot` 无需密钥地回放已提交的 fixture;`test:snapshot:record` 使用真实 API 并重写采集到的会话日志和 stdout golden。fixture 缺失时立即报错。每个场景携带 `input.json`、`stdout.golden.jsonl` 和 `session.jsonl`;无模型场景使用仅含 header 的日志。`replay.override.json` 仅在标记为 `overridden` 的场景中必需,因为它的存在会替换推导出的回放。fixture 守卫拒绝缺失、不匹配和遗留的文件。两个命令均接受场景过滤器。
|
||||
|
||||
## 曾考虑的替代方案
|
||||
|
||||
- **手工编写的模型 chunk `llm.json`**:早期草案的做法。复用真实会话日志使 fixture 成为系统的真实产物而非手工构建的 mock,并兼作行为 golden。
|
||||
- **字节级 HTTP 录制库(Polly/nock/MSW)**:否决。与适配器耦合,处理流式 SSE(Server-Sent Events)时笨拙,且层级低于被测对象。
|
||||
- **从 `turn/end {kind:'error'|'aborted'}` 合成 throw/cancel 条目**:否决。这会将 `llm-replay` 耦合到 loop 内部的轮次关闭语义,且 `turn/end` 原因是有损的(无法区分抛出的 401 与 finish-error);显式的 `replay.override.json` 伴随文件是更清晰的 seam。
|
||||
|
||||
## 后果
|
||||
|
||||
新测试层为每个场景增加了经评审的 input、session、stdout、可选 override 和可选 workspace fixture。workspace 种子在录制和回放时都会被复制到临时 cwd。作为回报,该层通过真实的 Loader 和 tool 组合提供确定性的无密钥 transcript 覆盖。子进程、input、workspace、归一化和回放 harness 可以支持 ACP 之外的示例。
|
||||
|
||||
本 RFC 与[拟议的确定性 RFC](../../proposed/testing/2026-06-11-deterministic-and-stress-testing.md) 相关但不取代它:该提案的「通用回放 fixture」在每次测试后重新推导会话的*消息历史*(一项内部一致性不变式),而快照测试固定的是*外部协议输出*。二者互补:一个守护事件溯源不变式,另一个守护面向编辑器的契约。
|
||||
@@ -0,0 +1,6 @@
|
||||
# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write
|
||||
2026-06-19-real-api-e2e-ci.md: cc3e14e2d411dfa4cc68132f649f4ed26ab1de1d
|
||||
2026-06-19-real-api-e2e-ci.zh.md: 58a2a87541fd5b73b8272aa729f9dd09a429e4b4
|
||||
@@ -0,0 +1,102 @@
|
||||
# Agent Note: Real-API e2e in CI against the external DeepSeek API
|
||||
|
||||
Status: implemented
|
||||
|
||||
English | [中文](2026-06-19-real-api-e2e-ci.zh.md)
|
||||
|
||||
## Problem
|
||||
|
||||
The harness leans hard on real-API tests by policy: [docs/testing.md](../../../../docs/testing.md) argues that a no-key suite proves the plumbing but not the product, and the [ACP inject postmortem](../../../../docs/postmortem/0001-acp-default-export-drops-inject.md) is the standing proof — 178 keyless tests stayed green while a real editor session crashed instantly. The real-API e2e suite (`pnpm run test:e2e`, the `*.e2e.ts` files) exists precisely to close that gap: it drives the agent against the live DeepSeek API — real model calls, real bash tools, multi-turn, resume, ACP-over-stdio.
|
||||
|
||||
The default gate ([.github/workflows/ci.yml](../../../../.github/workflows/ci.yml)) is deliberately keyless: it carries no secret and runs for forks. `test:e2e` self-skips without a key (`describe.skipIf(!process.env.DEEPSEEK_API_KEY)`), so adding it there would report green without exercising the real suite. A separate secret-bearing workflow is required to make real-API coverage a merge signal.
|
||||
|
||||
This Agent Note records the decision to add a **second, secret-consuming workflow** that runs the real-API suite in CI, and — because introducing the first CI secret into a repo that may later go public is a security/isolation decision — the threat model it relies on and what changes when the repo becomes public.
|
||||
|
||||
## Decision
|
||||
|
||||
Add a dedicated workflow, [.github/workflows/e2e.yml](../../../../.github/workflows/e2e.yml), separate from ci.yml. It runs only `pnpm run test:e2e` against the external API using a repo secret, on trusted events, with a preflight that converts a missing secret into a loud failure instead of a false green. The keyless workflow remains separate so forkable quality gates and secret-consuming real-API gates keep different trigger and credential policies.
|
||||
|
||||
### A separate workflow, not a job in ci.yml
|
||||
|
||||
ci.yml's value is that it is keyless, forkable, and always-green: any contributor (including an outside fork) gets a complete keyless signal with no secret in the blast radius. Adding a secret-consuming job there would couple that always-green gate to credential availability and a different trigger policy. Keeping the secret-bearing work in its own file isolates the secret, trigger, and concurrency policy, and preserves ci.yml's property for forks. Different lifecycles → different files.
|
||||
|
||||
### Cost is not the constraint; reliability is
|
||||
|
||||
Internal inference cost is not the limiting constraint, so the workflow optimizes for coverage and signal. It runs every matching `*.e2e.ts` file on multiple triggers and every trusted PR, implementing the [docs/testing.md](../../../../docs/testing.md) with-key policy.
|
||||
|
||||
### Triggers: trusted events only
|
||||
|
||||
`workflow_dispatch` + `push` to `main`/`master` + nightly `schedule` (`17 0 * * *`, 08:17 Asia/Shanghai) + `pull_request`. Push gives a post-merge signal; schedule catches external-API drift; dispatch is the manual escape hatch; and trusted pull requests get a pre-merge gate. That pre-merge signal deliberately accepts the larger key-exposure surface described under § Security.
|
||||
|
||||
### The untrusted-PR gate
|
||||
|
||||
GitHub withholds repo secrets from two kinds of PR: those from **forks**, and **Dependabot** PRs (same-repo branch, so `head.repo.fork == false`, but secrets are still withheld). A job-level `if:` skips the whole job for both:
|
||||
|
||||
```
|
||||
github.event_name != 'pull_request'
|
||||
|| !(github.event.pull_request.head.repo.fork || github.event.pull_request.user.login == 'dependabot[bot]')
|
||||
```
|
||||
|
||||
The Dependabot clause keys on the PR **author** (`pull_request.user.login`), not `github.actor` (the run trigger): a maintainer who reopens or re-runs a Dependabot PR would make `github.actor` a human while the PR is still keyless, and an author-based test stays correct across that. A job skipped by a **job-level** `if:` reports as a *successful* check (unlike a workflow/trigger-level skip, which stays pending), so this workflow is safe to mark as a required status check if desired — a fork/Dependabot PR's skipped-but-green check does not block the merge.
|
||||
|
||||
The gate is a *clean-skip nicety*, not the secret's security boundary (see § Security — the boundary is GitHub's own fork-secret withholding under `pull_request`). Without the gate, forks still could not read the key; they would just hit a confusing preflight hard-fail and waste compute.
|
||||
|
||||
### Preflight: fail loud, never false-green
|
||||
|
||||
Because the job only runs on trusted events where the secret is expected, the preflight is an unconditional presence check: empty key → `exit 1` with a `::error::` annotation naming the secret to configure. This is the crux that makes a self-skipping suite safe to gate on. Without it, a deleted/renamed/misconfigured secret would make `test:e2e` skip every real suite and report all-green — a silent regression of the entire safety net. The guard turns "secret missing" from an invisible false pass into a visible failure. (Its correctness was verified live: the run before the secret existed failed at exactly this step.)
|
||||
|
||||
### Secret mapping and hygiene
|
||||
|
||||
The repo secret is named `DEEPSEEK_API_KEY_EXTERNAL`; it is mapped to the `DEEPSEEK_API_KEY` env var the adapters and tests read (`process.env.DEEPSEEK_API_KEY`). The distinct secret name documents intent (this is the *external* public-API key, not an internal-endpoint key) and lets an internal-endpoint key coexist later without collision. Hygiene choices, each defensive:
|
||||
|
||||
- **Step-scoped secret.** `DEEPSEEK_API_KEY` is set in the `env:` of only the preflight and e2e steps, never job-level — so checkout/setup-node/install never see it. A compromised install-time lifecycle script in a dependency cannot read a secret that isn't in its environment.
|
||||
- **`permissions: contents: read`.** The job only reads the repo to run tests; it needs no write scopes (no PR comments, no status writes), so the `GITHUB_TOKEN` is dropped to least privilege.
|
||||
- **`DEEPSEEK_BASE_URL` pinned** to `https://api.deepseek.com` on the e2e step. The adapter would default to this when unset ([packages/llm/llm-deepseek/src/index.ts](../../../../packages/llm/llm-deepseek/src/index.ts) `PUBLIC_BASE_URL`), but pinning is self-documenting and hermetic — a stray repo-root `.env` (which `vitest.e2e.config.ts` loads if present) cannot silently redirect the run to another endpoint.
|
||||
- **No secret echoed.** The preflight prints only `DEEPSEEK_API_KEY present.` — not the value or its length.
|
||||
|
||||
### Scope, runtime shape
|
||||
|
||||
The job runs only `test:e2e` on Node 24; keyless gates and version compatibility belong to the main CI workflow. Tests run unbuilt through the workspace paths map with a bounded configurable worker pool, per-test retries, and a job timeout. Superseded PR runs are cancelled, while push and scheduled runs complete for post-merge signal.
|
||||
|
||||
The DeepSeek native `web_search` probe is registered but skipped. The live Anthropic-compatible endpoint can return a successful response without structured source blocks, so its positive-source assertion is not a reliable merge signal; unit coverage still pins response parsing, but CI does not prove the live source-block wire shape.
|
||||
|
||||
## Security
|
||||
|
||||
The repository's first CI secret requires a recorded threat model because access differs between same-repository, fork, and Dependabot pull requests and changes when the repository becomes public.
|
||||
|
||||
### Who can reach the secret today (private repo)
|
||||
|
||||
- **No write access (fork PRs): cannot.** Two independent facts block it. First, the workflow uses `pull_request`, **not** `pull_request_target` — GitHub does not pass repo secrets to fork-PR runs of `pull_request`, so `secrets.DEEPSEEK_API_KEY_EXTERNAL` resolves to empty on a fork runner. Second, the `if:` gate skips fork PRs entirely. The withholding is the real boundary; the gate is defense-in-depth and UX.
|
||||
- **Write (push) access: can.** A same-repo branch PR receives secrets, so a write-access author could modify test code (or an install lifecycle script, or the workflow YAML on their branch) to exfiltrate the key. This is **inherent to GitHub Actions, not introduced here**: anyone with push access to any repo can already exfiltrate any of its Actions secrets by authoring a workflow. Write access ⇒ secret access, always. The mitigation lives in who is granted write and in branch protection, not in this file.
|
||||
|
||||
So "everyone who could open a PR can steal it" is false: only the write-access set can, and that set could already steal any secret the repo holds.
|
||||
|
||||
### The residual exposure the `pull_request` trigger adds
|
||||
|
||||
Because PR runs are enabled, the key is handed to **the code on a write-access author's PR branch** before merge. This is a larger surface than `push` + `schedule` + `workflow_dispatch`, accepted for a pre-merge signal within the trusted write set. If that calculus changes, drop the `pull_request` trigger while retaining post-merge, nightly, and on-demand coverage.
|
||||
|
||||
### What changes when the repo goes public
|
||||
|
||||
The secret stays protected from the public **through this workflow**: `pull_request` behaves identically on a public repo — fork PRs (now openable by anyone) still receive no secret, and on public repos GitHub additionally gates fork-PR runs behind maintainer approval, where even an approved run gets no secret (approving the run is not the same as handing over the key). The write-access set is unchanged by visibility, so the insider reality is also unchanged.
|
||||
|
||||
What gets worse is the *surrounding* model, and these are the things to address before flipping visibility:
|
||||
|
||||
- **Logs become world-readable.** A careless secret echo that today leaks to org members would leak to the entire internet and be scraped within minutes. Secret-handling discipline (no value/length echoes — already done) matters far more.
|
||||
- **The `pull_request_target` footgun becomes catastrophic.** If anyone ever "fixes" PR runs by switching the trigger to `pull_request_target`, the workflow would run untrusted fork code in the base-repo context **with** secrets — a full key-leak vector. This is benign-ish on a private repo and disastrous on a public one. A `SECURITY —` comment on the trigger in e2e.yml forbids the change and points here.
|
||||
- **Rotate on flip.** The key lived in a private repo's CI; treat going-public as "assume exposed" and rotate `DEEPSEEK_API_KEY_EXTERNAL` at that moment.
|
||||
- **Settle the secret behind controls.** Confirm Settings → Actions → *"Send secrets to workflows from fork pull requests"* stays **off** (the one setting that would actually break the fork boundary), and consider moving the key into a GitHub **Environment** with required reviewers so even merged code uses it only under controlled conditions and rotation has a single home.
|
||||
|
||||
None of these require changing the workflow to go public; they are operational steps plus the already-added `pull_request_target` guard comment.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
- **A secret-consuming job inside ci.yml** — rejected: it would couple the keyless, forkable, always-green gate to credential availability and a different trigger/concurrency policy; different lifecycles, different files.
|
||||
- **Omitting the `pull_request` trigger** (the smaller key-exposure surface) — rejected for the pre-merge signal; the Security section carries the accepted exposure analysis.
|
||||
|
||||
## Consequences
|
||||
|
||||
A second CI workflow and the first repo secret to maintain. The real-API suite now gates merges (pre-merge on trusted PRs, post-merge on the main branch) and runs nightly, so a real break in the agent's interaction with the external API surfaces in CI rather than only in a developer's local run — at the cost of real (but internally free) API calls on every trusted PR and merge. The preflight makes secret misconfiguration self-announcing instead of silently disabling the net.
|
||||
|
||||
The design carries a documented constraint surface: the `pull_request` trigger's key-exposure tradeoff (drop it to harden), the `if:` gate's dependence on the author-based Dependabot test, and the hard prohibition on `pull_request_target`. The going-public checklist above is the operational companion — this Agent Note is the place a future maintainer should re-read before changing the trigger set or flipping repo visibility, rather than re-deriving the fork/secret model from scratch.
|
||||
|
||||
The scheduled trigger auto-disables after 60 days of repo inactivity (a GitHub behavior); push/PR/dispatch are backstops, and an active monorepo will not hit it. Runner egress to `https://api.deepseek.com` is assumed — GitHub-hosted `ubuntu-latest` has it; an egress-restricted self-hosted runner would need connectivity confirmed before relying on the nightly.
|
||||
@@ -0,0 +1,100 @@
|
||||
# RFC: 在 CI 中对外部 DeepSeek API 运行真实 API e2e 测试
|
||||
|
||||
Status: implemented
|
||||
|
||||
[English](2026-06-19-real-api-e2e-ci.md) | 中文
|
||||
|
||||
## 问题
|
||||
|
||||
按照策略,harness 高度依赖真实 API 测试:[docs/testing.md](../../../testing.md) 论证了无密钥套件只能验证管道连通性而非产品本身,[ACP inject 事后分析](../../../postmortem/0001-acp-default-export-drops-inject.md)是现成的证据——178 个无密钥测试全绿,而真实编辑器会话一启动就崩溃。真实 API e2e 套件(`pnpm run test:e2e`,即 `*.e2e.ts` 文件)正是为弥合这一差距而存在的:它驱动 agent(智能体)对接实时 DeepSeek API——真实模型调用、真实 bash 工具、多轮次对话、恢复、ACP-over-stdio。
|
||||
|
||||
默认门禁([.github/workflows/ci.yml](../../../../.github/workflows/ci.yml))刻意无密钥:不携带 secret,可供 fork 运行。`test:e2e` 在无密钥时自动跳过(`describe.skipIf(!process.env.DEEPSEEK_API_KEY)`),因此将其加入该工作流只会报绿而不会真正执行真实套件。要让真实 API 覆盖率成为合并信号,需要一个独立的、携带 secret 的工作流。
|
||||
|
||||
本 RFC 记录的决策是:添加一个**第二个、消费 secret 的工作流**来在 CI 中运行真实 API 套件。由于这是向一个未来可能公开的仓库引入首个 CI secret,属于安全/隔离决策,本文同时记录其依赖的威胁模型以及仓库公开后的变化。
|
||||
|
||||
## 决策
|
||||
|
||||
添加一个专用工作流 [.github/workflows/e2e.yml](../../../../.github/workflows/e2e.yml),与 ci.yml 分离。它仅使用 repo secret 对外部 API 运行 `pnpm run test:e2e`,仅在可信事件上触发,并带有一个 preflight 检查:将缺失的 secret 转化为明确的失败而非虚假的绿色。无密钥工作流保持独立,使可 fork 的质量门禁与消费 secret 的真实 API 门禁各自拥有不同的触发和凭证策略。
|
||||
|
||||
### 独立工作流,而非 ci.yml 中的一个 job
|
||||
|
||||
ci.yml 的价值在于它无密钥、可 fork、始终为绿:任何贡献者(包括外部 fork)都能获得完整的无密钥信号,secret 不在爆炸半径内。在其中添加消费 secret 的 job 会将这个始终为绿的门禁耦合到凭证可用性和不同的触发策略上。将携带 secret 的工作放在独立文件中,隔离了 secret、触发和并发策略,并为 fork 保留了 ci.yml 的特性。不同的生命周期→不同的文件。
|
||||
|
||||
### 约束不是成本,而是可靠性
|
||||
|
||||
内部推理(inference)成本不是限制因素,因此工作流以覆盖率和信号为优化目标。它在多个触发条件和每个可信 PR(Pull Request)上运行所有匹配的 `*.e2e.ts` 文件,落实 [docs/testing.md](../../../testing.md) 的有密钥策略。
|
||||
|
||||
### 触发条件:仅限可信事件
|
||||
|
||||
`workflow_dispatch` + `push` 到 `main`/`master` + 每夜 `schedule`(`17 0 * * *`,即北京时间 08:17)+ `pull_request`。push 提供合并后信号;schedule 捕捉外部 API 漂移;dispatch 是手动逃生通道;可信 pull request 获得合并前门禁。该合并前信号有意接受 § 安全性中描述的更大密钥暴露面。
|
||||
|
||||
### 不可信 PR 的门禁
|
||||
|
||||
GitHub 对两类 PR 扣留 repo secret:来自 **fork** 的 PR,以及 **Dependabot** PR(同仓库分支,`head.repo.fork == false`,但 secret 仍被扣留)。一个 job 级 `if:` 对两者都跳过整个 job:
|
||||
|
||||
```
|
||||
github.event_name != 'pull_request'
|
||||
|| !(github.event.pull_request.head.repo.fork || github.event.pull_request.user.login == 'dependabot[bot]')
|
||||
```
|
||||
|
||||
Dependabot 子句基于 PR **作者**(`pull_request.user.login`)而非 `github.actor`(运行触发者):维护者重新打开或重跑 Dependabot PR 时,`github.actor` 会变成人类,但该 PR 仍然无密钥;基于作者的判断在这种情况下依然正确。被 **job 级** `if:` 跳过的 job 报告为*成功*检查(不同于工作流/触发级跳过会保持 pending),因此如果需要将此工作流标记为 required status check 也是安全的——fork/Dependabot PR 的跳过但绿色的检查不会阻塞合并。
|
||||
|
||||
该门禁是一个*干净跳过的便利措施*,而非 secret 的安全边界(见 § 安全性——边界是 GitHub 自身在 `pull_request` 下对 fork 的 secret 扣留机制)。没有该门禁,fork 仍然无法读取密钥;只是会遇到令人困惑的 preflight 硬失败并浪费计算资源。
|
||||
|
||||
### Preflight:大声失败,绝不虚假为绿
|
||||
|
||||
由于 job 仅在 secret 应当存在的可信事件上运行,preflight 是一个无条件的存在性检查:密钥为空→`exit 1` 并附带 `::error::` 注解指明需要配置的 secret 名称。这是让自跳过套件可以安全地作为门禁的关键。没有它,被删除/重命名/错误配置的 secret 会让 `test:e2e` 跳过所有真实套件并报告全绿——整个安全网的静默退化。该守卫将「secret 缺失」从不可见的虚假通过转化为可见的失败。(其正确性已在实际中验证:secret 存在之前的运行恰好在此步骤失败。)
|
||||
|
||||
### Secret 映射与卫生
|
||||
|
||||
repo secret 命名为 `DEEPSEEK_API_KEY_EXTERNAL`;映射到适配器和测试读取的 `DEEPSEEK_API_KEY` 环境变量(`process.env.DEEPSEEK_API_KEY`)。独立的 secret 名称记录了意图(这是*外部*公开 API 密钥,不是内部端点密钥),并允许内部端点密钥日后无冲突地共存。以下卫生选择均为防御性设计:
|
||||
|
||||
- **Step 级 secret。** `DEEPSEEK_API_KEY` 仅在 preflight 和 e2e 步骤的 `env:` 中设置,从不在 job 级设置——因此 checkout/setup-node/install 永远看不到它。依赖中被入侵的安装时生命周期脚本无法读取不在其环境中的 secret。
|
||||
- **`permissions: contents: read`。** job 仅读取仓库以运行测试;不需要写权限(无 PR 评论、无 status 写入),因此 `GITHUB_TOKEN` 降至最小权限。
|
||||
- **`DEEPSEEK_BASE_URL` 固定**为 e2e 步骤上的 `https://api.deepseek.com`。适配器在未设置时会默认使用此值([packages/llm/llm-deepseek/src/index.ts](../../../../packages/llm/llm-deepseek/src/index.ts) `PUBLIC_BASE_URL`),但显式固定具有自文档性和密封性——仓库根目录的 `.env`(`vitest.e2e.config.ts` 存在时会加载)无法静默地将运行重定向到其他端点。
|
||||
- **不回显 secret。** preflight 仅打印 `DEEPSEEK_API_KEY present.`——不打印值或长度。
|
||||
|
||||
### 范围与运行时形态
|
||||
|
||||
job 仅在 Node 24 上运行 `test:e2e`;无密钥门禁和版本兼容性属于主 CI 工作流。测试通过 workspace paths 映射以未构建形式运行,使用有界的可配置 worker 池、逐测试重试和 job 超时。被取代的 PR 运行会被取消,而 push 和 schedule 运行完整执行以提供合并后信号。
|
||||
|
||||
## 安全性
|
||||
|
||||
仓库的首个 CI secret 需要一份记录在案的威胁模型,因为同仓库 PR、fork PR 和 Dependabot PR 的访问权限各不相同,且仓库公开后会发生变化。
|
||||
|
||||
### 当前谁能触及 secret(私有仓库)
|
||||
|
||||
- **无写权限(fork PR):不能。** 两个独立事实阻止了它。第一,工作流使用 `pull_request` 而**非** `pull_request_target`——GitHub 不会将 repo secret 传递给 fork PR 的 `pull_request` 运行,因此 `secrets.DEEPSEEK_API_KEY_EXTERNAL` 在 fork runner 上解析为空。第二,`if:` 门禁完全跳过 fork PR。secret 扣留是真正的边界;门禁是纵深防御和用户体验。
|
||||
- **有写(push)权限:能。** 同仓库分支 PR 会收到 secret,因此有写权限的作者可以修改测试代码(或安装生命周期脚本,或其分支上的工作流 YAML)来窃取密钥。这**是 GitHub Actions 的固有特性,并非本文引入的**:任何对任何仓库有 push 权限的人都可以通过编写工作流来窃取该仓库的任何 Actions secret。写权限⇒secret 访问权,始终如此。缓解措施在于谁被授予写权限以及分支保护,而非本文件。
|
||||
|
||||
因此「任何能开 PR 的人都能窃取它」是错误的:只有写权限集合内的人能,而这些人本来就能窃取仓库持有的任何 secret。
|
||||
|
||||
### `pull_request` 触发器增加的残余暴露面
|
||||
|
||||
由于启用了 PR 运行,密钥会在合并前被交给**写权限作者 PR 分支上的代码**。这比 `push` + `schedule` + `workflow_dispatch` 的暴露面更大,为在可信写权限集合内获得合并前信号而接受。如果这一权衡发生变化,可移除 `pull_request` 触发器,同时保留合并后、每夜和按需覆盖。
|
||||
|
||||
### 仓库公开后的变化
|
||||
|
||||
**通过本工作流**,secret 对公众仍然受保护:`pull_request` 在公开仓库上行为一致——fork PR(现在任何人都能开)仍然收不到 secret,且在公开仓库上 GitHub 额外要求维护者批准 fork PR 运行,即使批准后运行也不会获得 secret(批准运行不等于交出密钥)。写权限集合不因可见性改变而改变,因此内部人员的现实也不变。
|
||||
|
||||
变差的是*周边*模型,以下是翻转可见性之前需要处理的事项:
|
||||
|
||||
- **日志变为全球可读。** 今天泄露给组织成员的粗心 secret 回显,公开后会泄露给整个互联网并在数分钟内被爬取。secret 处理纪律(不回显值/长度——已做到)的重要性大幅提升。
|
||||
- **`pull_request_target` 陷阱变为灾难性的。** 如果有人为了「修复」PR 运行而将触发器切换为 `pull_request_target`,工作流将在 base-repo 上下文中运行不可信的 fork 代码并**携带** secret——完整的密钥泄露向量。在私有仓库中这勉强无害,在公开仓库中则是灾难。e2e.yml 中触发器上的 `SECURITY —` 注释禁止此更改并指向本文。
|
||||
- **翻转时轮换密钥。** 密钥曾存在于私有仓库的 CI 中;将公开视为「假定已暴露」,在那一刻轮换 `DEEPSEEK_API_KEY_EXTERNAL`。
|
||||
- **将 secret 置于控制之下。** 确认 Settings → Actions → *"Send secrets to workflows from fork pull requests"* 保持**关闭**(这是唯一真正会打破 fork 边界的设置),并考虑将密钥移入带有 required reviewers 的 GitHub **Environment**,使即使已合并的代码也只在受控条件下使用它,且轮换有单一归属。
|
||||
|
||||
以上均不需要修改工作流即可公开;它们是运维步骤加上已添加的 `pull_request_target` 守卫注释。
|
||||
|
||||
## 曾考虑的替代方案
|
||||
|
||||
- **在 ci.yml 中添加消费 secret 的 job**:否决。会将无密钥、可 fork、始终为绿的门禁耦合到凭证可用性和不同的触发/并发策略上;不同的生命周期,不同的文件。
|
||||
- **省略 `pull_request` 触发器**(更小的密钥暴露面):为获得合并前信号而否决;安全性章节承载了已接受的暴露分析。
|
||||
|
||||
## 后果
|
||||
|
||||
新增一个 CI 工作流和仓库的首个需要维护的 secret。真实 API 套件现在作为合并门禁(可信 PR 上的合并前门禁、主分支上的合并后门禁)并每夜运行,因此 agent 与外部 API 交互中的真实故障会在 CI 中浮现,而非仅在开发者的本地运行中出现——代价是每个可信 PR 和合并都会产生真实的(但内部免费的)API 调用。preflight 使 secret 配置错误变为自我通告而非静默禁用安全网。
|
||||
|
||||
本设计携带一个记录在案的约束面:`pull_request` 触发器的密钥暴露权衡(移除以加固)、`if:` 门禁对基于作者的 Dependabot 判断的依赖,以及对 `pull_request_target` 的硬性禁止。上述公开清单是运维伴侣——本 RFC 是未来维护者在更改触发器集合或翻转仓库可见性之前应重读的地方,而非从头重新推导 fork/secret 模型。
|
||||
|
||||
schedule 触发器在仓库不活跃 60 天后会自动禁用(GitHub 行为);push/PR/dispatch 是后备,活跃的 monorepo 不会触及此限制。假设 runner 对 `https://api.deepseek.com` 有出站连通性——GitHub 托管的 `ubuntu-latest` 具备此条件;受出站限制的自托管 runner 需要在依赖每夜运行之前确认连通性。
|
||||
+6
@@ -0,0 +1,6 @@
|
||||
# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write
|
||||
2026-06-20-remove-redundant-snapshot-log-goldens.md: 18d0a4491eb10a3b4dc56d3d63285c219ba6a00a
|
||||
2026-06-20-remove-redundant-snapshot-log-goldens.zh.md: 5675c69862b6052ed3f3e4710461cc1478b9fa7d
|
||||
+35
@@ -0,0 +1,35 @@
|
||||
# Agent Note: Use `session.jsonl` as the only snapshot session-log artifact
|
||||
|
||||
Status: implemented
|
||||
|
||||
## Problem
|
||||
|
||||
Model-driving ACP snapshot scenarios ship both `session.jsonl` and `session.expected.jsonl`. For normal recorded scenarios, `session.jsonl` is the replay fixture harvested from a real run, and the replay test normalizes the newly persisted log and compares it to `session.expected.jsonl`. In the current fixtures, the two normalized logs are identical for ordinary recorded scenarios.
|
||||
|
||||
Authored override scenarios (`error-finish`, `cancel`) currently use `replay.override.json` to drive model behavior and keep `session.jsonl` as a minimal dummy fixture, while `session.expected.jsonl` holds the expected persisted log. The override file is a JSON array of `ReplayEntry` objects: `{ "kind": "chunks", "chunks": StreamChunk[] }`, `{ "kind": "throw", "chunks": StreamChunk[], "message": string, "code": string }`, or `{ "kind": "hang" }`. That split is also unnecessary: when an override sidecar exists, `llm-replay` replaces the derived script and does not need `session.jsonl` for model chunks, so `session.jsonl` can still be the expected session-log artifact for the scenario.
|
||||
|
||||
## Decision
|
||||
|
||||
The `session.expected.jsonl` concept is removed entirely. Every scenario has at most one committed session-log artifact, `session.jsonl`:
|
||||
|
||||
- For recorded scenarios, `session.jsonl` remains the raw harvested log. Replay still derives model chunks from it, and the snapshot test compares the replay run's normalized persisted log against normalized `session.jsonl`.
|
||||
- For authored override scenarios, `replay.override.json` drives model behavior and `session.jsonl` holds the expected produced session log. The replay adapter ignores the fixture for model chunks when the override exists, so the same file can be the expected log without affecting replay behavior.
|
||||
- For no-model scenarios, `session.jsonl` can stay as the minimal fixture needed to boot `llm-replay`; no session-log comparison is needed unless the scenario creates a persisted session.
|
||||
|
||||
Stdout expected outputs remain unchanged; they are the editor-facing projection and are not redundant with the session fixture.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
**Normalizing both sides against a shared (replay-run) context** — rejected: `normalizeSessionLog` scrubs cwd by exact string match, so the fixture's recorded cwd would survive unscrubbed and every compare would fail. Each side normalizes against its own header-derived context — the implementation note below carries the mechanics.
|
||||
|
||||
## Verification
|
||||
|
||||
`session.expected.jsonl` appears nowhere in the snapshot harness, fixtures, orphan guards, or docs; the snapshot test derives the expected session log from `session.jsonl` for every model scenario; authored sidecar scenarios commit their expected produced log as `session.jsonl` with `replay.override.json` as the model-behavior override; and the orphan-fixture guards know which files each scenario kind requires. The [ACP snapshot tests Agent Note](2026-06-19-acp-snapshot-tests.md) describes the reduced fixture set.
|
||||
|
||||
## Consequences
|
||||
|
||||
Reviewers lose one artifact name that made the expected persisted log visually separate from the replay fixture. The stdout expected output still protects the editor transcript, and comparing replay output to `session.jsonl` preserves the loop/persistence regression check without duplicating files.
|
||||
|
||||
## Implementation note
|
||||
|
||||
Each side is normalized against its own header values because recording and replay have different ids, paths, and timestamps. `fixtureContext()` derives the fixture context from its header, making already-normalized fixtures idempotent. Session logs use plain equality rather than file-snapshot updates, so comparison never rewrites fixtures.
|
||||
+37
@@ -0,0 +1,37 @@
|
||||
# RFC: 使用 `session.jsonl` 作为唯一的快照会话日志产物
|
||||
|
||||
Status: implemented
|
||||
|
||||
[English](2026-06-20-remove-redundant-snapshot-log-goldens.md) | 中文
|
||||
|
||||
## 问题
|
||||
|
||||
模型驱动的 ACP(Agent Client Protocol)快照场景同时包含 `session.jsonl` 和 `session.golden.jsonl`。对于普通录制场景,`session.jsonl` 是从真实运行中采集的回放 fixture(测试前置数据),回放测试对新持久化的日志做归一化后与 `session.golden.jsonl` 比较。在当前 fixture 中,普通录制场景的归一化录制日志与归一化 golden 完全一致。
|
||||
|
||||
手工编写的覆盖场景(`error-finish`、`cancel`)目前使用 `replay.override.json` 驱动模型行为,并保留 `session.jsonl` 作为最小占位 fixture,而 `session.golden.jsonl` 存放预期的持久化日志。覆盖文件是一个 `ReplayEntry` 对象的 JSON 数组:`{ "kind": "chunks", "chunks": StreamChunk[] }`、`{ "kind": "throw", "chunks": StreamChunk[], "message": string, "code": string, "status"?: number }` 或 `{ "kind": "hang" }`。这种拆分同样是多余的:当覆盖 sidecar 存在时,`llm-replay` 会替换派生脚本,不需要从 `session.jsonl` 获取模型分片,因此 `session.jsonl` 仍可作为该场景的预期会话日志产物。
|
||||
|
||||
## 决策
|
||||
|
||||
彻底移除 `session.golden.jsonl` 概念。每个场景最多只有一个提交到仓库的会话日志产物,即 `session.jsonl`:
|
||||
|
||||
- 对于录制场景,`session.jsonl` 仍是原始采集的日志。回放仍从中派生模型分片,快照测试将回放运行归一化后的持久化日志与归一化后的 `session.jsonl` 进行比较。
|
||||
- 对于手工编写的覆盖场景,`replay.override.json` 驱动模型行为,`session.jsonl` 存放预期产出的会话日志。当覆盖文件存在时,回放适配器不从 fixture 获取模型分片,因此同一个文件既可作为预期日志,又不影响回放行为。
|
||||
- 对于无模型场景,`session.jsonl` 可保留为引导 `llm-replay` 所需的最小 fixture;除非场景创建了持久化会话,否则无需进行会话日志比较。
|
||||
|
||||
stdout golden 保持不变;它们是面向编辑器的投影,与会话 fixture 不构成冗余。
|
||||
|
||||
## 曾考虑的替代方案
|
||||
|
||||
**对两侧基于共享的(回放运行)上下文做归一化**:否决。`normalizeSessionLog` 通过精确字符串匹配擦除 cwd,因此 fixture 中录制的 cwd 不会被擦除,每次比较都会失败。两侧各自基于自身 header 派生的上下文做归一化——下方的实现说明描述了具体机制。
|
||||
|
||||
## 验证
|
||||
|
||||
`session.golden.jsonl` 在快照 harness、fixture、遗留文件守卫和文档中均不再出现;快照测试对每个模型场景都从 `session.jsonl` 派生预期会话日志;手工编写的 sidecar 场景将预期产出的日志作为 `session.jsonl` 提交,并以 `replay.override.json` 作为模型行为覆盖;遗留 fixture 守卫知道每种场景类型需要哪些文件。[ACP 快照测试 RFC](../../implemented/testing/2026-06-19-acp-snapshot-tests.md) 描述了精简后的 fixture 集合。
|
||||
|
||||
## 后果
|
||||
|
||||
评审者失去了一个让预期持久化日志在视觉上与回放 fixture 分离的产物名称。stdout golden 仍保护编辑器 transcript(文本记录),将回放输出与 `session.jsonl` 比较则在不重复文件的前提下保留了循环/持久化的回归检查。
|
||||
|
||||
## 实现说明
|
||||
|
||||
两侧各自基于自身 header 值做归一化,因为录制与回放具有不同的 id、路径和时间戳。`fixtureContext()` 从 fixture 的 header 派生上下文,使已归一化的 fixture 具有幂等性。会话日志使用普通相等比较而非文件快照更新,因此比较过程不会改写 fixture。
|
||||
@@ -0,0 +1,6 @@
|
||||
# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write
|
||||
2026-06-22-fork-child-replay-seed-boundary.md: 28ce76309da2dca7076dd11211229a0631d11db3
|
||||
2026-06-22-fork-child-replay-seed-boundary.zh.md: 92b60589b4cfe9668a1542405693cec8d29eceaf
|
||||
@@ -0,0 +1,49 @@
|
||||
# Agent Note: Persist the seed boundary so fork-child replay routes correctly
|
||||
|
||||
Status: implemented
|
||||
|
||||
English | [中文](2026-06-22-fork-child-replay-seed-boundary.zh.md)
|
||||
|
||||
## Problem
|
||||
|
||||
The [per-session snapshot replay Agent Note](2026-06-22-subagent-snapshot-replay.md) made the snapshot tier express a nested-agent shape: a parent plus one recorded log per in-process subagent, each replayed as its own script keyed by calling session. It noted (§ Scope, final bullet) that a fork snapshot was "a trivial future addition, not a gap in the keying." That was wrong about a fork child specifically — not the keying, but the *script derivation*.
|
||||
|
||||
A subagent script is derived from a recorded session log by [`deriveReplayScript`](../../../../packages/support/llm-replay): it groups the log's `assistant/chunk` events by `(turn, step)` into one replay entry per `stream()` call. This is correct for a **spawn** child, whose log contains only its own model calls.
|
||||
|
||||
A **fork** child is different. The fork backend seeds the child session with a *balanced completed-turn prefix of the parent's log* ([`dsh-subagent-inprocess`](../../../../packages/subagent/subagent-inprocess)), and that seed becomes the child session's persisted `log` (`Session`'s constructor copies the seed into `this.log`). So a fork child's `.jsonl` begins with the **parent's** events — including the parent's `assistant/chunk` events — and only then carries the child's own turn.
|
||||
|
||||
Deriving the child script from the whole fork-child log therefore replays the **parent's** recorded responses as the **child's** model calls: the live fork child's first `stream()` would receive the parent's first recorded chunk sequence instead of its own. The recorded scenarios are all spawn today, so this never fired — but a fork snapshot would have mis-routed silently, exactly the class of bug the snapshot tier exists to catch.
|
||||
|
||||
## Decision
|
||||
|
||||
Record where a session's **inherited** prefix ends, persist it, and have the replay harness derive a child's script from its **own** events only.
|
||||
|
||||
### 1. `seedLength` on the session header
|
||||
|
||||
`SessionHeader` gains an optional `seedLength: number` — how many leading events were inherited via a seed rather than produced by this session. The fork backend stamps it (= the seeded-prefix length) when it creates the child; a fresh spawn leaves it absent (≡ 0). It is threaded through `CreateSessionOptions.meta` (and `CreateAgentOptions.meta`), set in `SessionStore.prepare`.
|
||||
|
||||
`seedLength` is **explicit**, never inferred from `seed.length`. A reconstruction (resume/load) seeds the session with its WHOLE stored log, so `seed.length` there is the full length, not the original boundary — the resume path passes the persisted `seedLength` back from the loaded header instead. (Same shape as `createdAt`, which is also explicitly preserved on reconstruction rather than re-defaulted to now.)
|
||||
|
||||
### 2. Both persistence backends round-trip it
|
||||
|
||||
- **JSONL**: a `seedLength` field on the header line (`toHeaderLine`/`fromHeaderLine`).
|
||||
- **SQLite**: a `seed_length` column on the `sessions` table.
|
||||
|
||||
The SQLite layout containing `seed_length`, `source_event_seqs`, and `surface_op` is schema version 4. Earlier version 3 layouts were ambiguous, so every non-current `user_version` is rejected without migration under the pre-release policy.
|
||||
|
||||
### 3. Replay derives a child script after the boundary
|
||||
|
||||
`dsh-llm-replay`'s `parseSessionHeader` now also reads `seedLength` (absent ⇒ 0), and `loadSessionScripts` derives a child's entries from `parseSessionLog(text).slice(seedLength)` — the events at or after the boundary, i.e. the child's own model calls. For a spawn child `seedLength` is 0 and this is a no-op, so spawn scenarios are byte-for-byte unchanged.
|
||||
|
||||
This closes the routing correctness gap, and two recorded fork scenarios exercise it end to end — see [Record fork and mixed spawn+fork snapshot scenarios](2026-06-22-fork-snapshot-scenarios.md).
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
- **Derive the boundary heuristically in `llm-replay`** (the seeded prefix is contiguous parent events ending at the last `turn/end` before the child's first `user/message`). Rejected: a brittle heuristic in the test harness that re-derives a fact the producer already knows. Persisting the boundary at its source (the fork backend) is the "explicit > implicit at package seams" rule applied across the persistence boundary — the reader of a child fixture never has to reconstruct where the inheritance ended.
|
||||
- **Pin the format version instead of bumping** (the `SESSION_FORMAT_VERSION = 0` "unstable" stance the event log uses). Rejected for the SQLite *table* layout: `SCHEMA_VERSION` is the monotonic bump-and-reject knob (a small enumerable set of revisions worth telling apart), distinct from the event-vocabulary `version`. Adding a column is precisely the breaking table change it versions, so it bumps.
|
||||
|
||||
## Consequences
|
||||
|
||||
- A new persisted header field across core + both backends; the core-data-structures catalog (`persistence.md`) is updated in the same change (its `SessionHeader` / `CreateSessionOptions` `type-equiv` blocks).
|
||||
- Existing SQLite databases at schema v2 are rejected on open (no user data pre-release).
|
||||
- Spawn replay is unchanged (`seedLength` 0). Fork replay now routes a child to its own script; covered by a regression in `llm-replay`'s tests (a child fixture whose seeded prefix carries a parent chunk — the derived child script must exclude it, proven red without the slice) and a persistence round-trip test (both backends, via the shared coordinator contract).
|
||||
@@ -0,0 +1,49 @@
|
||||
# RFC: 持久化 seed 边界以确保 fork 子会话回放正确路由
|
||||
|
||||
Status: implemented
|
||||
|
||||
[English](2026-06-22-fork-child-replay-seed-boundary.md) | 中文
|
||||
|
||||
## 问题
|
||||
|
||||
[逐会话快照回放 RFC](2026-06-22-subagent-snapshot-replay.md) 让快照层表达了嵌套 agent(智能体)的形状:一个父会话加上每个进程内 subagent 各一份已录制的日志,各自作为独立脚本回放、以调用方会话为键。该 RFC 指出(§ Scope 末尾条目)fork 快照是「一个平凡的后续补充,不是键控方案的缺口」。这对 fork 子会话而言是错的——问题不在键控,而在*脚本推导*。
|
||||
|
||||
subagent 脚本由 [`deriveReplayScript`](../../../../packages/support/llm-replay) 从已录制的会话日志推导:它按 `(turn, step)` 对日志中的 `assistant/chunk` 事件分组,每次 `stream()` 调用对应一条回放条目。对 **spawn** 子会话而言这是正确的,因为其日志只包含自身的模型调用。
|
||||
|
||||
**fork** 子会话不同。fork 后端用*父日志的一段平衡的已完成轮次前缀*([`dsh-subagent-inprocess`](../../../../packages/subagent/subagent-inprocess))来播种子会话,而该 seed 会成为子会话持久化的 `log`(`Session` 构造函数将 seed 复制进 `this.log`)。因此 fork 子会话的 `.jsonl` 以**父会话**的事件开头——包括父会话的 `assistant/chunk` 事件——之后才是子会话自身的轮次。
|
||||
|
||||
从 fork 子会话的完整日志推导脚本,会把**父会话**的已录制响应当作**子会话**的模型调用来回放:实际运行的 fork 子会话第一次调用 `stream()` 时,会收到父会话的第一段 chunk 序列而非自身的。目前已录制的场景全部是 spawn,所以这从未触发——但 fork 快照会静默地错误路由,恰好属于快照层存在的意义所要捕获的那类 bug。
|
||||
|
||||
## 决策
|
||||
|
||||
记录会话**继承**前缀的结束位置,将其持久化,并让回放 harness 仅从子会话**自身**的事件推导脚本。
|
||||
|
||||
### 1. 会话头部的 `seedLength`
|
||||
|
||||
`SessionHeader` 新增可选字段 `seedLength: number`——表示有多少前导事件是通过 seed 继承而来、而非本会话产生的。fork 后端在创建子会话时设置它(= 播种前缀的长度);全新的 spawn 子会话不设置(等同于 0)。它通过 `CreateSessionOptions.meta`(及 `CreateAgentOptions.meta`)传递,在 `SessionStore.prepare` 中设置。
|
||||
|
||||
`seedLength` 是**显式**的,绝不从 `seed.length` 推断。重建(resume/load)时用会话的完整已存储日志作为 seed,此时 `seed.length` 是全长而非原始边界——resume 路径改为从加载的 header 中取回持久化的 `seedLength`。(形状与 `createdAt` 相同:重建时显式保留,而非重新默认为当前时间。)
|
||||
|
||||
### 2. 两个持久化后端均完整往返
|
||||
|
||||
- **JSONL**:header 行上的 `seedLength` 字段(`toHeaderLine`/`fromHeaderLine`)。
|
||||
- **SQLite**:`sessions` 表上的 `seed_length` 列。
|
||||
|
||||
包含 `seed_length`、`source_event_seqs` 和 `surface_op` 的 SQLite 布局为 schema version 4。更早的 version 3 布局存在歧义,因此在预发布策略下,所有非当前 `user_version` 均直接拒绝,不做迁移。
|
||||
|
||||
### 3. 回放从边界之后推导子会话脚本
|
||||
|
||||
`dsh-llm-replay` 的 `parseSessionHeader` 现在也读取 `seedLength`(缺失则为 0),`loadSessionScripts` 从 `parseSessionLog(text).slice(seedLength)` 推导子会话条目——即边界及之后的事件,也就是子会话自身的模型调用。对 spawn 子会话而言 `seedLength` 为 0,此操作是空操作,spawn 场景逐字节不变。
|
||||
|
||||
这关闭了路由正确性的缺口,两个已录制的 fork 场景对其进行端到端验证——见 [Record fork and mixed spawn+fork snapshot scenarios](2026-06-22-fork-snapshot-scenarios.md)。
|
||||
|
||||
## 曾考虑的替代方案
|
||||
|
||||
- **在 `llm-replay` 中启发式推导边界**(播种前缀是连续的父事件,止于子会话第一条 `user/message` 之前的最后一个 `turn/end`)。否决:在测试 harness 中用脆弱的启发式重新推导一个生产者已经知道的事实。在源头(fork 后端)持久化边界,是「在包(package)seam 处显式优于隐式」这条规则跨越持久化边界的应用——子会话 fixture(测试前置数据)的读取者永远不需要重建继承在哪里结束。
|
||||
- **固定格式版本而不递增**(事件日志使用的 `SESSION_FORMAT_VERSION = 0`「不稳定」姿态)。对 SQLite *表*布局否决:`SCHEMA_VERSION` 是单调递增并拒绝旧版的旋钮(一组小的、值得区分的修订),与事件词汇表的 `version` 不同。新增列正是它所版本化的那种破坏性表变更,因此需要递增。
|
||||
|
||||
## 后果
|
||||
|
||||
- core 与两个后端新增一个持久化 header 字段;核心数据结构目录(`persistence.md`)在同一变更中更新(其 `SessionHeader` / `CreateSessionOptions` 的 `type-equiv` 块)。
|
||||
- 既有的 schema v2 SQLite 数据库在打开时被拒绝(预发布阶段无用户数据)。
|
||||
- spawn 回放不变(`seedLength` 为 0)。fork 回放现在将子会话路由到自身的脚本;由 `llm-replay` 测试中的一个回归用例覆盖(一个子会话 fixture,其播种前缀包含父会话的 chunk——推导出的子会话脚本必须排除它,不做 slice 时该用例为红)以及一个持久化往返测试(两个后端,通过共享的 coordinator 契约)。
|
||||
@@ -0,0 +1,6 @@
|
||||
# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write
|
||||
2026-06-22-fork-snapshot-scenarios.md: a5324cbfa13b79c0ea60b74b689f1b19db99a725
|
||||
2026-06-22-fork-snapshot-scenarios.zh.md: 543382db86eb50b5f278a99de74586a13bff9eb7
|
||||
@@ -0,0 +1,31 @@
|
||||
# Agent Note: Record fork and mixed spawn+fork snapshot scenarios
|
||||
|
||||
Status: implemented
|
||||
|
||||
English | [中文](2026-06-22-fork-snapshot-scenarios.zh.md)
|
||||
|
||||
## Problem
|
||||
|
||||
The [seed-boundary Agent Note](2026-06-22-fork-child-replay-seed-boundary.md) made fork-child replay route correctly: `dsh-llm-replay` derives a child's script from the events at or after its persisted `seedLength` boundary, so a fork child's inherited parent prefix is not replayed as the child's own model calls. But it shipped with **no recorded fork scenario** — the slice was exercised only by `llm-replay`'s unit tests (a synthetic child fixture) and a persistence round-trip test. The full-transcript snapshot tier, the one net that boots the real `acp-agent` and replays an end-to-end nested transcript, had only spawn children (`subagent-spawn`, `subagent-multi`). A fork-routing regression that left the unit tests green would still have escaped the tier built to catch transcript regressions.
|
||||
|
||||
The snapshot infrastructure to express a fork scenario was already in place — both in-process backends are wired into `cordis.yml` / `cordis.snapshot.yml` as two model-facing tools (`subagent` → spawn, `subagent_fork` → fork), the harness harvests every child log, and replay forwards per-child fixtures keyed by `seedLength`. What was missing was a *recorded scenario* that drives a fork child through it.
|
||||
|
||||
## Decision
|
||||
|
||||
Record two scenarios against the real API, both replayed keyless in the default gate:
|
||||
|
||||
- **`subagent-fork`** — the parent completes a turn that establishes a fact, then delegates one subtask via `subagent_fork`. The fork child inherits the conversation (its log carries a non-zero `seedLength`), so it can answer from the parent's context. This is the focused regression: the child fixture's `seedLength` is the boundary the replay slice depends on, recorded from a real fork rather than hand-synthesized.
|
||||
- **`subagent-mixed`** — the parent completes a turn, then delegates once via `subagent` (a fresh spawn child, `seedLength` 0) and once via `subagent_fork` (a fork child, non-zero `seedLength`) in one transcript. This is the mixed spawn+fork scenario the seed-boundary and per-session-replay Agent Notes both named as a future addition: one transcript exercises both transports and both branches of the slice (`seedLength` 0 = no-op, `seedLength > 0` = trim the inherited prefix), with the two children ordered spawn-then-fork by `createdAt`.
|
||||
|
||||
### Why a completed turn-1 is required
|
||||
|
||||
The fork backend seeds the child with the parent's **balanced completed-turn prefix**. A parent that forks on its very first turn has no completed turn to inherit, so the seed is empty (≡ a fresh spawn, `seedLength` 0) — which would NOT exercise the slice. Both scenarios therefore use a two-prompt input: the first prompt completes a turn (establishing a codeword the child is later asked to recall), the second delegates the fork. The recalled codeword in the child's transcript is incidental to the model's behavior; the load-bearing artifact is the child fixture's recorded `seedLength`, which the replay slice consumes.
|
||||
|
||||
## Consequences
|
||||
|
||||
- The fork-routing slice is now guarded at the full-transcript tier, not just by unit tests. Removing the `slice(seedLength)` (replaying the whole child log) turns **both** new scenarios red — the fork child receives the parent's recorded chunks instead of its own — proving the guard bites (verified red→green when the scenarios landed).
|
||||
- `subagent-mixed` is the first snapshot scenario to drive two *different* subagent backends in one transcript, exercising the per-session replay keying across a spawn and a fork child simultaneously.
|
||||
- Out-of-process (ACP) subagent replay remains a different shape (each child is its own process with its own replay) and is still tracked as `TODO(acp-subagent-replay)` — these scenarios are in-process only.
|
||||
- Re-recording (`pnpm run test:snapshot:record`) regenerates all four fork/spawn fixtures from the live API; the two new scenarios self-skip without a key like every recorded scenario.
|
||||
|
||||
<!-- agent-note-format: alternatives-not-recorded (pre-format Agent Note) -->
|
||||
@@ -0,0 +1,31 @@
|
||||
# RFC: 记录 fork 与混合 spawn+fork 快照场景
|
||||
|
||||
Status: implemented
|
||||
|
||||
[English](2026-06-22-fork-snapshot-scenarios.md) | 中文
|
||||
|
||||
## 问题
|
||||
|
||||
[seed-boundary RFC](2026-06-22-fork-child-replay-seed-boundary.md) 使 fork 子会话的回放路由正确运作:`dsh-llm-replay` 从子会话持久化的 `seedLength` 边界处或之后的事件推导出子会话的脚本,因此 fork 子会话继承的父会话前缀不会被当作子会话自身的模型调用来回放。但该 RFC 交付时**没有记录 fork 场景**——该切片仅由 `llm-replay` 的单元测试(一个合成的子会话 fixture(测试前置数据))和一个持久化往返测试覆盖。全 transcript(文本记录)快照层(即启动真实 `acp-agent` 并回放端到端嵌套 transcript 的那张网)只有 spawn 子会话(`subagent-spawn`、`subagent-multi`)。如果一个 fork 路由回归让单元测试保持绿色,它仍然会逃过专为捕获 transcript 回归而建的那一层。
|
||||
|
||||
表达 fork 场景所需的快照基础设施已经就位:两个进程内后端都在 `cordis.yml` / `cordis.snapshot.yml` 中以两个面向模型的工具接入(`subagent` → spawn、`subagent_fork` → fork),harness 会收集每个子会话的日志,回放按 `seedLength` 为键转发各子会话的 fixture。缺少的是一个*已记录的场景*来驱动 fork 子会话走完这条路径。
|
||||
|
||||
## 决策
|
||||
|
||||
针对真实 API 记录两个场景,均在默认门禁中以无密钥方式回放:
|
||||
|
||||
- **`subagent-fork`**:父会话完成一个轮次以建立一个事实,然后通过 `subagent_fork` 委派一个子任务。fork 子会话继承对话(其日志携带非零 `seedLength`),因此可以从父会话的上下文中作答。这是聚焦的回归守卫:子会话 fixture 的 `seedLength` 就是回放切片所依赖的边界,来自真实 fork 的记录而非手工合成。
|
||||
- **`subagent-mixed`**:父会话完成一个轮次,然后在同一个 transcript 中分别通过 `subagent`(全新的 spawn 子会话,`seedLength` 为 0)和 `subagent_fork`(fork 子会话,`seedLength` 非零)各委派一次。这是 seed-boundary 和 per-session-replay 两份 RFC 都列为后续补充的混合 spawn+fork 场景:一个 transcript 同时覆盖两种传输方式和切片的两个分支(`seedLength` 0 = 无操作,`seedLength > 0` = 裁剪继承的前缀),两个子会话按 `createdAt` 排序为先 spawn 后 fork。
|
||||
|
||||
### 为什么需要一个已完成的第一轮次
|
||||
|
||||
fork 后端用父会话的**已完成轮次的平衡前缀**([`completedTurnPrefix`](../../../../packages/subagent/subagent-fork))来初始化子会话。如果父会话在第一轮次就 fork,则没有已完成的轮次可继承,seed 为空(等价于全新 spawn,`seedLength` 为 0),这不会覆盖切片逻辑。因此两个场景都使用双 prompt 输入:第一个 prompt 完成一个轮次(建立一个 codeword,子会话稍后被要求回忆它),第二个 prompt 委派 fork。子会话 transcript 中回忆出的 codeword 只是模型行为的附带产物;真正承载验证的产物是子会话 fixture 中记录的 `seedLength`,回放切片消费的正是它。
|
||||
|
||||
## 后果
|
||||
|
||||
- fork 路由切片现在由全 transcript 层守卫,而不仅仅是单元测试。移除 `slice(seedLength)`(回放整个子会话日志)会让**两个**新场景变红——fork 子会话收到的是父会话记录的 chunk 而非自己的——证明守卫确实生效(场景落地时已验证红→绿)。
|
||||
- `subagent-mixed` 是第一个在同一个 transcript 中驱动两种*不同* subagent 后端的快照场景,同时覆盖了跨 spawn 和 fork 子会话的 per-session 回放键控。
|
||||
- 进程外(ACP)subagent 回放形态不同(每个子会话是独立进程、有自己的回放),仍以 `TODO(acp-subagent-replay)` 跟踪——本文场景仅限进程内。
|
||||
- 重新录制(`pnpm run test:snapshot:record`)会从真实 API 重新生成全部四个 fork/spawn fixture;两个新场景在无密钥时自动跳过,与所有已录制场景一致。
|
||||
|
||||
<!-- rfc-format: alternatives-not-recorded (pre-format RFC) -->
|
||||
@@ -0,0 +1,6 @@
|
||||
# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write
|
||||
2026-06-22-subagent-snapshot-replay.md: 89fc4e8d4d267fd4df373a7fd82b8c6e742be6ea
|
||||
2026-06-22-subagent-snapshot-replay.zh.md: 7dc234ab9c6e3bb1facd78e98aad15005d158325
|
||||
@@ -0,0 +1,58 @@
|
||||
# Agent Note: Per-session snapshot replay for nested agents
|
||||
|
||||
Status: implemented
|
||||
|
||||
English | [中文](2026-06-22-subagent-snapshot-replay.zh.md)
|
||||
|
||||
## Problem
|
||||
|
||||
The snapshot tier (`pnpm run test:snapshot`) boots the real `acp-agent` subprocess, replays a recorded session through [`dsh-llm-replay`](../../../../packages/support/llm-replay), and diffs the normalized stdout transcript + re-persisted session log against committed expected outputs. It is the only tier that exercises the full editor-facing transcript end to end.
|
||||
|
||||
It was built for ONE session per process, and that assumption is wired into two places:
|
||||
|
||||
- **`dsh-llm-replay` keyed nothing.** It served the Nth `llm/stream` call the Nth recorded entry from a single global cursor. With a parent agent AND an in-process subagent both streaming on one context, the calls interleave and the single cursor hands the child the parent's script (and vice versa).
|
||||
- **The harness harvested one log.** `findSessionLog` walked the sessions root and returned the FIRST `.jsonl` it found. A subagent runs as a second `Session` with its own log in the same cwd bucket, so the child's transcript was silently dropped.
|
||||
|
||||
This was the `TODO(subagent-snapshots)` deferral recorded in the [subagent seam Agent Note](../feature/2026-06-21-subagent-capability-seam.md): the in-process backends (PR2) shipped with unit + e2e coverage, but the full-transcript snapshot tier could not express a nested-agent shape until this infrastructure landed. This Agent Note is that stacked follow-up.
|
||||
|
||||
## Decision
|
||||
|
||||
Replay is keyed **per calling session**, and the harness harvests **every** session log.
|
||||
|
||||
### 1. The calling session id rides on the model request
|
||||
|
||||
`GenerateOptions` gains an optional `sessionId`, stamped from `agent.session.id` during request assembly. Adapters ignore it; an `llm/stream` listener uses it to route by the issuing session. Its type is `Branded<'SessionId'>` (from `dsh-brand`) rather than `SessionId` from `dsh-session`, because that package imports `Message` from `dsh-llm` and importing back would create a cycle. The types are equivalent, so a session id assigns without a cast. Moving the brand to a dedicated ids package remains separate work because it would touch every id import.
|
||||
|
||||
### 2. Replay binds live sessions to recorded scripts by first-call order
|
||||
|
||||
A nested scenario records more than one log: the parent (`session.jsonl`) plus one per subagent child (`session.1.jsonl`, …). `dsh-llm-replay` loads them all, derives one script per recorded session, and orders the scripts by header `createdAt` (the parent is created before its children).
|
||||
|
||||
Live session ids are freshly random every run and never equal the recorded ones, so a live session cannot bind to a script by id equality. Instead it binds by **first-call order**: the first live session to make any model call claims the first ordered script (the parent — earliest `createdAt`, and necessarily the first to stream, because it must run a turn before it can delegate), the next new live session claims the next script, and so on. Each session then advances its own cursor independently.
|
||||
|
||||
This keys by WHO calls, not by global call order — so it stays correct even if subagents ever run concurrently or in the background (a global cursor would interleave them). A call carrying no `sessionId` (a direct unit-test `stream()`) is treated as one anonymous session bound to the primary script, so the single-session path is byte-for-byte the old behavior. More distinct live sessions than recorded scripts is a fail-loud error (an unrecorded subagent appeared), never a silent mis-route.
|
||||
|
||||
Child fixtures sort by `createdAt`, which matches call order while siblings run strictly sequentially. The id tiebreak only makes degenerate collisions deterministic. Concurrent or background children must introduce an explicit first-call ordinal instead of relying on timestamps.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
The alternative considered and rejected was a **call-ordered merge of the parent and child logs** into one global script (sound only because in-process subagent execution is strictly nested — the parent blocks on the child). It is simpler for today's synchronous cut but bakes in the parent-blocks-on-child invariant that a future backgrounded/concurrent subagent would break; per-session keying does not.
|
||||
|
||||
### 3. The harness harvests every log, primary-first
|
||||
|
||||
`harvestSessionLogs` collects every `.jsonl` across every cwd bucket under the sessions root (the JSONL backend puts a parent and its same-cwd child in the same bucket), parses each header, and orders them primary-first: the top-level session (no `parentSession`) leads, then each child by ascending `createdAt`. `RunResult.sessionLogs` is the plural result; the spec writes each back to its fixture on record (`session.jsonl` + `session.<n>.jsonl`) and diffs each harvested log against its fixture on replay. The normalizer already accepted plural session ids and collapses any stray UUID, so no normalizer change was needed.
|
||||
|
||||
### 4. Scenarios
|
||||
|
||||
Two nested scenarios were added and recorded against the real API:
|
||||
|
||||
- **`subagent-spawn`** — the parent delegates one subtask via the `subagent` tool to a fresh spawn child (2 sessions).
|
||||
- **`subagent-multi`** — the parent delegates two subtasks, each to its own spawn child (3 sessions), stressing the per-session keying with three concurrent scripts and the `createdAt` ordering of two children under one parent.
|
||||
|
||||
Both replay keyless in the default gate.
|
||||
|
||||
## Consequences
|
||||
|
||||
- The `TODO(subagent-snapshots)` deferral is resolved: nested-agent transcripts are now a first-class snapshot shape.
|
||||
- `GenerateOptions.sessionId` is a small, honest core-seam addition useful beyond replay (telemetry, request routing).
|
||||
- The `subagent` tool is bound to a single provider, so both children in `subagent-multi` are spawn (fresh). The keying routes by session, not by backend, so it is already correct for fork. The script *derivation* was not: a fork child's log begins with the seeded parent prefix (the parent's `assistant/chunk` events), so deriving its script from the whole log would replay the parent's responses as the child's. That correctness gap is closed by persisting a seed boundary — see [Persist the seed boundary so fork-child replay routes correctly](2026-06-22-fork-child-replay-seed-boundary.md) — and recorded fork + mixed spawn+fork scenarios now exercise both transports through one transcript (see [Record fork and mixed spawn+fork snapshot scenarios](2026-06-22-fork-snapshot-scenarios.md)).
|
||||
- Out-of-process (ACP) subagents are a different replay shape entirely (each child is its own PROCESS with its own replay), tracked as `TODO(acp-subagent-replay)` in the PR3 plan.
|
||||
@@ -0,0 +1,58 @@
|
||||
# RFC: 嵌套 agent 的逐会话快照回放
|
||||
|
||||
Status: implemented
|
||||
|
||||
[English](2026-06-22-subagent-snapshot-replay.md) | 中文
|
||||
|
||||
## 问题
|
||||
|
||||
快照测试层(`pnpm run test:snapshot`)启动真实的 `acp-agent` 子进程,通过 [`dsh-llm-replay`](../../../../packages/support/llm-replay) 回放录制的会话,并将归一化后的 stdout transcript(文本记录)与重新持久化的会话日志对已提交的金标文件做 diff。这是唯一一个端到端验证完整编辑器侧 transcript 的测试层。
|
||||
|
||||
该层最初为每个进程只有一个会话而构建,这一假设硬编码在两处:
|
||||
|
||||
- **`dsh-llm-replay` 没有做任何键控。** 它用一个全局游标,将第 N 次 `llm/stream` 调用对应到单一录制序列的第 N 条。当父 agent 和一个进程内 subagent 在同一个上下文上同时流式输出时,调用交错,单一游标会把子 agent 的脚本发给父 agent(反之亦然)。
|
||||
- **harness 只收集一份日志。** `findSessionLog` 遍历 sessions 根目录,返回找到的第一个 `.jsonl`。subagent 作为第二个 `Session` 运行,在同一个 cwd bucket 下有自己的日志,因此子 agent 的 transcript 被静默丢弃。
|
||||
|
||||
这就是 [subagent seam RFC](../../implemented/feature/2026-06-21-subagent-capability-seam.md) 中记录的 `TODO(subagent-snapshots)` 延期项:进程内后端(PR2)已有单元测试和 e2e 覆盖,但全 transcript 快照层在本基础设施就绪之前无法表达嵌套 agent 的形态。本 RFC 即为该堆叠后续。
|
||||
|
||||
## 决策
|
||||
|
||||
回放按**调用方会话**键控,harness 收集**所有**会话日志。
|
||||
|
||||
### 1. 调用方会话 id 附着在模型请求上
|
||||
|
||||
`GenerateOptions` 新增可选字段 `sessionId`,在请求组装时从 `agent.session.id` 赋值。适配器忽略它;`llm/stream` 监听器用它按发起会话路由。其类型为 `Branded<'SessionId'>`(来自 `dsh-brand`)而非 `dsh-session` 的 `SessionId`,因为后者所在包(package)导入了 `dsh-llm` 的 `Message`,反向导入会形成循环。两个类型等价,因此会话 id 赋值无需类型转换。将 brand 移到一个专用 ids 包属于独立工作,因为它会影响所有 id 导入。
|
||||
|
||||
### 2. 回放按首次调用顺序将活跃会话绑定到录制脚本
|
||||
|
||||
嵌套场景录制多份日志:父会话(`session.jsonl`)加每个 subagent 子会话各一份(`session.1.jsonl`……)。`dsh-llm-replay` 全部加载,为每个录制会话派生一份脚本,并按 header 中的 `createdAt` 排序(父会话先于子会话创建)。
|
||||
|
||||
活跃会话 id 每次运行都是全新随机值,永远不等于录制时的 id,因此活跃会话无法通过 id 相等绑定到脚本。取而代之的是**首次调用顺序**绑定:第一个发起任何模型调用的活跃会话认领第一份有序脚本(即父会话:`createdAt` 最早,且必然最先流式输出,因为它必须先运行一个轮次才能委派),下一个新活跃会话认领下一份脚本,依此类推。此后每个会话独立推进自己的游标。
|
||||
|
||||
这种方式按谁在调用键控,而非按全局调用顺序。因此即使 subagent 将来并发或在后台运行(全局游标会导致交错),它仍然正确。不携带 `sessionId` 的调用(直接在单元测试中调用 `stream()`)被视为一个匿名会话、绑定到主脚本,因此单会话路径与旧行为逐字节一致。活跃会话数多于录制脚本数时会快速失败报错(出现了未录制的 subagent),绝不会静默错误路由。
|
||||
|
||||
子 fixture(测试前置数据)按 `createdAt` 排序,在兄弟会话严格顺序执行时与调用顺序一致。id 平局打破仅使退化碰撞具有确定性。并发或后台子会话必须引入显式的首次调用序号,而非依赖时间戳。
|
||||
|
||||
## 曾考虑的替代方案
|
||||
|
||||
曾考虑但否决的方案是:**将父子日志按调用顺序合并**为一份全局脚本(仅在进程内 subagent 执行严格嵌套——父 agent 阻塞等待子 agent——时才正确)。对当前的同步裁剪而言更简单,但将「父阻塞于子」这一不变式固化了进去;未来若引入后台/并发 subagent 就会失效。逐会话键控则不会。
|
||||
|
||||
### 3. harness 收集所有日志,主会话优先
|
||||
|
||||
`harvestSessionLogs` 收集 sessions 根目录下每个 cwd bucket 中的所有 `.jsonl`(JSONL 后端将父会话与同 cwd 的子会话放在同一个 bucket),解析各自的 header,并按主会话优先排序:顶层会话(无 `parentSession`)在前,各子会话按 `createdAt` 升序排列。`RunResult.sessionLogs` 是复数结果;spec 在录制时将每份日志写回对应 fixture(`session.jsonl` + `session.<n>.jsonl`),在回放时将每份收集到的日志与其 fixture 做 diff。归一化器已支持复数会话 id 并会折叠任何游离 UUID,因此无需修改归一化器。
|
||||
|
||||
### 4. 场景
|
||||
|
||||
新增两个嵌套场景,均对真实 API 录制:
|
||||
|
||||
- **`subagent-spawn`**:父 agent 通过 `subagent` 工具将一个子任务委派给一个新 spawn 的子 agent(2 个会话)。
|
||||
- **`subagent-multi`**:父 agent 委派两个子任务,各自交给自己的 spawn 子 agent(3 个会话),以三份并行脚本和同一父 agent 下两个子会话的 `createdAt` 排序来压测逐会话键控。
|
||||
|
||||
两者均在默认门禁中以 keyless 方式回放。
|
||||
|
||||
## 后果
|
||||
|
||||
- `TODO(subagent-snapshots)` 延期项已解决:嵌套 agent 的 transcript 现在是快照层的一等形态。
|
||||
- `GenerateOptions.sessionId` 是一个小而诚实的 core-seam 新增,在回放之外同样有用(遥测、请求路由)。
|
||||
- `subagent` 工具绑定到单一提供方,因此 `subagent-multi` 中的两个子 agent 都是 spawn(全新创建)。键控按会话路由而非按后端路由,因此对 fork 同样正确。但脚本*派生*逻辑此前不正确:fork 子会话的日志以种子化的父前缀(父会话的 `assistant/chunk` 事件)开头,如果从完整日志派生脚本,就会把父 agent 的响应当作子 agent 的来回放。这一正确性缺口通过持久化种子边界来弥合——见 [Persist the seed boundary so fork-child replay routes correctly](2026-06-22-fork-child-replay-seed-boundary.md)——录制的 fork 与混合 spawn+fork 场景现在通过一份 transcript 同时验证两种传输方式(见 [Record fork and mixed spawn+fork snapshot scenarios](2026-06-22-fork-snapshot-scenarios.md))。
|
||||
- 进程外(ACP)subagent 是完全不同的回放形态(每个子 agent 是自己的进程、有自己的回放),作为 `TODO(acp-subagent-replay)` 记录在 PR3 计划中。
|
||||
@@ -0,0 +1,6 @@
|
||||
# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write
|
||||
2026-07-04-hook-snapshot-matrix.md: b365992c01e081e5698e81a9ff9682e9b8166ce6
|
||||
2026-07-04-hook-snapshot-matrix.zh.md: bcb8e5ba55a14dd0c299dac161146a19f18201eb
|
||||
@@ -0,0 +1,52 @@
|
||||
# Agent Note: Hook snapshot matrix — end-to-end expected outputs for both bridges
|
||||
|
||||
Status: implemented
|
||||
|
||||
English | [中文](2026-07-04-hook-snapshot-matrix.zh.md)
|
||||
|
||||
## Problem
|
||||
|
||||
The hook bridges — [`dsh-hooks-claude`](../../../../packages/hooks/hooks-claude) (7 Claude Code hook points) and [`dsh-hooks-codex`](../../../../packages/hooks/hooks-codex) (5 Codex points) — map external hook commands onto the harness interception seams. They carry deep unit and coverage-spec coverage (every decision arm, every payload dialect, driven against a mocked seam) plus one key-gated e2e (`hooks.e2e.ts`, a live `PreToolUse` block). But the full-transcript snapshot tier — the one net that boots the real `acp-agent` subprocess, replays a recorded session keyless, and diffs the normalized ACP stdout + re-persisted log against committed expected outputs — covered exactly ONE hook: a Claude `UserPromptSubmit` block (`hook-cc-promptsubmit-block`).
|
||||
|
||||
That is the tier a mocked unit test structurally cannot be: it exercises the REAL bridge translating a REAL hook process's outcome into the REAL seam decision, then the REAL loop's reaction, rendered exactly as an editor sees it. A bridge-translation or loop-structure regression that left every unit green would still escape it for every hook point but one — and for the Codex bridge, the ACP example did not even LOAD it, so no Codex hook could fire end-to-end at all.
|
||||
|
||||
## Decision
|
||||
|
||||
The implementation has two coupled parts:
|
||||
|
||||
### 1. The ACP example ships BOTH hook bridges
|
||||
|
||||
`examples/acp-agent/cordis.yml` and `cordis.snapshot.yml` now load `dsh-hooks-codex` alongside `dsh-hooks-claude`, each pointed at its own config file (`./hooks.json` for Claude, `./codex-hooks.json` for Codex — the two dialects cannot share one file). This is a genuine product-surface change, not test-only wiring: the shipped ACP server (and the `demo:acp` front door) now carries both bridges.
|
||||
|
||||
It is safe because a bridge whose config file is absent is a **silent no-op**: `apply()` catches the read failure, logs through `ctx.logger`, and registers nothing — zero listeners, zero session events. The `acp-agent` app ships no stdout logger, so the warning cannot reach the ACP JSON-RPC channel. A scenario (or a real project) that wants only Claude hooks ships only `hooks.json`; the Codex bridge sees no `codex-hooks.json` and vanishes. This was verified empirically: with both bridges loaded, all pre-existing snapshots (none of which ship a `codex-hooks.json`) are byte-identical.
|
||||
|
||||
Loading both is the minimum that lets the snapshot tier exercise each dialect against the same real app the product ships. Recording (which boots `cordis.yml`) loads both by construction, and replay inherits them the same way: `cordis.snapshot.yml` is an include-overlay of `cordis.yml` that swaps only the llm entry (see [single-source the acp-agent replay config](2026-07-04-single-source-acp-replay-config.md)), so a bridge added to the live tree is in the replay tree with no second edit.
|
||||
|
||||
### 2. A snapshot scenario per hook point × its headline outcome, both dialects
|
||||
|
||||
Thirteen scenarios under `examples/acp-agent/tests/snapshots/`, naming `hook-<dialect>-<point>-<outcome>`:
|
||||
|
||||
- **Authored, no model turn** (keyless, no sidecar — the derived replay script is empty; the `rejected` turn carrying `hook/*` events is compared): `hook-cc-promptsubmit-block`, `hook-codex-promptsubmit-block`.
|
||||
- **Recorded against the real API, hook active during recording** (the model's reaction to the decision is part of the captured transcript, replayed keyless thereafter): `hook-{cc,codex}-promptsubmit-context` (allow + additionalContext fold), `hook-cc-pretool-deny` / `hook-codex-pretool-block` (deny → `isError` tool result), `hook-cc-pretool-ask` (ask → degrades to deny with the approval-required reason), `hook-{cc,codex}-posttool-block` (block with feedback), `hook-{cc,codex}-posttool-context` (accept + additionalContext), `hook-{cc,codex}-stop-continue` (a blocking Stop hook forces one extra step via steering).
|
||||
|
||||
Each hook command emits only FIXED LITERAL strings (no timestamps/pids/`$RANDOM`/cwd echoes); the snapshot normalizer scrubs the one volatile field a `hook/result` carries (`durationMs`). The `Stop` scenarios self-limit with a marker file (`.stop_fired`) so the force-continue does not loop — the `stop_hook_active` loop-guard is still a bridge `TODO`, so an unconditional Stop hook would force-continue every step.
|
||||
|
||||
The `PostToolUse` block scenarios self-limit at the mechanism they prove. The Claude hook persists a workspace marker after its first rejection, so one recovery call is allowed; the Codex prompt makes one call and reports the injected result. Each expected output pins one blocked call without repeated block/retry cycles.
|
||||
|
||||
### Three hook points are deliberately NOT snapshotted
|
||||
|
||||
Discovered while building the matrix, and documented here because the omission is a decision, not an oversight:
|
||||
|
||||
- **`SessionStart` and `SubagentStart`** inject context through a detached, best-effort `void runPoint(...).then(agent.inject())` with NO turn binding. The resulting `context/message` races the work it precedes (the first model request / the child's first turn) and lands at a nondeterministic log position. A recorded expected output does not even reproduce on its own replay — a 10× replay stability check failed 10/10 for both. They stay on the bridges' unit coverage, which drives the seam directly without the timing race. (If the injection is ever made turn-bound and deterministic — the direction the `TODO(session-start-gating)` points — these become snapshottable.)
|
||||
- **`SubagentStop`** is observe-only: its `subagent/end` handler passes no turn (so no `hook/*` log events) and does no injection. It writes NOTHING to the transcript, so an expected output would be byte-identical to the no-hook run and could never be proven to fail — a guard that cannot bite. It stays on unit coverage (`bridge.spec.ts` already asserts the observe-only call).
|
||||
|
||||
The matrix therefore covers every hook point that has a DETERMINISTIC, OBSERVABLE transcript footprint, for both dialects.
|
||||
|
||||
## Consequences
|
||||
|
||||
- Every bridge seam mapping with an observable transcript is now guarded at the full-transcript tier, in the real app, for both dialects — including the Codex bridge, which had no end-to-end coverage at all. Recorded expected outputs capture the model's real reaction to a denied/blocked/force-continued turn, which a hand-authored transcript could only guess at.
|
||||
- The `UserPromptSubmit` block scenarios are authored keylessly (no model turn); the rest replay keylessly from recorded fixtures. `pnpm run test:snapshot:record` regenerates the recorded fixtures from the live API and self-skips without a key like every recorded scenario.
|
||||
- The prove-red discipline holds: tampering a hook config's output (e.g. changing a deny reason) turns its scenario red on replay — the hook process runs FOR REAL during replay (only the model is replayed), so the expected output guards the actual hook→seam→loop path, not a mock of it.
|
||||
- The `acp-agent` demo now loads a Codex bridge it will usually no-op (no `codex-hooks.json` in a typical project), which is the intended fail-soft behavior, not a cost.
|
||||
|
||||
<!-- agent-note-format: alternatives-not-recorded (pre-format Agent Note) -->
|
||||
@@ -0,0 +1,50 @@
|
||||
# RFC: Hook 快照矩阵——覆盖两种 bridge 的端到端 golden 测试
|
||||
|
||||
Status: implemented
|
||||
|
||||
[English](2026-07-04-hook-snapshot-matrix.md) | 中文
|
||||
|
||||
## 问题
|
||||
|
||||
hook bridge——[`dsh-hooks-claude`](../../../../packages/hooks/hooks-claude)(7 个 Claude Code hook 点)和 [`dsh-hooks-codex`](../../../../packages/hooks/hooks-codex)(5 个 Codex 点)——将外部 hook 命令映射到 harness 的拦截 seam 上。它们拥有深度的单元测试和 coverage-spec 覆盖率(每个决策分支、每种 payload 方言,均对 mock 的 seam 驱动),外加一个需要密钥的 e2e 测试(`hooks.e2e.ts`,一次真实的 `PreToolUse` 拦截)。但完整 transcript(文本记录)快照层:那张真正启动 `acp-agent` 子进程、无密钥回放录制会话、并将规范化的 ACP stdout 与重新持久化的日志与已提交 golden 做 diff 的网,只覆盖了一个 hook:Claude 的 `UserPromptSubmit` 拦截(`hook-cc-promptsubmit-block`)。
|
||||
|
||||
这正是 mock 单元测试在结构上无法替代的层级:它验证的是真实 bridge 将真实 hook 进程的结果翻译到真实 seam 决策,再到真实 agent loop(智能体循环)的反应,渲染结果与编辑器看到的完全一致。一个 bridge 翻译或 loop 结构的回归,即使让所有单元测试保持绿色,也会在除那一个 hook 点之外的所有点上逃逸;而对于 Codex bridge,ACP 示例甚至没有加载它,因此没有任何 Codex hook 能端到端触发。
|
||||
|
||||
## 决策
|
||||
|
||||
实现由两个耦合部分组成:
|
||||
|
||||
### 1. ACP 示例同时加载两种 hook bridge
|
||||
|
||||
`examples/acp-agent/cordis.yml` 和 `cordis.snapshot.yml` 现在同时加载 `dsh-hooks-codex` 与 `dsh-hooks-claude`,各自指向自己的配置文件(Claude 用 `./hooks.json`,Codex 用 `./codex-hooks.json`——两种方言无法共用一个文件)。这是一个真正的产品接口变更,而非仅用于测试的接线:交付的 ACP 服务器(以及 `demo:acp` 入口)现在同时携带两种 bridge。
|
||||
|
||||
这是安全的,因为配置文件不存在时 bridge 是**静默无操作**的:`apply()` 捕获读取失败、通过 `ctx.logger` 记录日志、不注册任何东西——零监听器、零会话事件。`acp-agent` 应用不附带 stdout logger,因此警告不会到达 ACP JSON-RPC 通道。只需要 Claude hook 的场景(或真实项目)只提供 `hooks.json`;Codex bridge 找不到 `codex-hooks.json` 便自动消失。这已通过实验验证:在两种 bridge 同时加载的情况下,所有既有快照(均不附带 `codex-hooks.json`)逐字节一致。
|
||||
|
||||
同时加载是让快照层能够在产品交付的同一个真实应用上验证每种方言的最低要求。录制(启动 `cordis.yml`)天然加载两者,回放以同样方式继承:`cordis.snapshot.yml` 是 `cordis.yml` 的 include-overlay,只替换 llm 入口(见[单一来源 acp-agent 回放配置](2026-07-04-single-source-acp-replay-config.md)),因此添加到运行时树的 bridge 无需第二次编辑即出现在回放树中。
|
||||
|
||||
### 2. 每个 hook 点 × 其主要结果各一个快照场景,覆盖两种方言
|
||||
|
||||
`examples/acp-agent/tests/snapshots/` 下共 13 个场景,命名为 `hook-<dialect>-<point>-<outcome>`:
|
||||
|
||||
- **手工编写、无模型轮次**(无密钥、无 sidecar——派生的回放脚本为空;比对的是携带 `hook/*` 事件的 `rejected` 轮次):`hook-cc-promptsubmit-block`、`hook-codex-promptsubmit-block`。
|
||||
- **对真实 API 录制、录制期间 hook 活跃**(模型对决策的反应是捕获的 transcript 的一部分,此后无密钥回放):`hook-{cc,codex}-promptsubmit-context`(allow + additionalContext 折叠)、`hook-cc-pretool-deny` / `hook-codex-pretool-block`(deny → `isError` 工具结果)、`hook-cc-pretool-ask`(ask → 降级为 deny 并附带 approval-required 原因)、`hook-{cc,codex}-posttool-block`(block 并附带反馈)、`hook-{cc,codex}-posttool-context`(accept + additionalContext)、`hook-{cc,codex}-stop-continue`(阻塞性 Stop hook 通过 steering(中途引导)强制多走一步)。
|
||||
|
||||
每个 hook 命令只输出固定字面量字符串(无时间戳/pid/`$RANDOM`/cwd 回显);快照规范化器擦除 `hook/result` 携带的唯一不稳定字段(`durationMs`)。`Stop` 场景通过标记文件(`.stop_fired`)自限,使 force-continue 不会循环——`stop_hook_active` 循环守卫仍是 bridge 的一个 `TODO`,因此无条件的 Stop hook 会在每一步都 force-continue。
|
||||
|
||||
### 三个 hook 点被有意排除在快照之外
|
||||
|
||||
在构建矩阵过程中发现,记录于此是因为这些遗漏是决策而非疏忽:
|
||||
|
||||
- **`SessionStart` 与 `SubagentStart`** 通过一个分离的、尽力而为的 `void runPoint(...).then(agent.inject())` 注入上下文,没有轮次绑定。由此产生的 `context/message` 与它所先于的工作(首次模型请求/子 agent 的首轮)存在竞争,落在日志中的位置不确定。录制的 golden 甚至无法在自身回放中复现——10 次回放稳定性检查对两者均 10/10 失败。它们留在 bridge 的单元覆盖率中,单元测试直接驱动 seam 而无时序竞争。(如果注入将来变为轮次绑定且确定性的——`TODO(session-start-gating)` 所指的方向——它们就可以纳入快照。)
|
||||
- **`SubagentStop`** 是纯观察性的:其 `subagent/end` 处理器不传递轮次(因此无 `hook/*` 日志事件)、不做注入。它对 transcript 不写入任何内容,因此 golden 与无 hook 运行逐字节一致,永远无法被证明失败——一道永远不会触发的守卫。它留在单元覆盖率中(`bridge.spec.ts` 已断言了纯观察调用)。
|
||||
|
||||
因此,该矩阵覆盖了所有具有确定性、可观测 transcript 足迹的 hook 点,涵盖两种方言。
|
||||
|
||||
## 后果
|
||||
|
||||
- 每个具有可观测 transcript 的 bridge seam 映射现在都在完整 transcript 层级、在真实应用中、对两种方言受到守护——包括此前完全没有端到端覆盖率的 Codex bridge。录制的 golden 捕获了模型对 deny/block/force-continue 轮次的真实反应,这是手工编写的 transcript 只能猜测的。
|
||||
- block 场景无需密钥(无模型轮次);其余场景从录制的 fixture(测试前置数据)无密钥回放。`pnpm run test:snapshot:record` 从真实 API 重新生成录制的 fixture,无密钥时自动跳过,与所有录制场景一致。
|
||||
- prove-red 纪律成立:篡改 hook 配置的输出(例如修改 deny 原因)会使其场景在回放时变红——hook 进程在回放期间真实运行(只有模型被回放),因此 golden 守护的是实际的 hook→seam→loop 路径,而非它的 mock。
|
||||
- `acp-agent` 演示现在加载了一个通常会无操作的 Codex bridge(典型项目中没有 `codex-hooks.json`),这正是预期的柔性失败行为,而非代价。
|
||||
|
||||
<!-- rfc-format: alternatives-not-recorded (pre-format RFC) -->
|
||||
@@ -0,0 +1,6 @@
|
||||
# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write
|
||||
2026-07-04-single-source-acp-replay-config.md: 51cbd54d45408df1548c9cc2522b07b5ffaac110
|
||||
2026-07-04-single-source-acp-replay-config.zh.md: 2aec0e0e46007be243c0386fc0fca92065ce3c9e
|
||||
@@ -0,0 +1,27 @@
|
||||
# Agent Note: Single-source the acp-agent replay config
|
||||
|
||||
Status: implemented
|
||||
|
||||
English | [中文](2026-07-04-single-source-acp-replay-config.zh.md)
|
||||
|
||||
## Problem
|
||||
|
||||
`examples/acp-agent` shipped two hand-maintained configs: `cordis.yml` (the live tree) and a `cordis.snapshot.yml` that mirrored it entry-for-entry with only the llm backend swapped — stripped of comments, the entire difference was the eight-line `llm-deepseek` stanza versus the two-line `llm-replay` stanza. Every app-shape change had to be made twice, and nothing gated the symmetry: if the copies drifted, the snapshot tier would silently exercise a different app than the one that ships — the ["green units, broken product" class of gap](../../../../docs/postmortem/0001-acp-default-export-drops-inject.md) the snapshot tier exists to close, reintroduced one level up, with reviewer vigilance as the only defense.
|
||||
|
||||
## Decision
|
||||
|
||||
`cordis.snapshot.yml` includes the live config, disables the named DeepSeek adapter by id and name, and inserts the replay adapter. Every other entry therefore comes from the shipping tree. Replay selects the overlay; recording still boots `cordis.yml`, and the load guard permits the intentionally disabled entry.
|
||||
|
||||
One vendored-plugin fact the overlay depends on, deliberately: the include applies `patches` when it loads the file — its `refresh()`/`internal/update` paths re-read without re-patching — which is exactly enough for a one-shot replay boot (the replay app loads no `hmr` and nothing rewrites the config mid-run). The snapshot suite is the proof: all scenarios pass unchanged on the overlay, byte-identical expected outputs included.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
### Why not the alternatives?
|
||||
|
||||
Keeping the full twin with a symmetry verify-gate was the recorded fallback — it would have removed the silent-drift class but kept a 125-line near-copy whose only content was one entry's difference, growing with every plugin the app gains. A bin-side swap (parse the config, replace the entry, delete the file) would have put YAML surgery inside a published artifact and moved the replay delta out of sight; the overlay keeps the delta declarative, readable, and next to the base config — the teaching value the twin's defenders actually wanted.
|
||||
|
||||
## Consequences
|
||||
|
||||
- A plugin added to `cordis.yml` is in the replay tree with no second edit; the drift class is structurally gone rather than gated.
|
||||
- The overlay depends on entries carrying stable `id:`s. The `name` assertion on the disable patch guards mis-targeting (a reused id skips the patch instead of disabling the wrong plugin). An id RENAME degrades the patch to a skip whose warning needs a logger the replay app deliberately lacks — the observable result is a futile keyless `llm-deepseek` entry alongside `llm-replay`, with replay output still correct (`llm-replay` owns the stream short-circuit); config rot for review to catch, not wrong snapshots. A top-level insert whose id collides with an existing entry resolves last-wins through the loader's id map — the current config has no collision, and a new patch line is where one would be introduced.
|
||||
- If a future replay tree needs a second divergence (another backend swapped), it is one more patch line, not a second fork of the file.
|
||||
@@ -0,0 +1,27 @@
|
||||
# RFC: 将 acp-agent 回放配置改为单一来源
|
||||
|
||||
Status: implemented
|
||||
|
||||
[English](2026-07-04-single-source-acp-replay-config.md) | 中文
|
||||
|
||||
## 问题
|
||||
|
||||
`examples/acp-agent` 曾维护两份手写配置:`cordis.yml`(正式运行树)和 `cordis.snapshot.yml`(逐条镜像前者,仅替换 LLM(大语言模型)后端)。去掉注释后,全部差异只是八行的 `llm-deepseek` 段落换成两行的 `llm-replay` 段落。每次应用结构变更都要改两遍,且没有门禁保障对称性:一旦两份副本漂移,快照层就会悄悄测试一个与实际交付不同的应用——正是快照层本要消除的["单元测试全绿、产品却坏了"这类缺口](../../../postmortem/0001-acp-default-export-drops-inject.md),在上一层被重新引入,唯一的防线是评审者的警觉。
|
||||
|
||||
## 决策
|
||||
|
||||
`cordis.snapshot.yml` include 正式配置,通过 id 和 name 禁用指定的 DeepSeek 适配器,并插入回放适配器。其余所有条目因此来自正式运行树。回放时选择 overlay;录制仍然启动 `cordis.yml`,加载守卫允许被有意禁用的条目。
|
||||
|
||||
overlay 依赖一个 vendor 插件的事实,这是有意为之:include 在加载文件时应用 `patches`,其 `refresh()`/`internal/update` 路径重读时不会重新打补丁。这恰好满足一次性回放启动的需要(回放应用不加载 `hmr`,也没有东西在运行中改写配置)。快照套件即为证明:所有场景在 overlay 上原样通过,包括逐字节一致的 golden 文件。
|
||||
|
||||
## 曾考虑的替代方案
|
||||
|
||||
### 为何不采用这些替代方案?
|
||||
|
||||
保留完整的双副本并加一道对称性校验门禁是记录在案的退路——它能消除静默漂移这一类问题,但仍保留一份 125 行的近乎复制品,其全部内容只是一个条目的差异,且随应用每增加一个插件而增长。在 bin 侧做替换(解析配置、替换条目、删除文件)则会把 YAML 手术放进发布产物,并把回放差异藏到视线之外;overlay 让差异保持声明式、可读,且紧邻基础配置——这正是双副本支持者真正看重的教学价值。
|
||||
|
||||
## 后果
|
||||
|
||||
- 向 `cordis.yml` 添加插件即自动进入回放树,无需第二次编辑;漂移这一类问题从结构上消失,而非靠门禁拦截。
|
||||
- overlay 依赖条目携带稳定的 `id:`。禁用补丁上的 `name` 断言防止误定位(id 被复用时补丁跳过而非禁用错误的插件)。如果 id 被重命名,补丁退化为跳过,其警告需要一个回放应用有意不具备的 logger——可观测结果是一条无效的无密钥 `llm-deepseek` 条目与 `llm-replay` 并存,回放输出仍然正确(`llm-replay` 拥有流的短路权);这属于配置腐烂,留给评审发现,不会产生错误的快照。顶层插入一个 id 与既有条目冲突的新条目时,loader 的 id map 以后者为准;当前配置无冲突,新增补丁行才是引入冲突的场所。
|
||||
- 如果未来回放树需要第二处差异(另一个后端被替换),只需多加一行补丁,而非再 fork 一份文件。
|
||||
+6
@@ -0,0 +1,6 @@
|
||||
# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write
|
||||
2026-07-06-pin-request-header-content-in-one-scenario.md: 0166b459fbb8d883f07fb195bdd5e025d70349de
|
||||
2026-07-06-pin-request-header-content-in-one-scenario.zh.md: 1ca7df68fc743b919c6769ae8fa40ea16eb3d88a
|
||||
+35
@@ -0,0 +1,35 @@
|
||||
# Agent Note: Pin request-header content in one snapshot scenario
|
||||
|
||||
Status: implemented
|
||||
|
||||
English | [中文](2026-07-06-pin-request-header-content-in-one-scenario.zh.md)
|
||||
|
||||
## Problem
|
||||
|
||||
An ACP snapshot suite needs to prove the exact composed system prompt and tool-schema list sent in each `request/header`, but duplicating that content inside every `session.jsonl` makes a prompt or schema edit rewrite dozens of giant one-line JSON records. Keeping one raw header avoids the duplication but still makes prompt review poor: prose is JSON-escaped onto one line and mixed with thousands of characters of tool schemas.
|
||||
|
||||
## Decision
|
||||
|
||||
Exactly one scenario per header-composition class is flagged `pinsHeader`. Its directory splits the pin by review format: `system-prompt.expected.md` contains the normalized full prompt sequence as ordinary Markdown, `tool-schemas.expected.json` contains the corresponding complete schema sequence as structured JSON, and `session.jsonl` retains config, reason, and any model-visible prefix while storing `header.system` and `header.tools` as `"{{system}}"` / `"{{tools}}"`. Every other JSONL uses the same prompt and tool tokens and also tokenizes session-prefix content. The pin mechanics live in [`dsh-acp-snapshot`](../../../../packages/support/acp-snapshot/README.md), whose suite factory enforces one pin per class.
|
||||
|
||||
The pure `scrubSystemPrompts` and `scrubToolSchemas` normalizers independently tokenize every stored full header. `scrubRequestHeaders` also tokenizes session-prefix content for non-pinning scenarios while retaining header count, field presence, config, reason, and prefix message count. Record and refresh write-back apply the appropriate scrub before writing JSONL and regenerate both sidecars from the normalized live full-header sequence, so neither path can reintroduce prompt/schema bulk into JSONL or leave a review artifact stale.
|
||||
|
||||
Guards make the split self-enforcing. On disk, every `session*.jsonl` is a fixed point of both prompt and schema scrubbers, only non-pinning fixtures must be fixed points of the full header scrub, both sidecars exist exactly beside pinning fixtures in canonical newline-terminated formats, and each class has one pin. Live, every `request/header` produced by a parent, spawn child, fork child, initial request, resume, or in-instance change must match the reconstructed class sequence after volatile-value normalization. A header without a string prompt, without an array-valued tool list, or beyond the pin's declared changed-header count fails loud.
|
||||
|
||||
One pin covers the whole suite because every session — parent, spawn child, fork child — composes the identical tool list and the identical prompt modulo cwd, and the uniformity guard fails the suite the moment that stops holding. If header composition ever becomes session-dependent by design (a restricted subagent toolset, say), the divergent shape gets its own pinning scenario.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
- **Re-record or hand-edit every fixture per change** — preserves exact headers but buries behavioral diffs under duplicated prompt and schema content.
|
||||
- **Scrub at compare time only, keeping fixtures raw** — lets compares pass while committed fixtures retain stale duplicate content and rewrite wholesale on the next recording. Stored tokens state honestly what each JSONL does not pin.
|
||||
- **Scrub everywhere, pin nowhere** — loses the only end-to-end record of the composed header as actually sent (prompt assembly, registered-tool order, full schemas). The generated tool catalog documents each tool in isolation; only a real fixture pins the composed set.
|
||||
- **Keep the one full pin entirely in JSONL** — removes suite-wide duplication but leaves prompt and schema changes as one escaped line. Markdown and structured JSON give each surface its natural review format without weakening the reconstructed-header assertion.
|
||||
- **Slim the session log itself (log a content digest, store the header elsewhere)** — violates the reconstructability contract: the product log must reproduce each request bit-for-bit ([reconstructable-requests Agent Note](../architecture/2026-07-05-reconstructable-requests.md)). Header bulk is a test-artifact concern, solved in test normalization; the live log is untouched.
|
||||
|
||||
## Verification
|
||||
|
||||
The suite replays every scenario against the split pins. Unit coverage exercises the independent and full scrubbers, both full-header sidecar formats, record/refresh regeneration, normalized prompt/schema extraction, fixed-point enforcement, required-file symmetry, reconstructed-header uniformity, and changed-header count rejection.
|
||||
|
||||
## Consequences
|
||||
|
||||
A system-prompt change produces a line-oriented Markdown diff in one file per affected composition class; a tool-description change produces a structured JSON diff in one file per class; ordinary behavioral fixtures remain untouched. Session fixtures display tokens for omitted content, and the live uniformity guard makes each split pin authoritative for every session in its class. Each pinning scenario carries two generated, newline-canonicalized sidecars.
|
||||
+35
@@ -0,0 +1,35 @@
|
||||
# RFC: 在单个快照场景中固定请求头内容
|
||||
|
||||
Status: implemented
|
||||
|
||||
[English](2026-07-06-pin-request-header-content-in-one-scenario.md) | 中文
|
||||
|
||||
## 问题
|
||||
|
||||
一个 ACP(Agent Client Protocol)快照测试套件需要证明每个 `request/header` 中实际发送的组合系统提示词与工具 schema 列表,但如果在每个 `session.jsonl` 中重复这些内容,一次提示词或 schema 编辑就会改写数十条巨大的单行 JSON 记录。保留一份原始 header 可以避免重复,但提示词的评审体验仍然很差:行文被 JSON 转义到一行中,与数千字符的工具 schema 混在一起。
|
||||
|
||||
## 决策
|
||||
|
||||
每个 header 组合类别恰好有一个场景被标记为 `pinsHeader`。其目录按评审格式拆分固定内容:`system-prompt.golden.md` 以普通 Markdown 存放归一化后的组合提示词,`tool-schemas.golden.json` 以结构化 JSON 存放完整的初始 schema 及后续 schema 变更,而 `session.jsonl` 保留 config、reason 及任何模型可见的前缀,同时将 `header.system` 和 `header.tools` 存为 `"{{system}}"` / `"{{tools}}"`。其余所有 JSONL 使用相同的提示词和工具 token,并同样对会话前缀内容做 token 化处理。固定机制实现在 [`dsh-acp-snapshot`](../../../../packages/support/acp-snapshot/README.md) 中,其套件工厂强制每个类别只有一个固定场景。
|
||||
|
||||
纯粹的 `scrubSystemPrompts` 和 `scrubToolSchemas` 归一化器应用于每个存储的会话 fixture(测试前置数据),独立地对初始 header 内容和 header-delta 批量内容做 token 化。`scrubRequestHeaders` 还为非固定场景的会话前缀内容做 token 化,同时保留结构性事实:system-delta 的位置与数量、新增/移除/变更的工具名称、前缀消息数量、字段存在性、config 和 reason。record 与 refresh 的回写操作在写入 JSONL 前应用相应的 scrub,并从归一化后的实时 header 和 delta 重新生成两个 sidecar 文件,因此两条路径都不会把提示词/schema 批量内容重新引入 JSONL,也不会让评审产物变陈旧。
|
||||
|
||||
守卫机制使这一拆分自我强制。在磁盘上:每个 `session*.jsonl` 都是提示词和 schema 两个 scrubber 的不动点;只有非固定 fixture 还必须是完整 header scrub 的不动点;两个 sidecar 文件恰好存在于固定 fixture 旁边,采用规范的换行终止格式;每个类别有且仅有一个固定场景。在运行时:由 parent、spawn 子会话、fork 子会话、初始请求或 resume 产生的每个 `request/header`,在经过易变值归一化后必须与重建的固定内容匹配;固定运行的提示词和 schema delta 也必须与其 sidecar 匹配。如果 header 没有字符串类型的 prompt、没有数组类型的工具列表,或包含未声明的 `request/header-delta`,则立即失败并报错。
|
||||
|
||||
一个固定场景覆盖整个套件,因为每个会话(parent、spawn 子会话、fork 子会话)组合出的工具列表完全相同、提示词除 cwd 外完全相同,而一致性守卫会在这一前提不再成立时立即使套件失败。如果 header 组合将来在设计上变为会话相关的(例如受限的 subagent 工具集),那么分歧的形态将获得自己的固定场景。
|
||||
|
||||
## 曾考虑的替代方案
|
||||
|
||||
- **每次变更重新录制或手动编辑所有 fixture**:保留了精确的 header,但行为差异被重复的提示词和 schema 内容淹没。
|
||||
- **仅在比较时 scrub,fixture 保持原始内容**:比较能通过,但已提交的 fixture 保留着陈旧的重复内容,下次录制时会整体重写。存储 token 诚实地表明每个 JSONL 没有固定什么。
|
||||
- **全部 scrub,不做任何固定**:丢失了组合 header 实际发送内容(提示词组装、已注册工具顺序、完整 schema)的唯一端到端记录。生成的工具目录只孤立地记录每个工具;只有真实 fixture 才能固定组合后的完整集合。
|
||||
- **将完整固定内容全部保留在 JSONL 中**:消除了套件范围的重复,但提示词和 schema 变更仍然是一行转义文本。Markdown 和结构化 JSON 为每种内容提供其自然的评审格式,同时不削弱重建 header 的断言。
|
||||
- **精简会话日志本身(记录内容摘要,将 header 存放在别处)**:违反可重建性契约:产品日志必须逐位重现每个请求([可重建请求 RFC](../architecture/2026-07-05-reconstructable-requests.md))。header 体积是测试产物的问题,在测试归一化中解决;线上日志不受影响。
|
||||
|
||||
## 验证
|
||||
|
||||
套件针对拆分后的固定内容回放每个场景。单元测试覆盖率涵盖独立 scrubber 和完整 scrubber、两种 sidecar 格式、record/refresh 重新生成、归一化提示词/schema 提取、不动点强制、必需文件对称性、重建 header 一致性以及 delta 拒绝。
|
||||
|
||||
## 后果
|
||||
|
||||
系统提示词变更在每个受影响的组合类别中产生一个面向行的 Markdown diff;工具描述变更在每个类别中产生一个结构化 JSON diff;普通行为 fixture 不受影响。会话 fixture 对省略的内容显示 token,运行时一致性守卫使每个拆分固定场景对其类别内的所有会话具有权威性。每个固定场景携带两个生成的、换行规范化的 sidecar 文件。
|
||||
@@ -0,0 +1,6 @@
|
||||
# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write
|
||||
2026-07-08-shared-acp-snapshot-package.md: c378222804251761a1b04f59c35799a97a1525f1
|
||||
2026-07-08-shared-acp-snapshot-package.zh.md: cda7a578d4643800736ff159a6427d3a0e3e0fae
|
||||
@@ -0,0 +1,38 @@
|
||||
# Agent Note: Extract the ACP snapshot suite into a support package
|
||||
|
||||
Status: implemented
|
||||
|
||||
## Problem
|
||||
|
||||
The ACP snapshot tier ([snapshot Agent Note](2026-06-19-acp-snapshot-tests.md)) was built from three modules living inside one example's test directory: `snapshot-harness.ts` (boot the real bin subprocess, drive it over ACP JSON-RPC, harvest the persisted logs), `snapshot-normalize.ts` (the pure expected-output normalizers), and the ~150-line scenario body plus fixture guards in `acp.snapshot.ts` (record/replay modes, the stdout expected-output and log comparisons, the pinned-header uniformity guard, the orphan/required-file/single-pin meta-tests).
|
||||
|
||||
A second ACP example wanting snapshot coverage — the sandbox/approval composition is the immediate consumer — could only copy those modules, forking exactly the logic that must not drift: record write-back, header scrubbing, child-session harvest ordering. The spawn/client glue was also triplicated across `acp.e2e.ts`, `hooks.e2e.ts`, and the harness. Location decided test rigor: the per-file 100% coverage gate measures `packages/*/*/src` only, so none of this machinery was measured — the same gap that had moved `dsh-llm-replay` out of `examples/` into [packages/support](../../../../packages/support/README.md). And the harness's ACP client hardcoded `requestPermission → cancelled`, so an approval round-trip — the headline behavior of the sandbox composition — could not be expressed at the snapshot tier at all.
|
||||
|
||||
## Decision
|
||||
|
||||
The machinery lives in [`packages/support/acp-snapshot`](../../../../packages/support/acp-snapshot/README.md) (`@deepseek-ai/dsh-acp-snapshot`); an example's `*.snapshot.ts` is its scenario table, its agent paths, and one factory call, over its own `snapshots/` fixtures and `cordis.snapshot.yml` overlay ([single-source replay config](2026-07-04-single-source-acp-replay-config.md)). Reading `DSH_SNAPSHOT` stays at that edge — the library takes a resolved `mode`.
|
||||
|
||||
**`src/launcher.ts`** — `launchAcpTestAgent` owns the common unbuilt-process boundary: absolute tsx loader resolution, `TSX_TSCONFIG_PATH`, isolated harness homes, stdio wiring, a raw-byte stdout tee, stderr and update capture, fail-closed permission fallback, update waiters, and graceful or signalled shutdown. Snapshot scenarios and ordinary e2e suites supply the same `AgentUnderTest` (`binScript`, `configPath`, `tsconfigPath`); a test that plays a user supplies only its permission handler. The ACP and hook e2e suites plus the sandbox/approval e2e suite use this launcher instead of rebuilding the SDK client boundary.
|
||||
|
||||
**`src/harness.ts`** — `runScenario` and the input-script/result types layer deterministic steps, temp workspaces, snapshot environment, and persisted-log harvest over the launcher. Its `session/request_permission` handler consumes an optional `InputScript.permissionAnswers` FIFO queue, each entry selecting by option **kind** (ids are agent-issued randoms a committed script cannot know; kinds are the ACP-stable vocabulary, mapped to the offered `optionId` at answer time); an absent or exhausted queue answers `cancelled`, and a kind the request never offered rejects the run — the agent itself is answered `cancelled`, so the scenario bug fails the harness rather than being absorbed as an agent-side denial. This is what lets an approval suite drive allow/reject round-trips deterministically from `input.json`.
|
||||
|
||||
**`src/normalize.ts`** — the pure normalizers, hook-free by policy: when a future event carries a new volatile field (an approval duration, say), the shared normalizer learns it in the same change, keeping one home for what "normalized" means rather than per-suite scrub extensions.
|
||||
|
||||
**`src/suite.ts`** — the `Scenario` type and `defineAcpSnapshotSuite(options)`, registering the per-scenario compares, record/refresh fixture write-back, the header pin with its live uniformity guard, and the fixture guard block (no orphan scenario dirs, required files present, exactly one pin per class, every JSONL a `scrubSystemPrompts` fixed point, non-pinning fixtures also `scrubRequestHeaders` fixed points). Refresh expands packed timing envelopes before aligning existing volatile event times, so switching between packed and unpacked layouts cannot shift later records; fresh chunk-fragment arrays remain authoritative because their boundaries are replay behavior. A scenario directory's `session.jsonl` plus contiguous `session.<n>.jsonl` siblings are its ordered primary/child inventory, so the scenario table declares policy without duplicating a child count. The pinned-header contract ([pinned-header Agent Note](2026-07-06-pin-request-header-content-in-one-scenario.md)) is per-suite: each header class flags exactly one `pinsHeader` scenario, whose `system-prompt.expected.md` and JSONL tool list split the composed header into reviewable artifacts; the uniformity guard compares both against every live header in that class. A pinning scenario declares any legitimate changed-header count, and its Markdown artifact records every full changed prompt. The pure helpers (`sessionFixtureNames`, `fixtureContext`, `normalizedHeaders`, `normalizedSystemPrompts`, `formatSystemPromptSnapshot`, `headerChangeCount`) are exported from the module for direct unit coverage.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
- **Copy the modules into each example** — the fork this Agent Note exists to prevent: the record/guard logic is exactly the code that must stay byte-identical across suites, and examples are outside the coverage gate, so each copy is also unmeasured.
|
||||
- **A shared module directory under `examples/`** — keeps the code outside the coverage gate and forces relative imports across example boundaries, against the package-name import convention; `examples/` leaves stay thin by design.
|
||||
- **A `/testing` subpath export of `dsh-acp-demo`** — couples test infrastructure into a product package's surface and dependency set; `packages/support/` exists precisely for real-but-lower-compatibility dev/test packages, with `dsh-llm-replay` as the precedent this package completes.
|
||||
- **Export raw test-body functions instead of a suite factory** — each example would re-own the `describe`/`it` skeleton (~80 lines of registration boilerplate per suite) for no flexibility gain; the factory keeps consumers to a scenario table plus one call, and the exported pure helpers preserve unit-testability inside the factory design.
|
||||
- **An injectable ACP `Client` factory instead of declarative `permissionAnswers`** — maximally flexible, but it leaks SDK client construction to every consumer and reopens per-example drift in exactly the layer being unified; a declarative queue keeps `input.json` the single scripting surface and compatible with expected-output normalization.
|
||||
- **Generalize beyond ACP (a transport-agnostic snapshot harness)** — no second transport exists; the harness is ACP-shaped end to end (SDK client, JSON-RPC frames, `session/update` waiters), and a speculative abstraction would be a seam split ahead of any consumer.
|
||||
|
||||
## Testing
|
||||
|
||||
Extraction parity was proven mechanically: after the move, `pnpm run test:snapshot` matched the base commit's result with zero byte changes under `examples/acp-agent/tests/snapshots/`. The package's `src/` holds per-file 100% statements/branches/functions/lines under the gating unit run, driven through the real launcher by a scripted fake ACP bin (`tests/fixtures/fake-acp-agent.ts`, behavior scripted per scenario via a `behavior.json` beside the fixture): `harness.spec.ts` directly covers launcher defaults, captures, update waiting, shutdown, and environment/config variants, then covers every scenario step op, both expect-error arms, the permission queue (selection, fallback, impossible-click), workspace seeding, and the harvest ordering/noise/fallback branches; `suite.spec.ts` runs the factory for real at collection time — a replay suite over committed synthetic fixtures and a record suite over a temp copy (write-back never touches the committed tree; `ACP_SNAPSHOT_SPEC_BOOTSTRAP=1` re-bootstraps it) — plus direct cases for the pure helpers. The fake bin substitutes the `session/new` cwd, not `process.cwd()`, into scripted logs, matching what the real bin's header carries (darwin realpaths `/var/folders/…` to `/private/var/folders/…`).
|
||||
|
||||
## Consequences
|
||||
|
||||
A new example gets the whole snapshot tier from a scenario table plus fixtures, while an ordinary ACP e2e gets the same tested process/client boundary from one launcher call. The costs: `suite.ts` imports vitest, so the package entry is importable only inside a vitest run — a shape no other package has, stated in its README; and each suite pins its own ~8 KB header fixture (a genuinely distinct composition deserves its own pin; an identical one would be caught by that suite's uniformity guard).
|
||||
@@ -0,0 +1,38 @@
|
||||
# RFC: 将 ACP 快照套件提取为支持包
|
||||
|
||||
Status: implemented
|
||||
|
||||
[English](2026-07-08-shared-acp-snapshot-package.md) | 中文
|
||||
|
||||
## 问题
|
||||
|
||||
ACP 快照层([快照 RFC](2026-06-19-acp-snapshot-tests.md))由位于某个示例测试目录中的三个模块构成:`snapshot-harness.ts`(启动真实 bin 子进程,通过 ACP JSON-RPC 驱动它,收集持久化日志)、`snapshot-normalize.ts`(纯粹的 golden 规范化器),以及 `acp.snapshot.ts` 中约 150 行的场景主体加 fixture(测试前置数据)守卫(record/replay 模式、stdout-golden 与日志比对、pinned-header 一致性守卫、orphan/required-file/single-pin 元测试)。
|
||||
|
||||
第二个 ACP 示例只能复制 record、规范化和收集逻辑,而这些逻辑必须保持一致。`examples/` 下的代码也不在包(package)覆盖率门禁范围内,且原始 harness 只能取消权限请求。共享包使这些机制纳入度量,并允许场景脚本化地提供审批答案。
|
||||
|
||||
## 决策
|
||||
|
||||
这些机制位于 [`packages/support/acp-snapshot`](../../../../packages/support/acp-snapshot/README.md)(`@deepseek-ai/dsh-acp-snapshot`);示例的 `*.snapshot.ts` 只包含场景表、agent 路径和一次工厂调用,依赖自己的 `snapshots/` fixture 与 `cordis.snapshot.yml` overlay([单源 replay 配置](2026-07-04-single-source-acp-replay-config.md))。读取 `DSH_SNAPSHOT` 留在边缘层——库接收的是已解析的 `mode`。
|
||||
|
||||
**`src/harness.ts`** 提供 `runScenario` 及其脚本/结果类型,以 agent 的 bin 和配置路径为参数。权限答案构成一个 FIFO 队列,以稳定的 option kind(而非随机的 option id)为键。缺少答案时取消该请求;不可用的 kind 取消 agent 请求并使场景失败。
|
||||
|
||||
**`src/normalize.ts`** 是纯规范化器,按策略不含钩子:当未来某个事件携带新的易变字段(例如审批耗时),共享规范化器在同一个变更中学会它,保持「规范化」的含义只有一个归属,而非各套件各自扩展清洗逻辑。
|
||||
|
||||
**`src/suite.ts`** 提供 `Scenario` 类型与 `defineAcpSnapshotSuite(options)`,注册逐场景比对、record/refresh 的 fixture 回写、header pin 及其实时一致性守卫,以及 fixture 守卫块(无 orphan 场景目录、必需文件齐全、每个 class 恰好一个 pin、每个 JSONL 是 `scrubSystemPrompts` 的不动点、非 pinning fixture 也是 `scrubRequestHeaders` 的不动点)。pinned-header 契约([pinned-header RFC](2026-07-06-pin-request-header-content-in-one-scenario.md))按套件划分:每个 header class 恰好标记一个 `pinsHeader` 场景,其 `system-prompt.golden.md` 与 JSONL 工具列表将组合后的 header 拆分为可评审的产物;一致性守卫将二者与该 class 中每个实时 header 进行比对。纯辅助函数(`childFixturePaths`、`fixtureContext`、`normalizedHeaders`、`normalizedSystemPrompts`、`formatSystemPromptSnapshot`、`headerDeltaCount`)从模块导出,以便直接进行单元覆盖。
|
||||
|
||||
## 曾考虑的替代方案
|
||||
|
||||
- **将模块复制到每个示例中**:正是本 RFC 要防止的 fork。record/守卫逻辑恰恰是必须在各套件间保持逐字节一致的代码,而示例不在覆盖率门禁范围内,因此每份副本也无法被度量。
|
||||
- **在 `examples/` 下建共享模块目录**:代码仍在覆盖率门禁之外,且需要跨示例边界的相对导入,违反包名导入约定;`examples/` 的叶子节点按设计应保持轻薄。
|
||||
- **`dsh-acp-demo` 的 `/testing` 子路径导出**:将测试基础设施耦合到产品包的对外服务接口与依赖集中;`packages/support/` 的存在正是为了真实但兼容性承诺较低的开发/测试包,`dsh-llm-replay` 是先例,本包与之配套。
|
||||
- **导出原始测试体函数而非套件工厂**:每个示例将重新拥有 `describe`/`it` 骨架(每套件约 80 行注册样板),却无灵活性收益;工厂使消费方只需一张场景表加一次调用,而导出的纯辅助函数在工厂设计内保留了可单元测试性。
|
||||
- **可注入的 ACP `Client` 工厂,而非声明式 `permissionAnswers`**:灵活性最大,但将 SDK 客户端构造泄露给每个消费方,并在正被统一的层面重新引入逐示例漂移;声明式队列使 `input.json` 成为唯一的脚本化界面,且可被 golden 规范化。
|
||||
- **泛化到 ACP 之外(传输无关的快照 harness)**:不存在第二种传输方式;harness 端到端都是 ACP 形态(SDK 客户端、JSON-RPC 帧、`session/update` 等待器),推测性的抽象将是一个超前于任何消费方的 seam 拆分。
|
||||
|
||||
## 测试
|
||||
|
||||
提取保留了所有既有 ACP golden 字节。包的 `src/` 通过脚本化的 ACP 子进程达到逐文件 100% 覆盖率:harness 测试覆盖每个步骤操作、两条预期错误分支、权限选择/回退/不可能选项、环境变量转发、workspace 种子注入、收集排序/噪声/回退;suite 测试对已提交的合成 fixture 执行 replay,并对临时副本执行 record,同时覆盖纯辅助函数。两个结构上不可达的守卫保留了有理由的覆盖率排除。fake agent 将 `session/new` 的 cwd 替换到日志中,包括 Darwin 的 `/var` realpath 行为,与真实 bin 一致。
|
||||
|
||||
## 后果
|
||||
|
||||
新示例只需一张场景表加 fixture 即可获得完整快照层——sandbox 分支从 master 合入后添加自己的套件(自己的 pin 场景、自己的 overlay、通过 `test:snapshot:record` 生成 fixture、通过 `permissionAnswers` 提供审批答案)。代价:`suite.ts` 导入 vitest,因此该包只能在 vitest 运行中导入——这是其他包没有的形态,已在其 README 中声明;每个套件 pin 自己约 8 KB 的 header fixture(真正不同的组合值得拥有自己的 pin;相同的组合会被该套件的一致性守卫捕获);e2e launcher 的重复仍然存在(`TODO(acp-test-harness)`)——当该迁移落地时,harness 即为提取目标。
|
||||
@@ -0,0 +1,6 @@
|
||||
# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write
|
||||
2026-07-18-tui-terminal-state-snapshots.md: 8e86588f69fdb9d615232252ecf57309d440f1cd
|
||||
2026-07-18-tui-terminal-state-snapshots.zh.md: b70a46830f44e9da663e30745fcdb7ad281592da
|
||||
@@ -0,0 +1,70 @@
|
||||
# Agent Note: Snapshot semantic terminal state for the TUI
|
||||
|
||||
Status: implemented
|
||||
|
||||
English | [中文](2026-07-18-tui-terminal-state-snapshots.zh.md)
|
||||
|
||||
## Problem
|
||||
|
||||
The TUI is a stateful renderer. Its user-visible result depends on ANSI parsing, differential frames, wrapping, scrollback, viewport position, terminal width, focus, cursor state, and each tool's presentation intent. Unit tests that collect `Terminal.write()` fragments can prove event handling, but they cannot prove the final screen a terminal displays. The same screen may also be emitted through different write fragments, so pinning those fragments creates false regressions.
|
||||
|
||||
Component-line snapshots stop before ANSI reaches a terminal and miss cursor movement, clearing, styling, overlay composition, and reflow. Raster screenshots include font and platform rendering noise that is unrelated to the TUI contract. A completed flow built by directly appending plausible session events has another blind spot: it proves the renderer accepts those shapes, not that the production agent loop and tool implementations produce them.
|
||||
|
||||
The TUI therefore needs a deterministic, reviewable representation of terminal state, recorded model journeys that execute the real downstream stack, and a smaller test at the real process and PTY boundary.
|
||||
|
||||
## Decision
|
||||
|
||||
TUI coverage has four complementary layers:
|
||||
|
||||
1. `packages/ui/tui/tests/tui.spec.ts` tests event mapping, input routing, disposal, and error behavior directly.
|
||||
2. `packages/ui/tui/tests/tui.snapshot.ts` mounts the production TUI against a headless terminal emulator for transient states that a completed session log cannot retain: in-flight streaming, pending tool calls, overlays, expansion, compaction reflow, errors, and shutdown.
|
||||
3. `examples/tui-agent/tests/tui.snapshot.ts` replays committed JSONL session logs through the production agent loop and real tools, then compares the resulting semantic terminal state.
|
||||
4. `examples/tui-agent/tests/tui-keyless-smoke.e2e.ts` boots the real Loader composition in a PTY, drives a scripted conversation through streaming and `ask_user_question`, and verifies startup, input, exit, failure reporting, and terminal restoration.
|
||||
|
||||
The runnable TUI has its own `examples/tui-agent` leaf beside the Headless and ACP leaves. It owns the interactive coding backends and tools directly and loads `@deepseek-ai/dsh-tui-demo`; TUI snapshots and PTY tests live with that leaf. The [redundant-agent removal](../simplification/2026-07-20-remove-stdio-and-echo-agents.md) owns this consolidation.
|
||||
|
||||
### Recorded-session replay
|
||||
|
||||
Each example-level scenario directory owns `session.jsonl`, optional child logs `session.<n>.jsonl`, and `terminal.expected.txt`. The primary log supplies user-authored `user/message` prompts and the recorded `assistant/chunk` sequence. `dsh-llm-replay` derives one model-call script per session, binds child logs to fresh child sessions, and is the only mocked boundary. The agent loop, bash and filesystem implementations, Code Mode worker, subagent provider, workflow worker, Cordis tools, presenters, and TUI are production implementations.
|
||||
|
||||
The suite rejects a journey when its tool-call sequence differs, an expected event count is missing, a tool result is an error, a turn ends in error, a workflow lifecycle is incomplete, or the live child-session count differs from the fixture set. These assertions prevent an attractive terminal expected output from hiding a failed or bypassed production path.
|
||||
|
||||
The live-model fixtures use `DSH_SNAPSHOT=record`; record mode rewrites their primary and child JSONL logs and terminal expected outputs. The deterministic Cordis toolchain keeps an authored complete JSONL script because reliably coercing a live model through five exact tool boundaries and two children is not a stable recording contract. `DSH_SNAPSHOT=refresh` replays every committed script keylessly and rewrites only derived terminal expected outputs. Plain replay compares without writing, and unknown mode values fail loud.
|
||||
|
||||
### Semantic terminal projection
|
||||
|
||||
The package-local `HeadlessTerminal` implements the same pi-tui `Terminal` interface as the process terminal and feeds every ANSI write into the pinned `@xterm/headless` parser. Snapshot code waits for synchronized frames to quiesce before reading state, so a checkpoint represents a completed screen rather than a timer-dependent write prefix.
|
||||
|
||||
Each expected output projects dimensions, active-buffer and viewport coordinates, lifecycle and cursor state, rows, wrap markers, and non-default style ranges into text. Scroll-heavy cards capture the used buffer; overlays capture the visible viewport. Text and style remain separate so a reviewer can distinguish content changes from presentation changes without decoding ANSI bytes.
|
||||
|
||||
Every checkpoint enforces theme independence across the complete terminal state: no RGB colors, no palette entries beyond ANSI 0–15, and no explicit background colors. Reverse video remains valid for selection because it uses terminal defaults. Both suites own closed inventories that reject missing scenarios, missing checkpoints, and orphaned expected output files.
|
||||
|
||||
### Required scenario matrix
|
||||
|
||||
| Layer | Scenario | Contract pinned |
|
||||
|---|---|---|
|
||||
| Recorded journey | Multi-turn conversation | Recorded reasoning/text chunks, two input turns, retained history, token totals, and idle editor state |
|
||||
| Recorded journey | Todo plan | Real `todo_write` execution, result card, and persistent plan rendering |
|
||||
| Recorded journey | Bash terminal card | Real local executor output, description, exit status, and completed terminal card |
|
||||
| Recorded journey | Parallel filesystem reads | Two calls from one assistant message, real file contents, ordering, and separate completed cards |
|
||||
| Recorded journey | Code Mode | Real `run_code` worker execution, two `tool/code-dispatch` events, captured program output, and completed card |
|
||||
| Recorded journey | Dynamic workflow | Real workflow worker, phase lifecycle, replayed child session, structured return value, and completed card |
|
||||
| Recorded journey | Cordis dynamic toolchain | Real mount, Code Mode inspect, direct subagent, workflow child, unmount, and all production presenters |
|
||||
| Transient state | Streaming and pending advanced calls | In-flight reasoning/text plus pending Code Mode, workflow, and Cordis cards that disappear from completed logs |
|
||||
| Transient state | Cards, interaction, layout, failure, and shutdown | Collapsed/expanded card families, question validation, compaction replacement, resize reflow, help/errors, cursor restoration, and terminal stop |
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
- **Snapshot raw terminal writes** — rejected because differential rendering may change write boundaries without changing the screen, while cursor and clear sequences are unreadable in review.
|
||||
- **Snapshot component render lines before terminal output** — rejected because it does not test ANSI parsing, cursor movement, overlays, viewport behavior, or independent components in one frame.
|
||||
- **Build every completed flow by appending session events** — rejected because a hand-authored event sequence can drift from the agent loop, tool execution, child-session binding, or worker behavior while its presentation test stays green. Direct event construction remains limited to transient renderer states.
|
||||
- **Reuse ACP stdout expected outputs as the TUI oracle** — rejected because a recorded model journey is transport-neutral but its presentation is not. TUI scenarios own terminal expected outputs while using the same JSONL replay vocabulary.
|
||||
- **Commit raster screenshots** — rejected because fonts, glyph metrics, antialiasing, and host terminal themes make them platform-sensitive and make semantic style changes difficult to review.
|
||||
- **Use only PTY end-to-end tests** — rejected because raw PTY output is a stream of historical drawing operations, not queryable final state. PTY tests retain the real Loader/input/teardown boundary, while the emulator owns broad state coverage.
|
||||
|
||||
## Consequences
|
||||
|
||||
- Completed advanced snapshots now fail when the real Code Mode, workflow, subagent, filesystem, bash, or Cordis path breaks, rather than accepting a fabricated result event.
|
||||
- TUI visual regressions produce readable cell-and-style diffs, while JSONL fixtures retain the exact model chunks that made the production path execute.
|
||||
- The emulator uses xterm's proposed buffer API. An xterm upgrade requires rerunning and reviewing the semantic projection; terminal-specific behavior still needs the PTY smoke.
|
||||
- Expected outputs deliberately encode wrapping and viewport behavior at fixed sizes. Intentional layout changes use keyless refresh, while model-journey changes use record mode and review both JSONL and terminal diffs.
|
||||
@@ -0,0 +1,70 @@
|
||||
# Agent Note: TUI 语义终端状态快照
|
||||
|
||||
Status: implemented
|
||||
|
||||
[English](2026-07-18-tui-terminal-state-snapshots.md) | 中文
|
||||
|
||||
## 问题
|
||||
|
||||
TUI 是有状态的渲染器。用户最终看到的结果取决于 ANSI 解析、差分帧、换行、回滚缓冲、视口位置、终端宽度、焦点、光标状态,以及各工具的呈现意图。收集 `Terminal.write()` 片段的单元测试可以验证事件处理,却无法验证终端最终显示的画面。同一画面也可能由不同的写入片段产生,因此固定这些片段会制造误报。
|
||||
|
||||
组件行快照止于 ANSI 进入终端之前,无法覆盖光标移动、清屏、样式、浮层组合和重排。栅格截图会带入与 TUI 契约无关的字体和平台渲染噪声。直接追加看似合理的会话事件来构造完整流程还存在另一处盲区:这种测试只能证明渲染器接受这些数据形态,无法证明生产环境的 agent loop(智能体循环)和工具实现会生成这些事件。
|
||||
|
||||
因此,TUI 既需要确定、便于评审的终端状态表示,也需要通过已录制模型流程执行真实下游组件,并保留一项范围更小、覆盖真实进程与 PTY 边界的测试。
|
||||
|
||||
## 决策
|
||||
|
||||
TUI 覆盖分为四个互补层次:
|
||||
|
||||
1. `packages/ui/tui/tests/tui.spec.ts` 直接测试事件映射、输入路由、资源释放和错误行为。
|
||||
2. `packages/ui/tui/tests/tui.snapshot.ts` 将生产 TUI 挂载到无界面终端模拟器,覆盖完整会话日志无法保留的瞬态:进行中的流式输出、待完成工具调用、浮层、展开状态、压缩重排、错误和关闭过程。
|
||||
3. `examples/tui-agent/tests/tui.snapshot.ts` 通过生产 agent loop 和真实工具回放已提交的 JSONL 会话日志,再比较生成的语义终端状态。
|
||||
4. `examples/tui-agent/tests/tui-keyless-smoke.e2e.ts` 在 PTY 中启动真实 Loader 组合,驱动一段经过流式输出和 `ask_user_question` 的脚本化会话,并验证启动、输入、退出、失败报告和终端恢复。
|
||||
|
||||
可运行 TUI 在 `examples/tui-agent` 中拥有独立叶节点,与 Headless 和 ACP 叶节点并列。它直接拥有交互式 coding 后端与工具,并加载 `@deepseek-ai/dsh-tui-demo`;TUI 快照和 PTY 测试也归属这个叶节点。[移除重复 agent 的决策](../simplification/2026-07-20-remove-stdio-and-echo-agents.md)负责此次整合。
|
||||
|
||||
### 已录制会话回放
|
||||
|
||||
每个示例级场景目录都包含 `session.jsonl`、可选的子会话日志 `session.<n>.jsonl`,以及 `terminal.expected.txt`。主日志提供用户来源的 `user/message` 提示词和已录制的 `assistant/chunk` 序列。`dsh-llm-replay` 为每个会话派生一份模型调用脚本,并将子日志绑定到新建的子会话;这是测试中唯一的 mock 边界。agent loop、bash 与文件系统实现、Code Mode worker、subagent 提供方、工作流 worker、Cordis 工具、呈现器和 TUI 都使用生产实现。
|
||||
|
||||
如果工具调用顺序不符、预期事件数量不足、工具结果报错、轮次以错误结束、工作流生命周期不完整,或者实时子会话数量与 fixture(测试前置数据)集合不一致,测试都会失败。即使终端预期输出表面正确,这些断言也能阻止失败或被绕过的生产路径混入结果。
|
||||
|
||||
真实模型 fixture 通过 `DSH_SNAPSHOT=record` 更新;录制模式会重写其主会话与子会话 JSONL 日志以及终端预期输出。确定性的 Cordis 工具链保留一份人工编写的完整 JSONL 脚本,因为要求真实模型稳定经过五个指定工具边界和两个子会话并不是可靠的录制契约。`DSH_SNAPSHOT=refresh` 会无密钥回放所有已提交脚本,并且只重写派生的终端预期输出。普通回放只比较而不写入,未知模式值会快速失败。
|
||||
|
||||
### 语义终端投影
|
||||
|
||||
包内的 `HeadlessTerminal` 实现与进程终端相同的 pi-tui `Terminal` 接口,并把每次 ANSI 写入交给固定版本的 `@xterm/headless` 解析器。读取状态前,快照代码会等待同步帧稳定,因此每个检查点表示已经完成的画面,而不是依赖计时的写入前缀。
|
||||
|
||||
每份预期输出把终端尺寸、活动缓冲区和视口坐标、生命周期与光标状态、各行、换行标记以及非默认样式区间投影为文本。滚动内容较多的卡片捕获已使用缓冲区;浮层捕获可见视口。文本和样式相互分离,评审人无需解码 ANSI 字节即可区分内容变化与呈现变化。
|
||||
|
||||
每个检查点还会对完整终端状态强制执行主题无关性:禁止 RGB 颜色、禁止 ANSI 0–15 以外的调色板项,也禁止显式背景色。选择行使用终端默认色进行反显,因此仍然有效。两套测试都拥有封闭清单,会拒绝缺失的场景、缺失的检查点和遗留预期输出文件。
|
||||
|
||||
### 必需场景矩阵
|
||||
|
||||
| 层次 | 场景 | 固定的契约 |
|
||||
|---|---|---|
|
||||
| 已录制流程 | 多轮会话 | 已录制的推理与文本分片、两轮输入、保留历史、token 总量和空闲编辑器状态 |
|
||||
| 已录制流程 | Todo 计划 | 真实 `todo_write` 执行、结果卡片和持久计划渲染 |
|
||||
| 已录制流程 | Bash 终端卡片 | 真实本地执行器输出、说明、退出状态和已完成终端卡片 |
|
||||
| 已录制流程 | 并行文件读取 | 同一条 assistant 消息中的两次调用、真实文件内容、顺序和两个独立完成卡片 |
|
||||
| 已录制流程 | Code Mode | 真实 `run_code` worker 执行、两条 `tool/code-dispatch` 事件、捕获的程序输出和已完成卡片 |
|
||||
| 已录制流程 | 动态工作流 | 真实工作流 worker、阶段生命周期、回放的子会话、结构化返回值和已完成卡片 |
|
||||
| 已录制流程 | Cordis 动态工具链 | 真实挂载、Code Mode 检查、直接 subagent、工作流子会话、卸载和全部生产呈现器 |
|
||||
| 瞬态 | 流式输出与待完成高级调用 | 进行中的推理和文本,以及完整日志中不会保留的待完成 Code Mode、工作流和 Cordis 卡片 |
|
||||
| 瞬态 | 卡片、交互、布局、失败和关闭 | 折叠与展开的卡片族、问题校验、压缩替换、尺寸重排、帮助与错误、光标恢复和终端停止 |
|
||||
|
||||
## 曾考虑的替代方案
|
||||
|
||||
- **快照原始终端写入**:不予采纳,因为差分渲染可能在画面不变时改变写入边界,而且光标与清屏序列难以评审。
|
||||
- **快照进入终端输出之前的组件渲染行**:不予采纳,因为它无法测试 ANSI 解析、光标移动、浮层、视口行为,也无法测试独立组件在同一帧中的相互作用。
|
||||
- **通过追加会话事件构造所有完整流程**:不予采纳,因为人工编写的事件序列可能与 agent loop、工具执行、子会话绑定或 worker 行为发生偏差,但呈现测试仍然保持绿色。直接构造事件只用于渲染器瞬态。
|
||||
- **复用 ACP stdout 预期输出作为 TUI 判定依据**:不予采纳,因为已录制模型流程与传输方式无关,其呈现方式却并非如此。TUI 场景使用同一套 JSONL 回放词汇,但拥有独立的终端预期输出。
|
||||
- **提交栅格截图**:不予采纳,因为字体、字形度量、抗锯齿和宿主终端主题会使结果依赖平台,也会增加语义样式变更的评审难度。
|
||||
- **只使用 PTY 端到端测试**:不予采纳,因为原始 PTY 输出是一系列历史绘制操作,而不是可查询的最终状态。PTY 测试保留真实 Loader、输入与清理边界,模拟器负责广泛的状态覆盖。
|
||||
|
||||
## 后果
|
||||
|
||||
- 当真实 Code Mode、工作流、subagent、文件系统、bash 或 Cordis 路径损坏时,已完成高级快照会失败,不会继续接受伪造的结果事件。
|
||||
- TUI 视觉回归会产生便于阅读的单元格和样式 diff,而 JSONL fixture 会保留触发生产路径的确切模型分片。
|
||||
- 模拟器使用 xterm 的拟议缓冲区 API。升级 xterm 时必须重新运行并评审语义投影;终端特有行为仍需由 PTY 冒烟测试覆盖。
|
||||
- 预期输出有意固定指定尺寸下的换行与视口行为。预期布局变更使用无密钥刷新;模型流程变更使用录制模式,并同时评审 JSONL 与终端 diff。
|
||||
@@ -0,0 +1,6 @@
|
||||
# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write
|
||||
2026-07-22-cross-platform-test-fixtures.md: 6217aabfdbe8f14f869004c8dafb7e19f4b7443a
|
||||
2026-07-22-cross-platform-test-fixtures.zh.md: 43942ec0468df822d04b39e318010c2b260c734f
|
||||
@@ -0,0 +1,33 @@
|
||||
# Agent Note: Keep supported-platform tests semantic
|
||||
|
||||
Status: implemented
|
||||
|
||||
English | [中文](2026-07-22-cross-platform-test-fixtures.zh.md)
|
||||
|
||||
## Problem
|
||||
|
||||
The unit and coverage suites run on Windows, macOS, and Linux, but a platform-neutral behavior can be hidden behind a platform-specific fixture. Literal POSIX paths become drive-relative paths on Windows, a hosted `file:` URI can be a valid UNC path there, and child-pipe closure or event-loop scheduling does not settle at the same point on every host. POSIX-only filesystem states such as FIFOs, executable mode bits, and directory search bits have no direct Windows fixture.
|
||||
|
||||
Treating fixture syntax as product behavior either reports false regressions or encourages production normalization that erases native path semantics.
|
||||
|
||||
## Decision
|
||||
|
||||
Tests of platform-neutral behavior construct absolute paths and `file:` URIs with the host's `node:path` and `node:url` APIs, then assert native absolute output or stable workspace-relative output as the contract requires. Invalid-URI fixtures use encodings rejected by `fileURLToPath()` on every supported platform.
|
||||
|
||||
Transport-failure tests inject the connection's message writer and deliver the same asynchronous write callback error that a real Node stream would report. The production writer still writes framed messages to child stdin. This keeps a real child alive while the test deterministically distinguishes transport failure from process exit without reaching into platform-specific pipe handles.
|
||||
|
||||
Language-server teardown targets the whole descendant tree through a negative process-group id on POSIX and synchronous `taskkill /T /F` on Windows. Windows suppresses only taskkill's already-absent-tree status; command, permission, and other tree-kill failures remain teardown failures. A read-only provider query retries once only when its selected pooled transport fails before or during that query; errors from a still-live server are not replayed. Terminal tests wait for their observable rendered output instead of assuming one event-loop turn is sufficient.
|
||||
|
||||
Tests for a genuinely POSIX-only primitive use a narrow Windows exclusion on that case. Adjacent cross-platform cases continue to pin non-regular file rejection, unavailable command rejection, and inaccessible working-directory rejection. Supported Windows paths remain inside the per-file coverage gate rather than being excluded with their test files.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
**Normalize all paths and URIs to POSIX strings.** This would make assertions uniform but would change correct Windows behavior: external paths are native absolute paths, UNC file URIs are valid, and configured homes resolve through the host path rules.
|
||||
|
||||
**Manipulate child-pipe internals until a write fails.** CRT descriptors and libuv handles have different ownership across hosts and Node versions, so this would test undocumented fixture machinery instead of the connection's write-failure contract.
|
||||
|
||||
**Skip whole files or packages on Windows.** Broad exclusions would hide supported behavior. Only the individual fixture whose state cannot exist on Windows is excluded; the surrounding contract remains covered.
|
||||
|
||||
## Consequences
|
||||
|
||||
Portable fixtures are slightly more explicit because expected paths derive from shared native constants and transport failures enter through a narrow writer seam. Platform-only exclusions require a neighboring cross-platform assertion for the product behavior they support. Windows teardown depends on the host `taskkill` command after graceful protocol shutdown has failed; a successful synchronous result keeps disposal bounded and makes descendant exit observable before cleanup returns, while a failed tree kill remains visible to the disposer.
|
||||
@@ -0,0 +1,33 @@
|
||||
# Agent Note: 让受支持平台的测试聚焦语义
|
||||
|
||||
Status: implemented
|
||||
|
||||
[English](2026-07-22-cross-platform-test-fixtures.md) | 中文
|
||||
|
||||
## 问题
|
||||
|
||||
单元测试与覆盖率测试套件会在 Windows、macOS 和 Linux 上运行,但平台无关行为可能被平台特有的 fixture(测试前置数据)掩盖。字面 POSIX 路径在 Windows 上会变成相对于驱动器的路径;带主机名的 `file:` URI 在 Windows 上可能是有效的 UNC 路径;子进程管道关闭或事件循环调度在不同宿主上的稳定时点也不一致。FIFO、可执行模式位和目录搜索权限位等仅存在于 POSIX 的文件系统状态,在 Windows 上没有可直接构造的 fixture。
|
||||
|
||||
把 fixture 语法当成产品行为,要么会误报回归,要么会促使生产代码引入抹去原生路径语义的归一化。
|
||||
|
||||
## 决策
|
||||
|
||||
测试平台无关行为时,使用宿主的 `node:path` 和 `node:url` API 构造绝对路径与 `file:` URI,再根据契约要求断言原生绝对输出或稳定的工作区相对输出。无效 URI fixture 使用一种在所有受支持平台上都会被 `fileURLToPath()` 拒绝的编码形式。
|
||||
|
||||
传输故障测试会注入连接的消息写入器,并传入与真实 Node 流相同的异步写入回调错误。生产写入器仍会把分帧消息写入子进程 stdin。这种方式让真实子进程保持存活,使测试无需触及平台特有的管道句柄,也能确定性地区分传输故障与进程退出。
|
||||
|
||||
语言服务器的资源清理会终止整棵后代进程树:POSIX 使用负数进程组 ID,Windows 同步执行 `taskkill /T /F`。Windows 只会忽略 taskkill 返回的「进程树已经不存在」状态;命令执行失败、权限错误及其他终止进程树的失败仍属于资源清理失败。只读的提供方查询仅在选定的池化传输于该次查询开始前或执行期间失效时重试一次;服务器仍存活时返回的错误不会重放。终端测试会等待可观察的渲染输出,不假设一次事件循环轮转已经足够。
|
||||
|
||||
对于真正仅存在于 POSIX 的原语,测试只在该用例上排除 Windows。相邻的跨平台用例仍会固定拒绝非普通文件、不可用命令和无法访问的工作目录的行为。Windows 上受支持的路径仍受逐文件覆盖率门禁约束,不会随测试文件一起排除。
|
||||
|
||||
## 曾考虑的替代方案
|
||||
|
||||
**将所有路径和 URI 归一化为 POSIX 字符串。**这会使断言保持一致,但也会改变正确的 Windows 行为:外部路径是原生绝对路径,UNC 文件 URI 有效,而且已配置的主目录会按照宿主路径规则解析。
|
||||
|
||||
**操纵子进程管道内部状态,直至写入失败。**CRT 描述符与 libuv 句柄在不同宿主和 Node 版本上的所有权不同,因此这种做法测试的是未文档化的 fixture 机制,而非连接的写入失败契约。
|
||||
|
||||
**在 Windows 上跳过整个测试文件或包。**过宽的排除会隐藏受支持的行为。只排除无法在 Windows 上构造相应状态的单项 fixture;相关契约仍保持覆盖。
|
||||
|
||||
## 后果
|
||||
|
||||
可移植 fixture 需要更显式地构造,因为预期路径要从共享的原生常量派生,传输故障则通过狭窄的写入器 seam 注入。仅适用于特定平台的排除项必须配有相邻的跨平台断言,以继续覆盖相应的产品行为。协议级优雅关停失败后,Windows 上的资源清理依赖宿主的 `taskkill` 命令;命令同步执行成功时,dispose 的完成边界明确,并确保清理返回前即可观察到后代进程退出;若进程树终止失败,dispose 的调用方仍能观察到该失败。
|
||||
Reference in New Issue
Block a user