Eight proposals grouped by category, each with problem statement, concrete plan, and risks: property-based testing over the protocol-shaped core (chunk streams, event logs, schema DSL); mutation testing as the counterweight to the 100%-coverage gate; deterministic tests + a universal replay-invariant fixture + nightly race stress; architectural conformance (dependency-cruiser rules and the LlmAdapter conformance kit); runtime arg validation at the model boundary with a structured error taxonomy and dev-mode invariants; doc-sync enforcement (typechecked doc snippets, API reports); supply-chain checks and nightly vendor-drift verification against the manifest; and deep-readonly public surfaces (logged-vs-in-flight mutability boundary). AGENTS.md points at docs/adr and docs/rfc.
39 lines
1.5 KiB
Markdown
39 lines
1.5 KiB
Markdown
# RFC 002: Mutation testing as the coverage counterweight
|
|
|
|
Status: proposed
|
|
|
|
## Problem
|
|
|
|
The per-file 100% coverage gate (ADR 0007) proves every line *executes* under
|
|
test — not that any assertion would notice if the line were wrong. Under
|
|
agent-written tests, coverage pressure can produce execution-without-assertion.
|
|
Mutation testing measures what coverage cannot: whether the suite *kills*
|
|
deliberately injected bugs.
|
|
|
|
## Proposal
|
|
|
|
Stryker (`@stryker-mutator/vitest-runner`) over `packages/*/src`:
|
|
|
|
- **PR-scoped incremental runs** (changed files only) as a CI job — fast
|
|
enough to gate merges once tuned.
|
|
- **Nightly full runs** with a tracked mutation score; start by recording,
|
|
then set the threshold at the observed baseline and ratchet upward (same
|
|
policy as coverage: thresholds only ever tighten).
|
|
- Surviving mutants are work items: an agent picks a survivor, writes the
|
|
killing test, repeats — a well-shaped autonomous loop.
|
|
- Equivalent mutants (provably behavior-preserving) get annotated exclusions
|
|
with reasons, mirroring the `/* v8 ignore */` policy.
|
|
|
|
## Plan
|
|
|
|
1. Add Stryker config scoped to one package (llm — smallest, most algorithmic)
|
|
and measure runtime.
|
|
2. Expand to all packages; record baseline scores in the config.
|
|
3. Wire the nightly job; add the incremental PR job once runtime is acceptable.
|
|
|
|
## Risks
|
|
|
|
Runtime: mutation testing is expensive; per-file 100% coverage helps (every
|
|
mutant is at least reached). If PR-scoped runs stay too slow, keep them
|
|
nightly-only and rely on the score ratchet.
|