Files
deepseek-harness/packages/session-persistence/session-persistence-sqlite
Tianyi Cui 774d460889 Expose audited hardcoded tunables as plugin config
The audit swept every packages/*/* plugin for the new AGENTS.md
convention (no hardcoded tunables in plugins) and exposes each finding
as a defaulted, validated Config field. Defaults are the previously
hardcoded values throughout, so no deployment or golden changes.

- tool-fs (had NO Config): readLimit, readMaxLineLength, readMaxBytes,
  readStreamMinSize. The caps thread through ReadToolCaps/ReadWindow —
  read-render already documented that the consumer applies the caps, so
  they become explicit per-request fields.
- tool-web: searchMaxResults (WEB_SEARCH_MAX_RESULTS stays as the
  schemastery default). Also fixes the stale GREP_LIMIT references in
  search.ts and the web-capability-seam RFC (no such constant exists).
- bash-local: graceMs (SIGTERM->SIGKILL escalation grace). The
  RunInternals.graceMs test seam is gone: graceMs is now a required
  SpawnSpec field filled from config, so tests exercise the real
  config path and the defaults live in exactly one place.
- subagent-acp: disposeEofGraceMs / disposeGraceMs. The AcpRunSpec
  fields become required for the same one-defaulting-layer reason.
- session-persistence-sqlite: journalMode ('wal' default; the
  rollback-journal modes serve filesystems where WAL's shared-memory
  files do not work, e.g. network mounts).
- hooks-claude + hooks-codex: stderrSummaryMaxChars for the persisted
  hook/result stderr summary. The duplicated summarize() helpers merge
  into hook-protocol's summarizeStderr(stderr, maxChars), beside the
  HookResultRecord field it feeds, with the bound parameterized the
  same way runHook's defaultTimeoutMs already is.
- compact-basic: charsPerToken for the token estimator (default 4, the
  English-text heuristic; CJK-heavy deployments need ~1-2 or compaction
  fires far too late). Also corrects the BasicCompactService class doc,
  which claimed defaults the required-field config never had.
- fs-local: deletes the dead STREAM_MIN_SIZE constant and the dead
  FsIoInternals.streamMinSize seam — the read-routing bound lives in
  the consumer (tool-fs), where it is now config. This is item 1 of
  the proposed prune-write-only-fs-surface RFC, annotated accordingly.

Every new field gets range validation (following the existing
assertPositiveFinite pattern), a README row, and tests covering the
configured behavior, the schema default, and load-time rejection.
2026-07-04 17:37:23 +08:00
..

@deepseek-ai/dsh-session-persistence-sqlite

A SQLite durable session-persistence backend — a second SessionPersistence implementation (session persistence), built to validate that the abstract seam and the shared runPersistenceContract suite are genuinely backend-agnostic. It satisfies the SAME contract as dsh-session-persistence-jsonl (append-only, contiguous-seq, lazy materialization, interrupted-turn close on load), expressed over node:sqlite rows instead of file bytes.

TODO: this backend talks to node:sqlite directly. If a cordis database service (cordis/db / a @cordisjs SQL driver plugin) is adopted, route through that instead of holding a raw DatabaseSync here — the contract surface (SessionPersistence) would not change, only the storage driver.

Storage model

Each SessionEvent maps 1:1 onto a row in an events table (session_id, seq, type, time, data, source_event_seqs, surface_op)data is the event payload as JSON text, so the row shape is the event verbatim (including assistant/chunk, keeping seq contiguous). The two TEXT columns source_event_seqs and surface_op are nullable; they store the event's optional surface-metadata fields (see session surface). Out-of-log metadata (SessionHeader) lives in a sessions row. A sessions row is written only by the first append — its existence is the lazy-materialization signal (list reports exactly the sessions that have a row), so no separate column is needed.

The repo targets Node ≥ 24 (the root engines field), which includes the stable node:sqlite module. The database opens with foreign_keys = ON (so ON DELETE CASCADE drops a session's events with its row) and the configured journal_mode (default wal; pick a rollback-journal mode like delete on filesystems where WAL's shared-memory files do not work, e.g. network mounts). The table-layout version is stored in PRAGMA user_version and checked on open: a fresh database is stamped with the current SCHEMA_VERSION; a database written by any other, incompatible build (a non-current user_version, older or newer) is rejected rather than opened against an unknown layout — there is no migration (unreleased software).

Contract semantics over rows

  • Append = a transaction. append runs BEGIN/COMMIT around the batch: it materializes the sessions row (if still lazy) and INSERTs every event, asserting the contiguous-seq contract first (the first event's seq must equal the stored next-seq). A mid-batch failure (a UNIQUE violation on a duplicated seq) rolls back entirely, so the stored log and the in-memory cursor stay consistent. (load() already balanced the stored log, so append never has to repair a crash tail.)
  • Lazy materialization. create() records intent in memory only — no row is written until the first append. A created-but-never-appended session has no sessions row, so it is absent from list() (which reports exactly the sessions that have a row).
  • Interrupted-turn close on load. load() reads every stored event ordered by seq and finds the longest seq-contiguous, parseable prefix — INCLUDING the real events of an interrupted final turn after the last turn/end (the loop only flushes at turn/end, so a process killed mid-turn leaves real, fully-written rows past it). A single turn can be huge in a long-horizon task, so those events are preserved, never truncated: load() CLOSES the orphaned turn by durably appending the minimal synthetic boundary events (an error tool/result for every assistant tool call left unanswered, a step/end if a step was open, then a turn/end carrying { kind: 'interrupted' }), inside one transaction that also DELETEs any never-fully-written torn tail row. load() is therefore mutating — after it the stored rows are balanced and the cursor is truthful, so the next append continues cleanly. The boundary (last turn/end, torn-tail detection) is computed from the seq/type columns so a malformed data in a torn tail row is never parsed (discarded, not unloadable). A parse error or seq gap inside the committed region (at or before the last real turn/end) makes the session unloadable. A session whose only turn never closed keeps its metadata row and stays present in list() — the same as the JSONL backend, whose file likewise survives a first append that never reached turn/end.

Configuration (schemastery)

interface Config {
  path: string   // SQLite database file path, or ':memory:' for an in-process DB
  journalMode?: 'wal' | 'delete' | 'truncate' | 'persist'   // journal_mode pragma; default 'wal'
}

Write path

Like the JSONL backend, the plugin also installs the session/event → buffer → session/flush drain: it snapshots each event when buffered (the live session.events object is mutable), persists a fork's seed once on session/created, keeps a per-session write cursor so a resumed session never re-appends stored events, and seeds existing live sessions on apply (HMR does not replay session/created). Dispose awaits every in-flight init + final drain and then closes the database, so no write lands after teardown.