Two CI-only test failures on the personal TUI/skills stack:
- packages/ui/tui/tests/tui.spec.ts anchored the dark-palette test at
process.cwd(), which is not guaranteed under $HOME; on CI runners the
prompt rendered an absolute path instead of the `~/` abbreviation. Anchor
cwd under homedir() so the assertion is deterministic.
- examples/tui-agent/tests/tui-keyless-smoke.e2e.ts asserted stale
dsh-customize / dsh-upgrade skill descriptions. Sync the expectations to
the bundled SKILL.md frontmatter.
Consolidates the personal dsh-tui customizations (module split into
components/session/extension, prompt template + running-glyph indicator,
copyable transcript, tool-card headers, timing placement, XML tool output,
status/footer rework) and ports upstream's model reasoning-effort selector
(Shift+Tab effort cycling, effort-aware /model, footer, and /status) onto
the personal module layout.
- subagent-sdk: the default registry name becomes `dsh-sdk` (the bare
`sdk` read ambiguously in configs); READMEs, config catalog, fixture,
and suites follow. The Loader fixture now omits providerName to exercise
the shipped default end to end.
- loader-composition.e2e: two full harness runtimes boot in sequence, so
the default 30s loader-smoke window times out under host load; raise the
subprocess deadline to 120s with matching vitest headroom (the
real-model.e2e precedent).
The web-fetch pinsHeader scenario landed on master after this branch, so its
tool-schemas.expected.json carried the old at-most-one-in_progress description
while replay assembles the new one. Refreshed keylessly with
DSH_SNAPSHOT=refresh; the Agent Note now records that every pinning scenario
carries its own copy of the description, so a branch changing it has to refresh
the pins that landed after it branched.
The local PTY readiness poll held its inferred_idle fallback for exactly
one pollIntervalMs after a prompt marker, so a bash foreground handoff
that lands on the silence boundary only wins the exact stdin_read
attribution when the kernel publishes it inside that single poll. On a
slow or loaded host it does not, and the attribution flips.
handoffGraceMs replaces the hardcoded one-poll window as a validated,
deployment-owned config field defaulting to 500ms, rejected at load when
it cannot contain one readiness poll. Real-shell tests that interrupt a
send now assert the session is usable again rather than which readiness
tier observed the handoff, because no fixed grace removes the race.
Stack the parallel-in_progress change on the web todo display (#497): the
GUI is now the surface where several active items are visible, so the two
land as a chain rather than colliding on tool-todo at merge time.
Conflicts combined rather than resolved to one side: tool-todo keeps this
branch's parallel-allowing validation AND web2-todo's additionalProperties
unknown-key rejection, in src/index.ts and both README sides; the spec
drops web2-todo's 'two in_progress' rejection case and keeps its unknown-key
case; the two headless advanced-toolchain session fixtures keep this
branch's parallel transcripts.
Review fix (ds-review-bot on #623): the ACP scenario runs at deployment
strength (the automation protocol has no session-scoped switch), so the
assembled-app path could not detect the delegation bypass itself. The new
keyless subagent-inheritance headless scenario closes that on the
semantic-checkpoint precedent: a seeded parent log carrying a real
sandbox/mode: read-only switch under a workspace-write deployment default
is resumed through the Loader-booted cli-demo app via a resume fixture
plugin and delegates through the real subagent tool; the child's real
write is denied by the real dsh-fs-sandbox fence (physical ENOENT
assertion), its persisted header carries the inherited baseline, and both
logs pin as expected outputs. Verified red: disabling the driver's capture
makes the scenario fail on the disk assertion (the child writes under the
deployment default).
Two review findings on the turndown swap, both verified empirically:
- Unclosed-tag nesting makes the synchronous turndown/domino walk
superlinear (measured: depth 512 ~0.15s, 2k ~2s, 20k ~5s), during
which the cooperative fetchTimeoutMs timer cannot fire. renderBody
now preflights nesting depth with a linear tag scan and passes
bodies past 512 levels through raw; the try/catch stays for markup
the scan cannot see (comment-hidden tags), simulated in tests via a
converter throw.
- Markdown escaping can expand converted HTML ~2x (100k underscores
render as 200k chars), so provider body caps no longer bounded the
model-visible result. formatFetchOutput now caps the complete output
(header + body + footer) under new fetchMaxOutputChars config
(default 200000 = 2x the local provider's default body cap), reusing
the truncation notice.
README EN+ZH, config catalog, Agent Note EN+ZH updated; the new
web-fetch fixture is migrated to the packed layout master now
requires; tool-web coverage stays 100% per-file.
The merge of origin/master at 9f218ce9d took master's re-recorded parent
session.jsonl wholesale, which reverted this branch's todo_write
description in that one file while the two child logs kept the new
parallel-in_progress text. The headless snapshot scrubs request headers
before comparison, so the three logs disagreed on the model-visible
tool contract without any test failing.
Re-record the scenario with test:snapshot:refresh, which replays the
committed scripts and rewrites all three persisted-log fixtures from the
live run. The parent regains the parallel-in_progress todo_write
description; both children pick up master's current run_code description
and its required `description` parameter, which they were stale on.
Fixture content only; no source or contract change, so the owning Agent
Note stands as written.
The simplified request path no longer anchors an unchanged request/header on
resume (agent.ts logs a header only when it differs from the folded baseline),
so resume-turn logs drop that event and later seqs shift down. Refresh the
keyless session-log fixtures to match; no model scripts changed.