Replace the full-innerHTML render loop with a static shell plus per-region updates, so composer drafts, fold state, focus, and scroll survive streaming turns. Fold session events into one trace graph consumed by Chat, Trajectory, Waterfall, and the shared inspector drawer, with live ACP updates patched into a keyed live-turn region. Align the visual system with a tokenized design spec: a 4px spacing base with fixed control/row height steps, foreground-derived text tiers and borders (color-mix), neutral interaction overlays, tiered motion durations with a reduced-motion collapse, hover-revealed scrollbars, and drawer-aware layout elasticity. Localize trajectory role chips and row previews. The renderer entry (app.ts) joins the coverage exclude list as a self-executing DOM bootstrap: jsdom lifecycle specs exercise its behavior, and extractable logic lives in covered modules (trace-graph.ts, renderer-content.ts).
@deepseek-ai/dsh-desktop
Desktop workbench for developing and studying DeepSeek Harness agents. The app is a first-party Electron client bound to this repository, not a generic workspace picker and not a continuation of the localhost prototype under /Users/tn.shen/Documents/原型.
The product loop is: run a task in chat, inspect what happened, modify Harness or a local plugin, restart only the runtime, replay a previous task, compare the new run against the baseline, and repeat.
Run it
From the repository root:
pnpm --dir packages/ui/desktop run dev
This starts a Vite renderer on http://127.0.0.1:5174, opens Electron, and starts the real Harness ACP runtime with:
node --import tsx packages/examples/acp-demo/src/bin.ts --config examples/acp-agent/cordis.yml
The renderer talks only to the Electron preload API. The main process owns the ACP subprocess, session JSONL reads, feedback writes, runtime restart, and diagnostics.
Useful checks:
node_modules/.bin/vitest run packages/ui/desktop/tests/index.spec.ts packages/ui/desktop/tests/acp-subprocess.spec.ts
node_modules/.bin/tsc --noEmit --ignoreConfig --module NodeNext --moduleResolution NodeNext --target ES2022 --lib ES2022,DOM --strict --allowImportingTsExtensions packages/ui/desktop/src/index.ts packages/ui/desktop/src/app.ts packages/ui/desktop/src/global.d.ts packages/ui/desktop/src/css.d.ts
User journey
The desktop app is for a Harness developer or researcher who wants to understand and improve one repo-bound agent runtime.
- Install and open the app. The app binds to this repository; there is no workspace picker.
- Start a session in
Chatand run one or two tasks to prove the runtime works. - Switch the same run to
Trajectory,Waterfall, orContextto understand what happened behind one message. - Click any interesting message, step, tool call, request, timing bar, or context section to open the
Inspectorwith complete facts. - Use
Devto ask the agent to modify Harness, a plugin, a prompt, a tool, or config. The first version seeds chat with repo paths and context instead of exposing a graphical code editor. - Restart only the managed Harness runtime subprocess when code/config changes require it.
- Replay a previous prompt or turn against the changed runtime.
- Compare the baseline run and candidate run. Keep the new behavior or ask the agent to revert/change it, then repeat.
Product shape
The Electron shell owns the desktop lifecycle while the Harness runtime runs as a managed subprocess. The renderer never talks directly to the filesystem or the model. It calls typed preload APIs; the main process starts the runtime, talks to the Agent Client Protocol (ACP) server over JSON-RPC stdio, and reads session logs through trusted query adapters.
The default runtime channel is ACP because dsh-acp already owns session/new, session/load, session/prompt, session/cancel, streaming session/update, permission prompts, and user elicitations. The desktop app may add a query side channel for persisted JSONL or ctx.sessionQuery, but live chat should not be implemented by polling JSONL.
The app is repository-bound. On startup it treats the package root as the Harness repository, launches the runtime from that directory, reads .sessions from that runtime's persistence config, and marks each run with repository state such as commit, dirty status, runtime config hash, and parent replay metadata.
Main surfaces
Sessions is the left navigation. It lists and searches live or persisted sessions, groups child sessions under their parent, opens or resumes a session, and can reveal the selected JSONL in Finder. Baseline pinning and replay remain part of the product shape below rather than shipped controls.
Chat is the reading and driving surface. One user turn has one assistant response shell containing ordered, independently selectable Thinking, paired Tool, and visible text blocks. Clicking a block opens its inspector target; the block's trailing arrow only expands or collapses its inline preview. The composer is always available for the selected live session.
Trajectory is the structural surface. A structure tree (turn -> step with durations, tool summaries, and error marks) navigates a step-grouped logical-object table; lifecycle events become tree nodes and sticky group headers instead of rows. A model response owns its effective request input and assistant output. A Tool row pairs tool/call with tool/result, so one row and inspector target own both Input and Output. Clicking the row opens the inspector; the trailing arrow only expands the inline payload.
Waterfall is the time surface. It answers where latency went across turns, steps, model calls, tool calls, and failures: a summary strip (total / LLM time / tool time / errors / slowest step / tokens) over an aligned time track. Clicking any bar, label, or stat keeps Waterfall active and opens the shared inspector; the inspector's explicit → Trajectory action performs cross-view navigation.
Context is the request-anatomy surface. It answers what the model saw at a selected request boundary: config, system prompt, session prefix, derived conversation surface, injected context/message entries, compaction summaries, visible tools, and request-header deltas. It may show section summaries and bounded previews, but the full raw data still belongs in the inspector.
Inspector is a right-side drawer, not a permanent column. It appears only after the user selects a message, trajectory node, waterfall bar, or context section. Its tabs are Input, Output, Metadata, and Feedback. It is the only place for complete JSON, JSONL, full schemas, complete system prompts, raw event windows, copy actions, and node-targeted feedback.
Surface behavior contract
| Surface | Primary job | Inline content | Inspector trigger |
|---|---|---|---|
Chat |
Drive and read the conversation | User messages, assistant messages, collapsed thinking rows, collapsed tool-use rows | Message or activity click |
Trajectory |
Navigate run structure | Turn/step tree plus User / Thinking / Model / paired Tool / Context logical records, statuses, durations, short previews | Logical row click |
Waterfall |
Diagnose latency | Spans, critical path, status, start/duration | Bar click |
Context |
Explain request anatomy | Config summary, system summary, prefix summary, derived-history summary, context-message summary, tool-schema list summary, deltas | Section click |
Compare |
Compare two run artifacts | Baseline/candidate diffs for output, context, tools, events, usage, duration, errors | Diff hunk click |
Dev |
Support repo-bound modification loop | Runtime state, dirty state, watched paths, suggested agent prompts, restart-needed flag | Config/plugin/path click |
The session shell keeps one composer mounted while Chat, Trajectory, and Waterfall switch in the middle pane, so a developer can continue the selected live session without losing a draft or leaving an inspection view. Non-session modules do not show it.
Information ownership
Middle surfaces locate and explain. The inspector preserves the complete facts.
Trajectory and Context should not duplicate inspector responsibilities. Trajectory shows where the user is in the run. Context shows which context sources contributed to a request. Inspector shows the selected object's complete input, output, metadata, and feedback.
The same selected object can be entered from multiple surfaces. A paired Tool selected from Trajectory, a Tool row selected from Chat, or a tool segment selected from Waterfall resolves to the same tool:<callId> target. Thinking uses reasoning:<turn>:<step> and model output uses assistant:<seq>. This keeps Input, Output, feedback, selection, and copying attached to the logical object, not to the view that opened it.
Why Trajectory still needs Inspector
Trajectory answers orientation questions:
- Which turn and step produced this behavior?
- Which request triggered which assistant output and tool calls?
- Did the run fail, cancel, or continue?
- What is the rough shape of this step before I inspect raw data?
Inline expansion answers those questions in place, but every row resolves to the same logical target used by Chat and Waterfall. Input/Output/Metadata format switching and Feedback history stay attached to that target, and the drawer stays open while switching between Chat, Trajectory, and Waterfall.
Why Context still needs Inspector
Context answers request-anatomy questions:
- What did this request include outside normal chat history?
- Did the system prompt, tools, call config, or session prefix change?
- Which
context/messageorsteering/messageentries were model-visible? - Was history compacted or replaced before this request?
- Which sections are large enough to matter for token cost?
Context should show section summaries and bounded previews. The complete system prompt, complete tool schemas, full derived message list, raw request/header, and raw request/header-delta belong in Inspector. This makes Context useful as a map while preserving the user's requirement that every fact remains reachable.
Plugin and development panel
The Dev panel is for changing Harness itself. It is not a full IDE and should not start as a graphical Cordis editor.
The first version lists Cordis config entries, local plugins/* packages when present, runtime status, repository dirty state, and whether a runtime restart is needed. It provides actions to ask the agent to add, modify, or disable a plugin by seeding the current chat with the relevant paths and config context.
Runtime changes are applied by restarting the managed ACP subprocess, not by restarting Electron. File watching should mark Restart needed for changes under packages/**, examples/**, plugins/**, cordis.yml, and package manifests. HMR can be explored later, but the reliable path is subprocess restart because ACP stdout is the JSON-RPC protocol channel.
Replay and compare
Trace is looking at the past. Replay is running a past task against the current runtime. Compare is inspecting two run artifacts.
Replay creates a new run with parentRunId or replayOf metadata. If the app can only replay the user prompt and cwd rather than reconstructing an exact intermediate state, the UI must label it as prompt replay. It must not imply bit-for-bit session replay unless the runtime can prove it restored the same context boundary.
Compare is not a normal tab inside one session. It is a mode over two runs: baseline and candidate. The first version compares final output, request header, context sections, tool sequence, tool inputs and outputs, event counts, step counts, usage, duration, and errors.
Backend integration plan
The first implementation uses two backend paths:
- ACP subprocess for live chat and runtime lifecycle. It maps to
session/new,session/load,session/prompt,session/cancel, and streamedsession/update. This is the authoritative path for driving the agent. - Session query/read adapters for exact inspection. They read live-preferred persisted logs through
listSessions(),listEvents(sessionId), and bounded raw-event windows for Inspector.
The renderer receives normalized view models, not raw filesystem paths. The main process owns:
- Runtime process start, stop, restart, status, stderr diagnostics, and JSON-RPC stdout framing.
- Session/run discovery and run metadata sidecars.
- Trace normalization into chat messages, trajectory nodes, waterfall spans, context sections, compare diffs, and inspector payload refs.
- Feedback writes attached to stable inspector target ids.
Trace and context data sources
The current session log can provide all core facts needed by the UI:
turn/startandturn/enddefine durable user-visible turn boundaries.step/startandstep/enddefine one model request plus its tool work.request/headerandrequest/header-deltareconstruct model config, system prompt, tools, andmessagePrefix.user/message,assistant/message,tool/result,context/message, andsteering/messagedefine the model-visible surface throughderiveMessages().assistant/chunkpreserves token-level streaming fidelity but should usually be folded in Chat and opened in Inspector only when needed.tool/callandtool/resultdefine tool input/output pairs and errors.sourceEventSeqsandsurfaceOpexplain provenance and compaction/replacement.
Minimum implementation plan
Start with a desktop package that defines shared UI contracts, then build the Electron shell around them.
- Add an Electron main process that launches the ACP runtime subprocess from the repository root and owns lifecycle commands: start, stop, restart, status.
- Add a preload API for sessions, prompts, trace reads, replay, compare, dev status, and feedback. Renderer code must not use Node globals.
- Build the renderer around the five middle surfaces and the inspector drawer. The renderer should treat all trace/context/compare objects as view models supplied by the main process.
- Implement session ingestion from ACP live updates and persisted session logs. Normalize selected objects to stable inspector targets.
- Implement Dev panel actions as agent-seeded chat prompts first. Direct file mutations and graphical Cordis editing are later features.
- Implement replay by creating a new run from a historical prompt or turn. Attach run metadata and open Compare after the candidate completes.
Phase-one acceptance criteria
- A developer can open the desktop app from this repository without selecting a workspace.
- A developer can create/load sessions and send prompts through the real ACP runtime.
- The same run can switch between
Chat,Trajectory,Waterfall, andContext. - Clicking any middle-surface object opens the Inspector drawer with
Input,Output,Metadata, andFeedback, with Feedback last. - Complete system prompts, tool schemas, raw event windows, and JSON/JSONL are reachable from Inspector.
- Feedback defaults the author to
shentuniand persists against the selected target. - Runtime restart does not close the Electron shell.
- Replay creates a separate candidate run with lineage metadata.
- Compare operates over two runs, not "inside" one session.
Model Experience
Composer prompt
What the model sees: The desktop composer sends the user's text to the managed ACP runtime as the session/prompt content. This package adds no extra system prompt, steering prose, or tool definition of its own.
Token effect: User-message tokens are data-dependent and then follow the active runtime's normal session-retention and compaction behavior. Electron shell chrome, inspector state, and the language toggle add zero model-context tokens.
Runtime context evidence
What the model sees: The visible system prompt, steering messages, tool set, message prefix, and config come from the active cordis.yml composition and its plugins. The desktop app only reads those emitted facts from ACP updates and persisted JSONL for inspection.
Token effect: No additional tokens are introduced by viewing Chat, Trajectory, Waterfall, or Develop. The displayed request/header, context/message, and steering/message data reflect tokens the runtime already assembled for the model.
Known Limitations and Deferred Work
- Development build only — this package ships a usable Electron/Vite app and a real ACP subprocess bridge, but it is not yet packaged as a signed distributable.
- ACP is the first runtime channel — direct in-process embedding could make context queries and restarts richer, but would make isolation, teardown, and hot reload harder.
- Develop is read-first — it exposes prompts, tools, plugins, config, runtime state, and the change loop as a source browser; direct graphical plugin/config editing is deferred.
- Trace refresh is mixed live/persisted — chat streams from ACP live updates (rendered incrementally, so composer input, fold state, and scroll survive streaming), while Trajectory and Waterfall read persisted JSONL after turns complete.
- Context and Compare surfaces are unimplemented — the session view ships
Chat,Trajectory, andWaterfall; theContext/Comparecontracts above and replay remain documented product shape for later work.