Files
deepseek-harness/packages/ui/desktop
NI0317 c4584b5c19 feat(ui): rebuild desktop renderer on a shared trace graph and design tokens
Replace the full-innerHTML render loop with a static shell plus
per-region updates, so composer drafts, fold state, focus, and scroll
survive streaming turns. Fold session events into one trace graph
consumed by Chat, Trajectory, Waterfall, and the shared inspector
drawer, with live ACP updates patched into a keyed live-turn region.

Align the visual system with a tokenized design spec: a 4px spacing
base with fixed control/row height steps, foreground-derived text
tiers and borders (color-mix), neutral interaction overlays, tiered
motion durations with a reduced-motion collapse, hover-revealed
scrollbars, and drawer-aware layout elasticity. Localize trajectory
role chips and row previews.

The renderer entry (app.ts) joins the coverage exclude list as a
self-executing DOM bootstrap: jsdom lifecycle specs exercise its
behavior, and extractable logic lives in covered modules
(trace-graph.ts, renderer-content.ts).
2026-07-19 01:37:04 +08:00
..

@deepseek-ai/dsh-desktop

Desktop workbench for developing and studying DeepSeek Harness agents. The app is a first-party Electron client bound to this repository, not a generic workspace picker and not a continuation of the localhost prototype under /Users/tn.shen/Documents/原型.

The product loop is: run a task in chat, inspect what happened, modify Harness or a local plugin, restart only the runtime, replay a previous task, compare the new run against the baseline, and repeat.

Run it

From the repository root:

pnpm --dir packages/ui/desktop run dev

This starts a Vite renderer on http://127.0.0.1:5174, opens Electron, and starts the real Harness ACP runtime with:

node --import tsx packages/examples/acp-demo/src/bin.ts --config examples/acp-agent/cordis.yml

The renderer talks only to the Electron preload API. The main process owns the ACP subprocess, session JSONL reads, feedback writes, runtime restart, and diagnostics.

Useful checks:

node_modules/.bin/vitest run packages/ui/desktop/tests/index.spec.ts packages/ui/desktop/tests/acp-subprocess.spec.ts
node_modules/.bin/tsc --noEmit --ignoreConfig --module NodeNext --moduleResolution NodeNext --target ES2022 --lib ES2022,DOM --strict --allowImportingTsExtensions packages/ui/desktop/src/index.ts packages/ui/desktop/src/app.ts packages/ui/desktop/src/global.d.ts packages/ui/desktop/src/css.d.ts

User journey

The desktop app is for a Harness developer or researcher who wants to understand and improve one repo-bound agent runtime.

  1. Install and open the app. The app binds to this repository; there is no workspace picker.
  2. Start a session in Chat and run one or two tasks to prove the runtime works.
  3. Switch the same run to Trajectory, Waterfall, or Context to understand what happened behind one message.
  4. Click any interesting message, step, tool call, request, timing bar, or context section to open the Inspector with complete facts.
  5. Use Dev to ask the agent to modify Harness, a plugin, a prompt, a tool, or config. The first version seeds chat with repo paths and context instead of exposing a graphical code editor.
  6. Restart only the managed Harness runtime subprocess when code/config changes require it.
  7. Replay a previous prompt or turn against the changed runtime.
  8. Compare the baseline run and candidate run. Keep the new behavior or ask the agent to revert/change it, then repeat.

Product shape

The Electron shell owns the desktop lifecycle while the Harness runtime runs as a managed subprocess. The renderer never talks directly to the filesystem or the model. It calls typed preload APIs; the main process starts the runtime, talks to the Agent Client Protocol (ACP) server over JSON-RPC stdio, and reads session logs through trusted query adapters.

The default runtime channel is ACP because dsh-acp already owns session/new, session/load, session/prompt, session/cancel, streaming session/update, permission prompts, and user elicitations. The desktop app may add a query side channel for persisted JSONL or ctx.sessionQuery, but live chat should not be implemented by polling JSONL.

The app is repository-bound. On startup it treats the package root as the Harness repository, launches the runtime from that directory, reads .sessions from that runtime's persistence config, and marks each run with repository state such as commit, dirty status, runtime config hash, and parent replay metadata.

Main surfaces

Sessions is the left navigation. It lists and searches live or persisted sessions, groups child sessions under their parent, opens or resumes a session, and can reveal the selected JSONL in Finder. Baseline pinning and replay remain part of the product shape below rather than shipped controls.

Chat is the reading and driving surface. One user turn has one assistant response shell containing ordered, independently selectable Thinking, paired Tool, and visible text blocks. Clicking a block opens its inspector target; the block's trailing arrow only expands or collapses its inline preview. The composer is always available for the selected live session.

Trajectory is the structural surface. A structure tree (turn -> step with durations, tool summaries, and error marks) navigates a step-grouped logical-object table; lifecycle events become tree nodes and sticky group headers instead of rows. A model response owns its effective request input and assistant output. A Tool row pairs tool/call with tool/result, so one row and inspector target own both Input and Output. Clicking the row opens the inspector; the trailing arrow only expands the inline payload.

Waterfall is the time surface. It answers where latency went across turns, steps, model calls, tool calls, and failures: a summary strip (total / LLM time / tool time / errors / slowest step / tokens) over an aligned time track. Clicking any bar, label, or stat keeps Waterfall active and opens the shared inspector; the inspector's explicit → Trajectory action performs cross-view navigation.

Context is the request-anatomy surface. It answers what the model saw at a selected request boundary: config, system prompt, session prefix, derived conversation surface, injected context/message entries, compaction summaries, visible tools, and request-header deltas. It may show section summaries and bounded previews, but the full raw data still belongs in the inspector.

Inspector is a right-side drawer, not a permanent column. It appears only after the user selects a message, trajectory node, waterfall bar, or context section. Its tabs are Input, Output, Metadata, and Feedback. It is the only place for complete JSON, JSONL, full schemas, complete system prompts, raw event windows, copy actions, and node-targeted feedback.

Surface behavior contract

Surface Primary job Inline content Inspector trigger
Chat Drive and read the conversation User messages, assistant messages, collapsed thinking rows, collapsed tool-use rows Message or activity click
Trajectory Navigate run structure Turn/step tree plus User / Thinking / Model / paired Tool / Context logical records, statuses, durations, short previews Logical row click
Waterfall Diagnose latency Spans, critical path, status, start/duration Bar click
Context Explain request anatomy Config summary, system summary, prefix summary, derived-history summary, context-message summary, tool-schema list summary, deltas Section click
Compare Compare two run artifacts Baseline/candidate diffs for output, context, tools, events, usage, duration, errors Diff hunk click
Dev Support repo-bound modification loop Runtime state, dirty state, watched paths, suggested agent prompts, restart-needed flag Config/plugin/path click

The session shell keeps one composer mounted while Chat, Trajectory, and Waterfall switch in the middle pane, so a developer can continue the selected live session without losing a draft or leaving an inspection view. Non-session modules do not show it.

Information ownership

Middle surfaces locate and explain. The inspector preserves the complete facts.

Trajectory and Context should not duplicate inspector responsibilities. Trajectory shows where the user is in the run. Context shows which context sources contributed to a request. Inspector shows the selected object's complete input, output, metadata, and feedback.

The same selected object can be entered from multiple surfaces. A paired Tool selected from Trajectory, a Tool row selected from Chat, or a tool segment selected from Waterfall resolves to the same tool:<callId> target. Thinking uses reasoning:<turn>:<step> and model output uses assistant:<seq>. This keeps Input, Output, feedback, selection, and copying attached to the logical object, not to the view that opened it.

Why Trajectory still needs Inspector

Trajectory answers orientation questions:

  • Which turn and step produced this behavior?
  • Which request triggered which assistant output and tool calls?
  • Did the run fail, cancel, or continue?
  • What is the rough shape of this step before I inspect raw data?

Inline expansion answers those questions in place, but every row resolves to the same logical target used by Chat and Waterfall. Input/Output/Metadata format switching and Feedback history stay attached to that target, and the drawer stays open while switching between Chat, Trajectory, and Waterfall.

Why Context still needs Inspector

Context answers request-anatomy questions:

  • What did this request include outside normal chat history?
  • Did the system prompt, tools, call config, or session prefix change?
  • Which context/message or steering/message entries were model-visible?
  • Was history compacted or replaced before this request?
  • Which sections are large enough to matter for token cost?

Context should show section summaries and bounded previews. The complete system prompt, complete tool schemas, full derived message list, raw request/header, and raw request/header-delta belong in Inspector. This makes Context useful as a map while preserving the user's requirement that every fact remains reachable.

Plugin and development panel

The Dev panel is for changing Harness itself. It is not a full IDE and should not start as a graphical Cordis editor.

The first version lists Cordis config entries, local plugins/* packages when present, runtime status, repository dirty state, and whether a runtime restart is needed. It provides actions to ask the agent to add, modify, or disable a plugin by seeding the current chat with the relevant paths and config context.

Runtime changes are applied by restarting the managed ACP subprocess, not by restarting Electron. File watching should mark Restart needed for changes under packages/**, examples/**, plugins/**, cordis.yml, and package manifests. HMR can be explored later, but the reliable path is subprocess restart because ACP stdout is the JSON-RPC protocol channel.

Replay and compare

Trace is looking at the past. Replay is running a past task against the current runtime. Compare is inspecting two run artifacts.

Replay creates a new run with parentRunId or replayOf metadata. If the app can only replay the user prompt and cwd rather than reconstructing an exact intermediate state, the UI must label it as prompt replay. It must not imply bit-for-bit session replay unless the runtime can prove it restored the same context boundary.

Compare is not a normal tab inside one session. It is a mode over two runs: baseline and candidate. The first version compares final output, request header, context sections, tool sequence, tool inputs and outputs, event counts, step counts, usage, duration, and errors.

Backend integration plan

The first implementation uses two backend paths:

  • ACP subprocess for live chat and runtime lifecycle. It maps to session/new, session/load, session/prompt, session/cancel, and streamed session/update. This is the authoritative path for driving the agent.
  • Session query/read adapters for exact inspection. They read live-preferred persisted logs through listSessions(), listEvents(sessionId), and bounded raw-event windows for Inspector.

The renderer receives normalized view models, not raw filesystem paths. The main process owns:

  • Runtime process start, stop, restart, status, stderr diagnostics, and JSON-RPC stdout framing.
  • Session/run discovery and run metadata sidecars.
  • Trace normalization into chat messages, trajectory nodes, waterfall spans, context sections, compare diffs, and inspector payload refs.
  • Feedback writes attached to stable inspector target ids.

Trace and context data sources

The current session log can provide all core facts needed by the UI:

  • turn/start and turn/end define durable user-visible turn boundaries.
  • step/start and step/end define one model request plus its tool work.
  • request/header and request/header-delta reconstruct model config, system prompt, tools, and messagePrefix.
  • user/message, assistant/message, tool/result, context/message, and steering/message define the model-visible surface through deriveMessages().
  • assistant/chunk preserves token-level streaming fidelity but should usually be folded in Chat and opened in Inspector only when needed.
  • tool/call and tool/result define tool input/output pairs and errors.
  • sourceEventSeqs and surfaceOp explain provenance and compaction/replacement.

Minimum implementation plan

Start with a desktop package that defines shared UI contracts, then build the Electron shell around them.

  1. Add an Electron main process that launches the ACP runtime subprocess from the repository root and owns lifecycle commands: start, stop, restart, status.
  2. Add a preload API for sessions, prompts, trace reads, replay, compare, dev status, and feedback. Renderer code must not use Node globals.
  3. Build the renderer around the five middle surfaces and the inspector drawer. The renderer should treat all trace/context/compare objects as view models supplied by the main process.
  4. Implement session ingestion from ACP live updates and persisted session logs. Normalize selected objects to stable inspector targets.
  5. Implement Dev panel actions as agent-seeded chat prompts first. Direct file mutations and graphical Cordis editing are later features.
  6. Implement replay by creating a new run from a historical prompt or turn. Attach run metadata and open Compare after the candidate completes.

Phase-one acceptance criteria

  • A developer can open the desktop app from this repository without selecting a workspace.
  • A developer can create/load sessions and send prompts through the real ACP runtime.
  • The same run can switch between Chat, Trajectory, Waterfall, and Context.
  • Clicking any middle-surface object opens the Inspector drawer with Input, Output, Metadata, and Feedback, with Feedback last.
  • Complete system prompts, tool schemas, raw event windows, and JSON/JSONL are reachable from Inspector.
  • Feedback defaults the author to shentuni and persists against the selected target.
  • Runtime restart does not close the Electron shell.
  • Replay creates a separate candidate run with lineage metadata.
  • Compare operates over two runs, not "inside" one session.

Model Experience

Composer prompt

What the model sees: The desktop composer sends the user's text to the managed ACP runtime as the session/prompt content. This package adds no extra system prompt, steering prose, or tool definition of its own.

Token effect: User-message tokens are data-dependent and then follow the active runtime's normal session-retention and compaction behavior. Electron shell chrome, inspector state, and the language toggle add zero model-context tokens.

Runtime context evidence

What the model sees: The visible system prompt, steering messages, tool set, message prefix, and config come from the active cordis.yml composition and its plugins. The desktop app only reads those emitted facts from ACP updates and persisted JSONL for inspection.

Token effect: No additional tokens are introduced by viewing Chat, Trajectory, Waterfall, or Develop. The displayed request/header, context/message, and steering/message data reflect tokens the runtime already assembled for the model.

Known Limitations and Deferred Work

  • Development build only — this package ships a usable Electron/Vite app and a real ACP subprocess bridge, but it is not yet packaged as a signed distributable.
  • ACP is the first runtime channel — direct in-process embedding could make context queries and restarts richer, but would make isolation, teardown, and hot reload harder.
  • Develop is read-first — it exposes prompts, tools, plugins, config, runtime state, and the change loop as a source browser; direct graphical plugin/config editing is deferred.
  • Trace refresh is mixed live/persisted — chat streams from ACP live updates (rendered incrementally, so composer input, fold state, and scroll survive streaming), while Trajectory and Waterfall read persisted JSONL after turns complete.
  • Context and Compare surfaces are unimplemented — the session view ships Chat, Trajectory, and Waterfall; the Context/Compare contracts above and replay remain documented product shape for later work.