feat(web): add durable session metrics (round 1)

This commit is contained in:
Hypatia May
2026-07-28 12:10:48 +08:00
parent 2a46685414
commit 9ca0241d5d
43 changed files with 1139 additions and 97 deletions
@@ -0,0 +1,6 @@
# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
# side as of the last confirmed-consistent state. Both languages carry equal authority;
# after editing either side, bring the other along and re-record with:
# pnpm run verify-translation-pairing --write .agents/notes/implemented/architecture/2026-07-28-host-owned-web-session-metrics.md
2026-07-28-host-owned-web-session-metrics.md: 04381c7443491fd9de101a87713fa5c183800d0e
2026-07-28-host-owned-web-session-metrics.zh.md: 6ad06ea61508d3f0703c19bc7c9f14969f119334
@@ -0,0 +1,35 @@
# Agent Note: Host-owned Web session metrics
Status: implemented
English | [中文](2026-07-28-host-owned-web-session-metrics.zh.md)
## Problem
A Web stats line derived from the currently loaded conversation nodes is window-dependent under pagination. Compaction can replace visible content without preserving historical usage, and route changes leave the browser without an authoritative context capacity. Cache-write tokens also risk being folded into a cache-hit formula whose denominator has different semantics.
## Decision
The Host owns one session-level metrics projection. It incrementally folds the complete durable event log, keys settled usage by `(turn, step)`, and replaces an earlier usage record for the same key instead of double-counting chunk and message forms. Uncached input, output, cache reads, and cache writes remain four disjoint cumulative buckets. Compaction can change the current prompt surface without erasing historical usage.
Current context pressure is a separate point-in-time value from `tokenMeter.measure(session).totalTokens`. Capacity comes only from `llm.resolveModelInfo(provider, model).context.contextWindow` for the agent's selected route. A route change immediately publishes metrics with capacity absent, then publishes the resolved capacity behind a route generation fence; stale metadata cannot label the new route.
The tail `session.history` response carries the projection, while older pages omit it. Live changes use `session/metrics` mux frames. Both forms carry a durable-log revision and a projection revision; the client accepts only nondecreasing revisions, preserves metrics across older-page prepend, and clears them at a new subscription baseline. Missing measurement or metadata stays absent.
The Web stats line treats the projection as its sole token source. It renders uncached input, output, and cache reads separately, computes cache hit as `cacheRead / (uncachedInput + cacheRead)`, and shows current context as a percentage of the exact route capacity. Cache writes never enter that percentage. Visible nodes continue to supply only turn and step counts.
## Alternatives considered
**Fold the loaded node window in React.** This cannot survive pagination or compaction and duplicates durable-log semantics in a presentation package.
**Send usage only with raw assistant events.** Reconnect and older-page stitching would still need the client to reconstruct a full-log aggregate, and duplicate usage forms would need protocol-specific repair there.
**Reuse one total-token field for cache hit.** Cache reads, cache writes, and uncached input represent distinct provider accounting buckets; combining them would make the displayed rate misleading.
**Keep the previous capacity until the new route resolves.** The old number would temporarily claim the wrong selected model. An explicit unknown state is honest and generation-safe.
## Consequences
Token totals remain stable across pagination, replay, compaction, and browser reconnect. The client stores a small detached projection instead of scanning the conversation window, and the status row remains readable for large histories through compact number formatting.
The Host performs one incremental log fold per session and schedules live projection updates only for usage, request-header, or surface-changing events; text and reasoning deltas do not publish metrics. Exact capacity resolution is asynchronous and may briefly render as unknown. Deployments without a token meter or model context metadata retain the row and label the unavailable value instead of fabricating one.
@@ -0,0 +1,35 @@
# Agent Note: Host 拥有的 Web 会话指标
Status: implemented
[English](2026-07-28-host-owned-web-session-metrics.md) | 中文
## 问题
Web 统计行若根据当前加载的会话节点推导指标,其结果会随分页窗口变化。压缩(compaction)可以替换可见内容,却无法保留历史用量;路由变更会让浏览器缺少权威的上下文容量。缓存写入 token 还可能被计入缓存命中率公式,而该公式的分母具有不同语义。
## 决策
Host 拥有一项会话级指标投影。它以增量方式归并完整的持久事件日志,按 `(turn, step)` 标识已结算用量;同一标识再次出现时,会替换较早的用量记录,而不会重复统计分片和消息两种形态。未缓存输入、输出、缓存读取与缓存写入保持为四个彼此独立的累计计数项。压缩可以改变当前提示词表层,但不会抹除历史用量。
当前上下文压力是一个独立的即时值,取自 `tokenMeter.measure(session).totalTokens`。容量仅来自 `llm.resolveModelInfo(provider, model).context.contextWindow`,并对应 agent(智能体)所选的路由。路由变更时,Host 会立即发布不带容量的指标,再通过路由代际围栏发布解析出的容量;陈旧元数据无法标记新的路由。
`session.history` 尾页响应携带该投影,较早页面则省略它。实时变更使用 `session/metrics` mux 帧。两种形式都携带持久日志修订号和投影修订号;客户端只接受不减小的修订号,在向前加载较早页面时保留指标,并在建立新的订阅基线时将其清除。测量值或元数据缺失时,对应字段保持缺失。
Web 统计行把该投影视为唯一的 token 数据来源。它分别呈现未缓存输入、输出与缓存读取,通过 `cacheRead / (uncachedInput + cacheRead)` 计算缓存命中率,并把当前上下文显示为精确路由容量的百分比。缓存写入绝不计入缓存命中率。可见节点仍然只提供轮次和步骤计数。
## 备选方案
**在 React 中归并已加载的节点窗口。** 此方案无法跨越分页或压缩保留数据,还会在展示包中重复实现持久日志语义。
**只随原始 assistant 事件发送用量。** 重连和较早页面拼接仍会要求客户端重建完整日志聚合,而且重复的用量形态需要在客户端按协议专门修复。
**为缓存命中率复用单一的 token 总数字段。** 缓存读取、缓存写入与未缓存输入是提供方记账中的不同计数项;将它们合并会使显示的比率产生误导。
**在新路由解析完成前保留旧容量。** 旧数值会在短时间内错误标示所选模型。显式的「未知」状态能如实反映情况,并避免跨代串扰。
## 后果
token 总量在分页、回放、压缩和浏览器重连期间保持稳定。客户端存储一项小型脱耦投影,无需扫描会话窗口;状态行采用紧凑数字格式,因此在较长的历史记录中仍然清晰易读。
Host 为每个会话执行一次增量日志归并,仅为用量事件、请求头事件或表层变更事件调度实时投影更新;文本与推理(reasoning)增量不会发布指标。精确容量解析为异步操作,因此可能短暂显示「未知」。未部署 token 计量器或缺少模型上下文元数据时,系统仍保留该行,并标示不可用的值,而不会虚构数据。
@@ -7,22 +7,30 @@
- tab "Trajectory"
- tab "Waterfall"
- text: "Use the bash tool to run exactly: echo WEB_E2E_OK. Then reply with the single word DONE and stop."
- button "复制":
- img
- button "在新对话中分支":
- img
- button "编辑":
- img
- button "▸ 上下文注入"
- button "Think The user wants me to run a simple bash command and reply with \"DONE\".":
- img
- text: Think The user wants me to run a simple bash command and reply with "DONE".
- text: Echo the test string
- img
- text: Bash Echo the test string
- button "Think The command executed successfully and output \"WEB_E2E_OK\". I just need to reply with \"DONE\".":
- img
- text: Think The command executed successfully and output "WEB_E2E_OK". I just need to reply with "DONE".
- paragraph: DONE
- text: cache hit 99% · 15,818 tokens · 1 turns · 2 steps
- text: 219 uncached input · 111 output · 15.5k cache read · cache hit 99% · context 6% of 128k · 1 turns · 2 steps
- textbox "Message the agent"
- button "Add attachment":
- img
- combobox "Access mode":
- option "Read-only" [selected]
- option "Read-write"
- button "选择模型,当前 deepseek-v4-flash":
- text: deepseek-v4-flash
- button "选择模型,当前 DeepSeek-V4-Flash":
- text: DeepSeek-V4-Flash
- img
- button "Send message" [disabled]
+1 -1
View File
@@ -9,7 +9,7 @@ This matrix shows which packages dispatch each harness-owned event and which pac
| --- | --- | --- | --- | --- |
| `agent-loop/config-start-failed` | `emit` | [`packages/core/agent-loop/src/index.ts:140`](../packages/core/agent-loop/src/index.ts) | [`agent-loop`](../packages/core/agent-loop) (`events.dispatch`) | [`tui`](../packages/ui/tui) |
| `agent/cancel-requested` | `emit` | [`packages/core/agent/src/types.ts:308`](../packages/core/agent/src/types.ts) | [`agent-loop`](../packages/core/agent-loop) (`emitAgentEvent`) | [`goal-session`](../packages/goal/goal-session) |
| `agent/created` | `emit` | [`packages/core/agent/src/types.ts:247`](../packages/core/agent/src/types.ts) | [`agent`](../packages/core/agent) (`events.dispatch`) | [`goal-session`](../packages/goal/goal-session), [`tui`](../packages/ui/tui) |
| `agent/created` | `emit` | [`packages/core/agent/src/types.ts:247`](../packages/core/agent/src/types.ts) | [`agent`](../packages/core/agent) (`events.dispatch`) | `apiproxy`, [`goal-session`](../packages/goal/goal-session), [`tui`](../packages/ui/tui) |
| `agent/disposed` | `emit` | [`packages/core/agent/src/types.ts:256`](../packages/core/agent/src/types.ts) | [`agent`](../packages/core/agent) (`events.dispatch`) | [`agent-loop`](../packages/core/agent-loop), [`goal-session`](../packages/goal/goal-session), [`tui`](../packages/ui/tui) |
| `agent/error` | `emit` | [`packages/core/agent/src/types.ts:423`](../packages/core/agent/src/types.ts) | [`agent-loop`](../packages/core/agent-loop) (`emitAgentEvent`) | `apiproxy`, [`goal-session`](../packages/goal/goal-session), [`session-telemetry`](../packages/telemetry/session-telemetry), [`tui`](../packages/ui/tui) |
| `agent/inbox/dequeue` | `emit` | [`packages/core/agent/src/types.ts:286`](../packages/core/agent/src/types.ts) | [`agent-loop`](../packages/core/agent-loop) (`emitAgentEvent`) | [`agent`](../packages/core/agent), `apiproxy`, [`tui`](../packages/ui/tui) |
+1 -1
View File
@@ -11,7 +11,7 @@ export type {
WorkspaceApi, WorkspaceId, WorkspaceView,
CommandsApi, CommandDescriptor, CommandExecuteResult, SkillsApi, SkillEntry,
ModelCatalogFailure, ModelCatalogModel, ModelProviderGroup, ModelReasoning,
ModelReasoningEffort, ModelTarget, SessionModels,
ModelReasoningEffort, ModelTarget, SessionMetrics, SessionModels,
} from '@deepseek-ai/dsh-host-apiproxy/api'
export type { ToolCallView, ToolResultView } from '@deepseek-ai/dsh-tools/presentation'
export type {
@@ -16,7 +16,7 @@ export type {
ToolCallView, ToolResultView, WorkspaceApi, WorkspaceId, WorkspaceView,
CommandsApi, CommandDescriptor, CommandExecuteResult, SkillsApi, SkillEntry,
ModelCatalogFailure, ModelCatalogModel, ModelProviderGroup, ModelReasoning,
ModelReasoningEffort, ModelTarget, SessionModels,
ModelReasoningEffort, ModelTarget, SessionMetrics, SessionModels,
RpcRequest, RpcResponse, RpcResult, RpcError, RpcErrorCode,
ClientRequest, ServerResponse, ServerRequest, ClientResponse, RpcMessage, RpcReceipt,
IApiClient, SessionId, SessionEvent, ContentBlock, StreamChunk,
+2 -2
View File
@@ -2,5 +2,5 @@
# side as of the last confirmed-consistent state. Both languages carry equal authority;
# after editing either side, bring the other along and re-record with:
# pnpm run verify-translation-pairing --write packages/client/runtime/README.md
README.md: 81261945cb2fd8b15f7c2f15cb1ae0b8e9928499
README.zh.md: cbbf6eded4a5375223791275f26f3bc7b6553200
README.md: fbb3a142b9efa50045f127d430bd2be2849f01f2
README.zh.md: b5bd3c458ed462e01dfae5bcd7399dde5552dc8d
+1 -1
View File
@@ -2,7 +2,7 @@
English | [中文](README.zh.md)
Client cordis boot and React-free object services: SlotsService wraps SlotCore and supplies renderer data sources; SessionsService owns Session objects, list/scope/history state; WorkspacesService depends on SessionsService and owns Workspace objects, list/actions, default-target derivation, and the New Session blank-reuse entry (`connectWorkspace`). The runtime fans the shared Host stream into both managers. Client sessions are always Host-born (Session+Agent+cwd in one `session.create`); the client holds no pre-entity session state — a session's Agent scope (the client mirror of host dsh-scope, keyed by the shared agent/session id) is born when its row enters the list mirror and dies with the prune. Contract: api-contracts v3 §4. `ConversationSnapshot` carries `todos` — the session's current todo projection: taken from the tail history page's full-log value (host-computed, independent of the page window), preserved across an older-page prepend, and overwritten by each live `todo/write` (last write wins). A tail response that omits the field means the log holds no `todo/write`, so the list resets to empty — a plan the log never kept (a write lost to a host crash) disappears on the next open or resync.
Client cordis boot and React-free object services: SlotsService wraps SlotCore and supplies renderer data sources; SessionsService owns Session objects, list/scope/history state; WorkspacesService depends on SessionsService and owns Workspace objects, list/actions, default-target derivation, and the New Session blank-reuse entry (`connectWorkspace`). The runtime fans the shared Host stream into both managers. Client sessions are always Host-born (Session+Agent+cwd in one `session.create`); the client holds no pre-entity session state — a session's Agent scope (the client mirror of host dsh-scope, keyed by the shared agent/session id) is born when its row enters the list mirror and dies with the prune. Contract: api-contracts v3 §4. `ConversationSnapshot` carries two Host-owned full-log projections. `todos` comes from the tail history page, survives older-page prepend, and follows live `todo/write` events. `metrics` comes from tail history and live `session/metrics` frames, survives older-page prepend, and accepts only nondecreasing log and projection revisions; a subscription baseline clears it before replay so a new stream generation can restart revisions safely. Missing metrics remain `null` rather than being inferred from the visible node window.
## Workspace and Session lists
+1 -1
View File
@@ -2,7 +2,7 @@
[English](README.md) | 中文
客户端 cordis 启动与不依赖 React 的对象服务:SlotsService 包装 SlotCore 并提供 renderer 数据源;SessionsService 拥有 Session 对象、列表/scopehistory 状态;WorkspacesService 依赖 SessionsService,拥有 Workspace 对象、列表/操作、默认目标派生,以及 New Session 空会话复用入口(`connectWorkspace`)。运行时把共享 Host 流分发给两个 manager。客户端 Session 一律由 Host 出生(一次 `session.create` 同瞬产出 Session+Agent+cwd);客户端不持有任何实体化之前的会话状态——Agent scopehost dsh-scope 的客户端镜像,以 agent/session 共用 id 为键)在会话行进入列表镜像时出生,随 prune 死亡。契约:api-contracts v3 §4。`ConversationSnapshot` 携带 `todos`——会话当前的 todo 投影:取自尾页 history 携带的全量 log 值(host 计算,独立于分页窗口),跨往前翻页保留,并被每次实时 `todo/write` 覆盖(后写胜出)。尾页响应省略该字段即表示 log 中没有任何 `todo/write`,因此列表复位为空——log 从未留下的计划(写入因 host 崩溃丢失)会在下一次打开或 resync 时消失
客户端 cordis 启动与不依赖 React 的对象服务:SlotsService 包装 SlotCore 并提供 renderer 数据源;SessionsService 拥有 Session 对象、列表/scopehistory 状态;WorkspacesService 依赖 SessionsService,拥有 Workspace 对象、列表/操作、默认目标派生,以及 New Session 空会话复用入口(`connectWorkspace`)。运行时把共享 Host 流分发给两个 manager。客户端 Session 一律由 Host 出生(一次 `session.create` 同瞬产出 Session+Agent+cwd);客户端不持有任何实体化之前的会话状态——Agent scopehost dsh-scope 的客户端镜像,以 agent/session 共用 id 为键)在会话行进入列表镜像时出生,随 prune 死亡。契约:api-contracts v3 §4。`ConversationSnapshot` 携带两项由 Host 拥有的完整日志投影。`todos` 来自 history 尾页,在向前加载较早页面时保留,并实时 `todo/write` 事件更新。`metrics` 来自 history 尾页和实时 `session/metrics` 帧,在向前加载较早页面时保留,并且只接受日志修订号与投影修订号均不减小的数据;订阅基线会在回放前将其清除,使新的流代次可以安全地从头开始计数修订号。缺失的 metrics 保持为 `null`,而不是根据可见节点窗口推断
## Workspace 与 Session 列表
@@ -6,7 +6,7 @@
import type { ContentBlock } from '@deepseek-ai/dsh-llm/types'
import type { TodoItem } from '@deepseek-ai/dsh-session/types'
import type {
RpcError, SessionId, ToolCallView, ToolResultView,
RpcError, SessionId, SessionMetrics, ToolCallView, ToolResultView,
} from '@deepseek-ai/dsh-client-connection/client'
import type { PendingInteraction } from './pending.ts'
@@ -246,4 +246,10 @@ export interface ConversationSnapshot {
/** Current whole-list `todo/write` projection — the tail page's full-log value, then each live
* write (last write wins); empty = the log holds no plan. */
todos: readonly TodoItem[]
/**
* Host-owned cumulative usage and current-context projection. Independent
* of `nodes` pagination; null until a tail response or live metrics frame
* supplies a current value.
*/
metrics: SessionMetrics | null
}
@@ -5,7 +5,7 @@ import type { ContentBlock } from '@deepseek-ai/dsh-llm/types'
import type { SessionEvent, TodoItem } from '@deepseek-ai/dsh-session/types'
import type {
HistoryEntry, IApiClient, MuxFrame, RpcError, RpcId, RpcResult,
SessionId, ToolEventView,
SessionId, SessionMetrics, ToolEventView,
} from '@deepseek-ai/dsh-client-connection/client'
// Value import from the inline-safe wire layer (not the connection plugin):
// plugin-to-plugin value imports are a bundle purity error.
@@ -102,6 +102,8 @@ export class Session implements ObservableSnapshot<ConversationSnapshot> {
/** Current whole-list todo/write projection: each tail history response replaces it (an omitted
* field is the authoritative empty list) and every live write overwrites it. */
private todos: readonly TodoItem[] = []
/** Host-owned metrics projection; ordering resets on each subscribed baseline. */
private metrics: SessionMetrics | null = null
/** `run_code` sub-dispatches by parent callId (window-derived, like openCalls). Appends
* copy-on-write the per-parent array so published snapshot references never mutate. */
private codeDispatches = new Map<string, readonly CodeSubCall[]>()
@@ -296,6 +298,7 @@ export class Session implements ObservableSnapshot<ConversationSnapshot> {
this.events = []
this.views = []
this.baseSeq = 0
this.metrics = null
// Superseded, not settled: the baseline replay re-sends still-pending requested frames verbatim
// (same rpcId), re-minting fresh waits; a stale reference's respond() still reaches the host.
this.pending.clear()
@@ -364,6 +367,14 @@ export class Session implements ObservableSnapshot<ConversationSnapshot> {
this.queueRev++
this.notifier.markDirty()
}
if (this.metrics !== null) {
this.metrics = null
this.notifier.markDirty()
}
return
}
case 'session/metrics': {
this.installMetrics(frame.metrics)
return
}
case 'approval/requested': {
@@ -482,13 +493,20 @@ export class Session implements ObservableSnapshot<ConversationSnapshot> {
this.openError = result.error
return
}
this.installWindow(result.value.events, result.value.hasMore, result.value.todos)
this.installWindow(result.value.events, result.value.hasMore, result.value.todos, result.value.metrics)
// Gap detection (§D.3-4): baseline past the window tail and liveBuffer did not cover it -> pull the tail page once more.
const tailSeq = this.windowTailSeq()
if (this.subscribedLastSeq !== null && tailSeq !== null && this.subscribedLastSeq > tailSeq) {
result = (await this.api.sessions.history({ sessionId: this.sessionId, maxMessages: PAGE_MESSAGES })).result
if (generation !== this.openGeneration) return
if (result.ok) this.installWindow(result.value.events, result.value.hasMore, result.value.todos)
if (result.ok) {
this.installWindow(
result.value.events,
result.value.hasMore,
result.value.todos,
result.value.metrics,
)
}
}
this.openState = 'open'
} catch (error) {
@@ -506,7 +524,12 @@ export class Session implements ObservableSnapshot<ConversationSnapshot> {
* Stitching MUST NOT route through acceptLiveEvent: openState is still 'loading' here
* (doOpen flips it after install), so recursing would push every buffered event straight
* back into liveBuffer where nothing ever drains it — a silent drop loop (audit S1). */
private installWindow(entries: HistoryEntry[], hasMore: boolean, todos: readonly TodoItem[] | undefined): void {
private installWindow(
entries: HistoryEntry[],
hasMore: boolean,
todos: readonly TodoItem[] | undefined,
metrics: SessionMetrics | undefined,
): void {
this.events = entries.map(e => e.event)
this.views = entries.map(e => e.view)
this.baseSeq = this.events[0]?.seq ?? 0
@@ -519,6 +542,7 @@ export class Session implements ObservableSnapshot<ConversationSnapshot> {
// field is the authoritative empty list, not a missing carrier. Assigning
// it clears a plan the log never kept (a write lost to a host crash).
this.todos = todos ?? []
if (metrics !== undefined) this.installMetrics(metrics)
this.foldAdapter.reset(this.events, this.baseSeq, this.views)
this.rebuildDerivedFromWindow()
const buffered = this.liveBuffer
@@ -569,7 +593,12 @@ export class Session implements ObservableSnapshot<ConversationSnapshot> {
const { result } = await this.api.sessions.history({ sessionId: this.sessionId, maxMessages: PAGE_MESSAGES })
// Failure or superseded by a full resync: drop — the resync path rebuilds and clears the buffer itself.
if (result.ok && generation === this.openGeneration && this.openState === 'open') {
this.installWindow(result.value.events, result.value.hasMore, result.value.todos)
this.installWindow(
result.value.events,
result.value.hasMore,
result.value.todos,
result.value.metrics,
)
}
} catch (error) {
console.error('[web-runtime] gap repair failed:', error)
@@ -761,6 +790,20 @@ export class Session implements ObservableSnapshot<ConversationSnapshot> {
return tail === undefined ? null : tail.seq
}
/** Install a metrics snapshot unless a newer durable or publication revision already landed. */
private installMetrics(metrics: SessionMetrics): void {
const current = this.metrics
if (
current !== null
&& (
metrics.logRevision < current.logRevision
|| metrics.projectionRevision < current.projectionRevision
)
) return
this.metrics = metrics
this.notifier.markDirty()
}
private buildSnapshot(): ConversationSnapshot {
const { nodes: folded, degraded } = this.foldAdapter.nodes()
// Frozen interrupted nodes ride fractional seqs: a stable merge keeps them in flow order.
@@ -811,6 +854,7 @@ export class Session implements ObservableSnapshot<ConversationSnapshot> {
blank: this.blankBit,
lastAgentError: this.lastAgentError,
todos: this.todos,
metrics: this.metrics,
}
}
}
+7 -2
View File
@@ -3,7 +3,7 @@
// deferred-controlled timing). Streams are hand pumps: pushMux/pushHost.
import type {
ClientResponse, CommandDescriptor, CommandExecuteResult, HostFrame, IApiClient, ModelTarget, MuxFrame,
RpcError, RpcReceipt, RpcRequest, RpcResponse, SessionId, SessionModels, SkillEntry,
RpcError, RpcReceipt, RpcRequest, RpcResponse, SessionId, SessionMetrics, SessionModels, SkillEntry,
WorkspaceId, WorkspaceView,
} from '@deepseek-ai/dsh-client-connection/client'
import { RpcId } from '@deepseek-ai/dsh-client-connection/client'
@@ -63,7 +63,12 @@ export class FakeApiClient implements IApiClient {
onCreate: (payload: unknown) => Promise<RpcResponse<{ sessionId: SessionId }>> = () => Promise.resolve(ok({ sessionId: 'fk-new' as SessionId }))
readonly defaultModel: ModelTarget = { provider: 'deepseek', model: 'deepseek-v4-flash' }
onHistory: (payload: { sessionId: SessionId; beforeSeq?: number; maxMessages?: number })
=> Promise<RpcResponse<{ events: never[]; hasMore: boolean; todos?: { content: string; status: 'pending' | 'in_progress' | 'completed' }[] }>> =
=> Promise<RpcResponse<{
events: never[]
hasMore: boolean
todos?: { content: string; status: 'pending' | 'in_progress' | 'completed' }[]
metrics?: SessionMetrics
}>> =
() => Promise.resolve(ok({ events: [], hasMore: false }))
onModels: (payload: unknown) => Promise<RpcResponse<SessionModels>> = () => Promise.resolve(ok({
@@ -85,6 +85,13 @@ describe('queue retirement (host queuedMirror rules)', () => {
expect(session.getSnapshot().queue).toHaveLength(1)
})
it('an unrelated durable event leaves the queue unchanged', () => {
const session = makeSession()
session.handleMuxEnvelope(rid('e1'), queuedFrame('留', 'p-1'))
session.handleMuxEnvelope(rid('e2'), { type: 'session/event', sessionId: SID, event: ev.user(0, 'unrelated') })
expect(session.getSnapshot().queue.map(row => row.key)).toEqual(['p-1'])
})
it('steering/message drains the source-matched steering row only', () => {
const session = makeSession()
session.handleMuxEnvelope(rid('e1'), queuedFrame('普通', 'p-1')) // idle → non-steering
+109 -3
View File
@@ -7,8 +7,9 @@
*/
import { describe, expect, it, vi } from 'vitest'
import { Context } from 'cordis'
import type { SessionEvent } from '@deepseek-ai/dsh-session/types'
import type { SessionId } from '@deepseek-ai/dsh-client-connection/client'
import type { SessionId, SessionMetrics } from '@deepseek-ai/dsh-client-connection/client'
import { Session } from '../src/client/sessions/session.ts'
import { FakeApiClient, deferred, err, ok } from './fake-api.ts'
import { entries, ev, plainTurn } from './event-script.ts'
@@ -22,9 +23,37 @@ function makeSession(api = new FakeApiClient()): { api: FakeApiClient; session:
return { api, session: new Session(SID, api) }
}
function histResponse(events: SessionEvent[], hasMore = false, todos?: { content: string; status: 'pending' | 'in_progress' | 'completed' }[]) {
function histResponse(
events: SessionEvent[],
hasMore = false,
todos?: { content: string; status: 'pending' | 'in_progress' | 'completed' }[],
metrics?: SessionMetrics,
) {
// history now returns HistoryEntry[] ({event, view?}); these tests are view-less.
return Promise.resolve(ok({ events: entries(events) as never[], hasMore, ...todos === undefined ? {} : { todos } }))
return Promise.resolve(ok({
events: entries(events) as never[],
hasMore,
...todos === undefined ? {} : { todos },
...metrics === undefined ? {} : { metrics },
}))
}
function metrics(
projectionRevision: number,
logRevision: number,
over: Partial<SessionMetrics> = {},
): SessionMetrics {
return {
projectionRevision,
logRevision,
uncachedInputTokens: 10,
outputTokens: 4,
cacheReadTokens: 90,
cacheWriteTokens: 3,
contextTokens: 35,
contextWindow: 100,
...over,
}
}
describe('open', () => {
@@ -40,6 +69,19 @@ describe('open', () => {
expect(snapshot.openState).toBe('open')
expect(snapshot.hasMore).toBe(true)
expect(snapshot.nodes.map(n => n.kind)).toEqual(['user', 'assistant'])
expect(snapshot.metrics).toBeNull()
})
it('installs full-log metrics independently of older history pages', async () => {
const { api, session } = makeSession()
const tailMetrics = metrics(4, 106)
api.onHistory = () => histResponse(plainTurn(100, 3, '问', '答'), true, undefined, tailMetrics)
await session.open()
expect(session.getSnapshot().metrics).toBe(tailMetrics)
api.onHistory = () => histResponse(plainTurn(94, 2, '旧问', '旧答'))
await session.loadOlder()
expect(session.getSnapshot().metrics).toBe(tailMetrics)
})
it('is idempotent: concurrent opens share one history call, reopening when open is a no-op', async () => {
@@ -104,6 +146,43 @@ describe('live event path', () => {
expect(session.getSnapshot().nodes).toEqual(before.nodes)
})
it('orders live metrics, rejects stale projections, and clears the value at a reconnect baseline', async () => {
const { session } = await opened()
const current = metrics(8, 10)
session.handleMuxEnvelope('m1' as never, {
type: 'session/metrics',
sessionId: SID,
metrics: current,
})
expect(session.getSnapshot().metrics).toBe(current)
session.handleMuxEnvelope('m2' as never, {
type: 'session/metrics',
sessionId: SID,
metrics: metrics(9, 9, { uncachedInputTokens: 1 }),
})
session.handleMuxEnvelope('m3' as never, {
type: 'session/metrics',
sessionId: SID,
metrics: metrics(7, 11, { uncachedInputTokens: 2 }),
})
expect(session.getSnapshot().metrics).toBe(current)
session.handleMuxEnvelope('sub' as never, {
type: 'session/subscribed',
sessionId: SID,
lastSeq: 5,
})
expect(session.getSnapshot().metrics).toBeNull()
const nextGeneration = metrics(0, 10, { contextTokens: 20 })
session.handleMuxEnvelope('m4' as never, {
type: 'session/metrics',
sessionId: SID,
metrics: nextGeneration,
})
expect(session.getSnapshot().metrics).toBe(nextGeneration)
})
it('accumulates chunks into partial, then finalize swaps partial out as the node lands', async () => {
const { session } = await opened()
const feed = (event: SessionEvent) => { session.handleMuxEnvelope('r' as never, { type: 'session/event', sessionId: SID, event }) }
@@ -332,6 +411,23 @@ describe('prompt and cancel errors', () => {
})
describe('pending interactions', () => {
it('routes an approval wait response through the original requested rpcId', async () => {
const { api, session } = makeSession()
session.handleMuxEnvelope('ra-answer' as never, {
type: 'approval/requested',
sessionId: SID,
approvalId: 'ap-answer' as never,
toolName: 'bash',
})
const wait = session.getSnapshot().pending[0]!
await wait.respond({ ok: true, value: { decision: 'allow' } })
expect(api.callsOf('respond')).toEqual([{
type: 'client-response',
rpcId: 'ra-answer',
result: { ok: true, value: { decision: 'allow' } },
}])
})
it('adds approval/question on requested and removes them on resolved', async () => {
const { session } = makeSession()
session.handleMuxEnvelope('ra' as never, { type: 'approval/requested', sessionId: SID, approvalId: 'ap1' as never, toolName: 'rm' })
@@ -374,6 +470,16 @@ describe('pending interactions', () => {
})
describe('remaining branches', () => {
it('rejects a second scope bind and allows rebinding after explicit release', () => {
const { session } = makeSession()
const first = new Context()
const second = new Context()
session.bindScope(first)
expect(() => { session.bindScope(second) }).toThrow(`session ${SID} already has a bound scope`)
session.unbindScope()
expect(() => { session.bindScope(second) }).not.toThrow()
})
it('prompt transport throw folds to internal promptError', async () => {
const { api, session } = makeSession()
api.onPrompt = () => Promise.reject(new Error('prompt wire down'))
@@ -2,5 +2,5 @@
# side as of the last confirmed-consistent state. Both languages carry equal authority;
# after editing either side, bring the other along and re-record with:
# pnpm run verify-translation-pairing --write packages/client/ui-conversation/README.md
README.md: 56a445ccfa86e0b11cf5aefc37819a30746f0739
README.zh.md: a7c160ecdd74074257c9d149630663dacd05c070
README.md: 9a8595e693e2b49691d0d130d4db0754e1e4829b
README.zh.md: f5080314488e807877ebf0c93ea82cdd9725e8ed
@@ -18,6 +18,8 @@ Per-session UI state for selection and the active view lives in the declared cha
The composer bar declares session-scoped single seats for `'conversation.input.plan'` and `'conversation.input.model'`, plus list slots for overlay, dock, left, and right input extensions. InputBar renders the model seat immediately before its pending indicator and send/stop button. Feature packages own each control and its state; ui-conversation supplies placement, the `locked` owner prop, and the standard slot shares. The resident no-session shell uses `DisabledInputBar` and therefore dispatches no session-scoped control seats.
The chat stats line reads durable token counters and current-context pressure only from `ConversationSnapshot.metrics`; visible nodes supply only the existing turn/step counts. It renders uncached input, output, and cache reads as separate compact values, computes cache hit as `cacheRead / (uncachedInput + cacheRead)` without cache writes, and shows context occupancy against the selected route's exact capacity. Missing host data is labeled unknown, never reconstructed from a paged window.
`src/client/` is organized for the future package split: `contract/` is the sole inter-domain shared face (`slots.ts` slot declarations + composed slot props including the tool-row contract, `views.ts` shared primitives, `tool-call-model.ts`); the `skeleton/`, `chat/`, and `toolviews/` (sample registrants) domain directories import contract files and never each other; `apply.ts` is the only assembly point allowed to import all three domains. The `/client` export surface is the contract only — `apply`/`inject`, the two service classes, and the `contract/` type families; implementation components (skeleton, chat rows) and the store factory stay internal and reach the page exclusively through apply's slot registrations (tests take them via the `./src/*` subpath).
## Model Experience
@@ -18,6 +18,8 @@ todo 两个面就是在该形状上的两个注册项,都是普通注册方插
输入栏为 `'conversation.input.plan'``'conversation.input.model'` 声明会话作用域的单实例 seat,并为 overlay、dock、left 和 right 输入扩展声明列表 slot。InputBar 将模型 seat 渲染在 pending 指示器与发送/停止按钮之前。各功能包拥有相应控件及其状态;ui-conversation 提供放置位置、`locked` owner prop 和标准 slot share。常驻无会话壳使用 `DisabledInputBar`,因此不会分发任何会话作用域的控件 seat。
聊天统计行只从 `ConversationSnapshot.metrics` 读取持久的 token 计数与当前上下文压力;可见节点仅提供既有的轮次和步骤计数。它以相互独立的紧凑值显示未缓存输入、输出与缓存读取,通过 `cacheRead / (uncachedInput + cacheRead)` 计算缓存命中率而不计入缓存写入,并根据所选路由的精确容量显示上下文占用率。Host 数据缺失时标为「未知」,绝不根据分页窗口重建。
`src/client/` 按未来的包拆分组织:`contract/` 是唯一的跨领域共享表层(`slots.ts` slot 声明 + 组合后的 slot props,包括工具行契约、`views.ts` 共享原语、`tool-call-model.ts`);`skeleton/``chat/``toolviews/`(示例注册方)领域目录只导入 contract 文件,彼此绝不导入;`apply.ts` 是唯一允许导入全部三个领域的组装点。`/client` 导出表层只包含契约:`apply``inject`、两个服务类和 `contract/` 类型家族;实现组件(骨架、聊天行)与 store factory 保持内部状态,只能通过 apply 的 slot 注册到达页面(测试通过 `./src/*` 子路径获取它们)。
## 模型体验
@@ -5,48 +5,63 @@ import type { ConversationSnapshot } from '@deepseek-ai/dsh-client-runtime/clien
import type { SnapshotSelectorHook } from '@deepseek-ai/dsh-client-ui-slots'
import css from './StatsLine.module.css'
interface UsageTotals {
type SessionMetrics = NonNullable<ConversationSnapshot['metrics']>
interface VisibleCounts {
turns: number
steps: number
tokens: number
cacheHitPct: number | null
}
/** Token accounting slice of assistant `usage` (typed upstream as unknown). */
interface UsageLike {
inputTokens?: number
outputTokens?: number
cacheReadTokens?: number
}
/**
* Fold assistant nodes into display totals.
* Count visible assistant turns and steps without treating the paged window
* as an accounting source.
* @param nodes - snapshot nodes.
* @returns totals; cacheHitPct null until any cache accounting arrives.
* @returns visible turn and step counts.
*/
export function deriveStats(nodes: ConversationSnapshot['nodes']): UsageTotals {
export function deriveVisibleCounts(nodes: ConversationSnapshot['nodes']): VisibleCounts {
const turns = new Set<number>()
let steps = 0
let tokens = 0
let input = 0
let cacheRead = 0
for (const node of nodes) {
if (node.kind !== 'assistant') continue
turns.add(node.turn)
steps += 1
const usage = node.usage as UsageLike | undefined
if (usage === undefined) continue
input += usage.inputTokens ?? 0
cacheRead += usage.cacheReadTokens ?? 0
tokens += (usage.inputTokens ?? 0) + (usage.outputTokens ?? 0) + (usage.cacheReadTokens ?? 0)
}
const denom = input + cacheRead
return {
turns: turns.size,
steps,
tokens,
cacheHitPct: denom === 0 ? null : Math.round((cacheRead / denom) * 100),
}
return { turns: turns.size, steps }
}
/**
* Format large token values with the status surfaces' compact suffix style.
* @param value - token count or model capacity.
* @returns locale-formatted count.
*/
export function formatMetricTokens(value: number): string {
if (value < 1_000) return value.toLocaleString('en-US')
return value.toLocaleString('en-US', {
notation: 'compact',
maximumFractionDigits: 1,
}).replace('K', 'k').replace('M', 'm').replace('B', 'b')
}
/**
* Existing Web cache-hit formula over disjoint uncached and cache-read input.
* @param metrics - Host-owned durable usage.
* @returns rounded integer percent, or null when no input was billed.
*/
export function cacheHitPercent(metrics: SessionMetrics): number | null {
const denominator = metrics.uncachedInputTokens + metrics.cacheReadTokens
return denominator === 0
? null
: Math.round(metrics.cacheReadTokens / denominator * 100)
}
/**
* Current context occupancy using the TUI's integer rounding and upper clamp.
* @param metrics - Host-owned current pressure and exact route capacity.
* @returns occupancy percent, or null when either input is unavailable.
*/
export function contextPercent(metrics: SessionMetrics): number | null {
if (metrics.contextTokens === undefined || metrics.contextWindow === undefined) return null
return Math.min(100, Math.round(metrics.contextTokens / metrics.contextWindow * 100))
}
/** Props: the conversation-snapshot selector hook (handed down by ChatView). */
@@ -54,12 +69,33 @@ export interface StatsLineProps { useSession: SnapshotSelectorHook<ConversationS
export const StatsLine = memo(function StatsLine({ useSession }: StatsLineProps) {
const nodes = useSession(s => s.nodes)
const stats = useMemo(() => deriveStats(nodes), [nodes])
if (stats.steps === 0) return null
const metrics = useSession(s => s.metrics)
const counts = useMemo(() => deriveVisibleCounts(nodes), [nodes])
if (counts.steps === 0 && (
metrics === null
|| (
metrics.uncachedInputTokens === 0
&& metrics.outputTokens === 0
&& metrics.cacheReadTokens === 0
&& (metrics.contextTokens ?? 0) === 0
)
)) return null
const parts: string[] = []
if (stats.cacheHitPct !== null) parts.push(`cache hit ${stats.cacheHitPct}%`)
parts.push(`${stats.tokens.toLocaleString('en-US')} tokens`)
parts.push(`${stats.turns} turns`)
parts.push(`${stats.steps} steps`)
if (metrics === null) {
parts.push('usage unknown')
parts.push('context unknown')
} else {
parts.push(`${formatMetricTokens(metrics.uncachedInputTokens)} uncached input`)
parts.push(`${formatMetricTokens(metrics.outputTokens)} output`)
parts.push(`${formatMetricTokens(metrics.cacheReadTokens)} cache read`)
const cacheHit = cacheHitPercent(metrics)
if (cacheHit !== null) parts.push(`cache hit ${cacheHit}%`)
const context = contextPercent(metrics)
parts.push(context === null
? 'context unknown'
: `context ${context}% of ${formatMetricTokens(metrics.contextWindow as number)}`)
}
parts.push(`${counts.turns} turns`)
parts.push(`${counts.steps} steps`)
return <div className={css.root}>{parts.join(' · ')}</div>
})
@@ -129,15 +129,23 @@ describe('small branch tails', () => {
})
it('StatsLine omits the cache-hit segment when no input accounting exists at all', () => {
// cacheHitPct is null only when input+cacheRead are both zero (pure
// output accounting) — any input makes it a real 0%.
const snap = {
nodes: [{ kind: 'assistant', seq: 1, turn: 1, step: 1, blocks: [], usage: { outputTokens: 10 } }],
metrics: {
logRevision: 2,
projectionRevision: 0,
uncachedInputTokens: 0,
outputTokens: 10,
cacheReadTokens: 0,
cacheWriteTokens: 5_000,
},
}
const source = { getSnapshot: () => snap, subscribe: () => () => {} }
const view = render(
<StatsLine useSession={bindSnapshotSelector(source) as unknown as StatsLineProps['useSession']} />,
)
expect(view.getByText('10 tokens · 1 turns · 1 steps')).toBeTruthy()
expect(view.getByText(
'0 uncached input · 10 output · 0 cache read · context unknown · 1 turns · 1 steps',
)).toBeTruthy()
})
})
@@ -58,7 +58,7 @@ function snapshotWith(
sessionId: SID, nodes, foldDegraded: false, partial: null, runningCalls, codeDispatches,
pending: [], queue: [], todos: [], running: runningCalls.length > 0, composerPhase: 'active', removed: false,
openState: 'open', openError: null,
hasMore: false, loadingOlder: false, promptError: null, blank: false, lastAgentError: null,
hasMore: false, loadingOlder: false, promptError: null, blank: false, lastAgentError: null, metrics: null,
}
}
@@ -1,5 +1,5 @@
// @vitest-environment jsdom
// StatsLine (rendered inside the chat view body): totals derivation + the RFC
// StatsLine (rendered inside the chat view body): durable metrics presentation + the RFC
// hard acceptance — zero renders during streaming. Bash sample row: the
// canonical sub-agent differential decided INSIDE the component off the
// standard useSessions kit (no registry predicates — tool ring dissolved).
@@ -12,7 +12,10 @@ import type {
import { createSnapshotStore } from '@deepseek-ai/dsh-client-runtime/client'
import { bindSnapshotSelector } from '@deepseek-ai/dsh-client-web-react'
import type { ToolRowProps } from '@deepseek-ai/dsh-client-ui-conversation/client'
import { StatsLine, deriveStats, type StatsLineProps } from '../src/client/chat/StatsLine.tsx'
import {
cacheHitPercent, contextPercent, deriveVisibleCounts, formatMetricTokens,
StatsLine, type StatsLineProps,
} from '../src/client/chat/StatsLine.tsx'
import { BashRow } from '../src/client/toolviews/bash-sample.tsx'
afterEach(cleanup)
@@ -28,7 +31,7 @@ function snapshotBase(): ConversationSnapshot {
return {
sessionId: SID, nodes: [], foldDegraded: false, partial: null, runningCalls: [], codeDispatches: new Map(),
pending: [], queue: [], todos: [], running: false, composerPhase: 'active', removed: false, openState: 'open', openError: null,
hasMore: false, loadingOlder: false, promptError: null, blank: false, lastAgentError: null,
hasMore: false, loadingOlder: false, promptError: null, blank: false, lastAgentError: null, metrics: null,
}
}
@@ -50,27 +53,49 @@ function makeSource(init?: Partial<ConversationSnapshot>) {
}
}
describe('deriveStats', () => {
it('folds turns/steps/tokens and cache hit percentage', () => {
const stats = deriveStats([
describe('stats derivation', () => {
it('counts visible turns and steps without reading node usage', () => {
const stats = deriveVisibleCounts([
assistant(1, 1, { inputTokens: 100, outputTokens: 50, cacheReadTokens: 900 }),
assistant(2, 1, { inputTokens: 100, outputTokens: 50 }),
assistant(3, 2),
])
expect(stats.turns).toBe(2)
expect(stats.steps).toBe(3)
expect(stats.tokens).toBe(1200)
expect(stats.cacheHitPct).toBe(82)
})
it('cache hit stays null with no cache accounting; non-assistant nodes ignored', () => {
it('ignores non-assistant nodes', () => {
const tool: ToolResultNode = {
kind: 'tool-result', seq: 5, time: 5_000, callId: 'c', call: null, callTime: null, content: [],
isError: false, callView: null, resultView: null,
}
const stats = deriveStats([tool, assistant(1, 1)])
const stats = deriveVisibleCounts([tool, assistant(1, 1)])
expect(stats.steps).toBe(1)
expect(stats.cacheHitPct).toBeNull()
})
it('keeps the cache formula disjoint from cache writes and rounds/clamps context like the TUI', () => {
const durable = {
logRevision: 20,
projectionRevision: 2,
uncachedInputTokens: 100,
outputTokens: 50,
cacheReadTokens: 900,
cacheWriteTokens: 50_000,
contextTokens: 34_500,
contextWindow: 100_000,
}
expect(cacheHitPercent(durable)).toBe(90)
expect(contextPercent(durable)).toBe(35)
expect(contextPercent({ ...durable, contextTokens: 200_000 })).toBe(100)
const { contextWindow: _contextWindow, ...withoutContextWindow } = durable
expect(contextPercent(withoutContextWindow)).toBeNull()
expect(cacheHitPercent({ ...durable, uncachedInputTokens: 0, cacheReadTokens: 0 })).toBeNull()
})
it('formats large values compactly in the existing en-US style', () => {
expect(formatMetricTokens(999)).toBe('999')
expect(formatMetricTokens(15_962)).toBe('16k')
expect(formatMetricTokens(2_172_544)).toBe('2.2m')
})
})
@@ -79,19 +104,65 @@ describe('StatsLine', () => {
return { useSession: bindSnapshotSelector(source) }
}
it('renders the joined stats row and hides with zero steps', () => {
it('renders separate durable counters, cache hit, context occupancy, and visible counts', () => {
const { source } = makeSource({
nodes: [assistant(1, 1, { inputTokens: 10, outputTokens: 5, cacheReadTokens: 90 })],
metrics: {
logRevision: 30,
projectionRevision: 4,
uncachedInputTokens: 120_237,
outputTokens: 13_881,
cacheReadTokens: 2_172_544,
cacheWriteTokens: 99_999,
contextTokens: 89_600,
contextWindow: 256_000,
},
})
const view = render(<StatsLine {...props(source)} />)
expect(view.getByText('cache hit 90% · 105 tokens · 1 turns · 1 steps')).toBeTruthy()
expect(view.getByText(
'120.2k uncached input · 13.9k output · 2.2m cache read · cache hit 95% · context 35% of 256k · 1 turns · 1 steps',
)).toBeTruthy()
const empty = makeSource()
const emptyView = render(<StatsLine {...props(empty.source)} />)
expect(emptyView.container.textContent).toBe('')
})
it('renders honest unknowns when the host projection is missing', () => {
const { source } = makeSource({ nodes: [assistant(1, 1)] })
const view = render(<StatsLine {...props(source)} />)
expect(view.getByText('usage unknown · context unknown · 1 turns · 1 steps')).toBeTruthy()
})
it.each([
{ uncachedInputTokens: 1, outputTokens: 0, cacheReadTokens: 0, contextTokens: 0 },
{ uncachedInputTokens: 0, outputTokens: 1, cacheReadTokens: 0, contextTokens: 0 },
{ uncachedInputTokens: 0, outputTokens: 0, cacheReadTokens: 1, contextTokens: 0 },
{ uncachedInputTokens: 0, outputTokens: 0, cacheReadTokens: 0, contextTokens: 1 },
])('keeps a metrics-only row visible for each nonzero projection bucket', (nonzero) => {
const { source } = makeSource({
metrics: {
logRevision: 1,
projectionRevision: 0,
cacheWriteTokens: 0,
...nonzero,
},
})
const view = render(<StatsLine {...props(source)} />)
expect(view.container.textContent).toContain('0 turns · 0 steps')
})
it('renders ZERO times during streaming chunk frames (RFC hard acceptance)', () => {
const { set, source } = makeSource({ nodes: [assistant(1, 1)] })
const { set, source } = makeSource({
nodes: [assistant(1, 1)],
metrics: {
logRevision: 4,
projectionRevision: 0,
uncachedInputTokens: 1,
outputTokens: 1,
cacheReadTokens: 0,
cacheWriteTokens: 0,
},
})
let renders = 0
function Counting(p: StatsLineProps) {
renders += 1
@@ -41,7 +41,7 @@ function snapshotWith(nodes: ToolResultNode[]): ConversationSnapshot {
return {
sessionId: SID, nodes, foldDegraded: false, partial: null, runningCalls: [], codeDispatches: new Map(),
pending: [], queue: [], todos: [], running: false, composerPhase: 'active', removed: false, openState: 'open', openError: null,
hasMore: false, loadingOlder: false, promptError: null, blank: false, lastAgentError: null,
hasMore: false, loadingOlder: false, promptError: null, blank: false, lastAgentError: null, metrics: null,
}
}
@@ -31,7 +31,7 @@ function snapshotBase(): ConversationSnapshot {
return {
sessionId: SID, nodes: [], foldDegraded: false, partial: null, runningCalls: [], codeDispatches: new Map(),
pending: [], queue: [], todos: [], running: false, composerPhase: 'active', removed: false, openState: 'open', openError: null,
hasMore: false, loadingOlder: false, promptError: null, blank: false, lastAgentError: null,
hasMore: false, loadingOlder: false, promptError: null, blank: false, lastAgentError: null, metrics: null,
}
}
@@ -20,7 +20,7 @@ function snapshotBase(): ConversationSnapshot {
return {
sessionId: SID, nodes: [], foldDegraded: false, partial: null, runningCalls: [], codeDispatches: new Map(),
pending: [], queue: [], todos: [], running: false, composerPhase: 'active', removed: false, openState: 'open', openError: null,
hasMore: false, loadingOlder: false, promptError: null, blank: false, lastAgentError: null,
hasMore: false, loadingOlder: false, promptError: null, blank: false, lastAgentError: null, metrics: null,
}
}
@@ -36,7 +36,7 @@ describe('render branch tails', () => {
expect(view.container.querySelector('[data-state="ok"]')).not.toBeNull()
})
it('StatsLine skips usage-less nodes and defaults each absent counter to zero', () => {
it('StatsLine takes durable counters from metrics while keeping visible node counts', () => {
const snap = {
nodes: [
{ kind: 'assistant', seq: 1, turn: 1, step: 1, blocks: [] },
@@ -44,12 +44,22 @@ describe('render branch tails', () => {
// outputTokens absent: the tokens sum's ?? 0 arm for output.
{ kind: 'assistant', seq: 3, turn: 2, step: 1, blocks: [], usage: { inputTokens: 5 } },
],
metrics: {
logRevision: 9,
projectionRevision: 1,
uncachedInputTokens: 9,
outputTokens: 6,
cacheReadTokens: 0,
cacheWriteTokens: 0,
},
}
const source = { getSnapshot: () => snap, subscribe: () => () => {} }
const view = render(
<StatsLine useSession={bindSnapshotSelector(source) as unknown as UseSession<ConversationSnapshot>} />,
)
expect(view.getByText('cache hit 0% · 15 tokens · 2 turns · 3 steps')).toBeTruthy()
expect(view.getByText(
'9 uncached input · 6 output · 0 cache read · cache hit 0% · context unknown · 2 turns · 3 steps',
)).toBeTruthy()
})
it('AssistantMarkdown reasoning as the streaming tail renders the running ring', () => {
@@ -23,7 +23,7 @@ function snapshotOf(overrides: Partial<ConversationSnapshot> = {}): Conversation
sessionId: SID, nodes: [], foldDegraded: false, partial: null, runningCalls: [], codeDispatches: new Map(),
pending: [], queue: [], todos: [], running: false, composerPhase: 'active', removed: false,
openState: 'open', openError: null, hasMore: false, loadingOlder: false,
promptError: null, blank: false, lastAgentError: null,
promptError: null, blank: false, lastAgentError: null, metrics: null,
...overrides,
}
}
@@ -26,7 +26,7 @@ function mountBar(shell: SessionInputShell, over?: { running?: boolean; disabled
sessionId: SID, nodes: [], foldDegraded: false, partial: null, runningCalls: [], codeDispatches: new Map(),
pending: [], queue: [], todos: [], running: over?.running ?? false, composerPhase: 'active',
removed: over?.disabled ?? false, openState: 'open', openError: null, hasMore: false,
loadingOlder: false, promptError: null, blank: false, lastAgentError: null,
loadingOlder: false, promptError: null, blank: false, lastAgentError: null, metrics: null,
})
const props: InputBarProps = {
sessionId: SID,
@@ -112,7 +112,7 @@ async function scopedBench(register?: (slash: SlashService) => void) {
sessionId, nodes: [], foldDegraded: false, partial: null, runningCalls: [], codeDispatches: new Map(),
pending: [], queue: [], todos: [], running: false, composerPhase: 'active', removed: false,
openState: 'open', openError: null, hasMore: false, loadingOlder: false,
promptError: null, blank: false, lastAgentError: null,
promptError: null, blank: false, lastAgentError: null, metrics: null,
})
const barProps: InputBarProps = {
sessionId,
@@ -20,7 +20,7 @@ function snapshotWith(queue: QueuedMessage[]): ConversationSnapshot {
return {
sessionId: SID, nodes: [], foldDegraded: false, partial: null, runningCalls: [], codeDispatches: new Map(),
pending: [], queue, todos: [], running: true, composerPhase: 'active', removed: false, openState: 'open', openError: null,
hasMore: false, loadingOlder: false, promptError: null, blank: false, lastAgentError: null,
hasMore: false, loadingOlder: false, promptError: null, blank: false, lastAgentError: null, metrics: null,
}
}
@@ -50,7 +50,7 @@ function conversationSnapshot(overrides: Partial<ConversationSnapshot> = {}): Co
sessionId: SID, nodes: [], foldDegraded: false, partial: null, runningCalls: [], codeDispatches: new Map(),
pending: [], queue: [], todos: [], running: false, composerPhase: 'active', removed: false,
openState: 'open', openError: null, hasMore: false, loadingOlder: false,
promptError: null, blank: false, lastAgentError: null,
promptError: null, blank: false, lastAgentError: null, metrics: null,
...overrides,
}
}
+2 -2
View File
@@ -2,5 +2,5 @@
# side as of the last confirmed-consistent state. Both languages carry equal authority;
# after editing either side, bring the other along and re-record with:
# pnpm run verify-translation-pairing --write packages/host/apiproxy/README.md
README.md: d6db9a9541b0727b61dbe501f7234564ffef139e
README.zh.md: 4175c8fdb98aad2882718a2c95cd9e45825d787d
README.md: fa6fcb17ad3f2332e71e00387034fd7196759d8f
README.zh.md: cfdf0d772ac81a4e569a799569b5de6e509f3b6f
+1 -1
View File
@@ -18,7 +18,7 @@ Workspace and Session lists are separate reconnect baselines. `workspace.create`
`host.pickDirectory` opens one native directory picker and returns its selected path, or `null` when the user cancels. Its host implementation invokes platform tools without a shell: `osascript` on macOS, an STA PowerShell `FolderBrowserDialog` on Windows, and Zenity with a KDialog fallback on Linux. The picker function is injectable for tests. This user-paced method is the sole unary call exempt from the default 30-second timeout; caller and connection aborts still propagate to the native process. The browser carrier separately restricts this privileged method to loopback, same-origin requests.
`session.history` pages on message boundaries, and its tail page (no `beforeSeq`) carries two session-level extras the page window cannot supply: the in-flight partial's chunk events, and `todos` the latest `todo/write` whole-list projection over the full log. Older pages omit `todos` because the projection is session-level, not per-page; a tail response that omits it means the whole log holds no `todo/write`, so clients read the absent field as the empty plan rather than as unchanged state.
`session.history` pages on message boundaries, and its tail page (no `beforeSeq`) carries session-level projections the page window cannot supply: the in-flight partial's chunk events; `todos`, the latest `todo/write` whole-list projection; and `metrics`, full-log usage deduplicated by `(turn, step)` plus current token-meter pressure and exact selected-route capacity when available. Older pages omit the session-level projections. Live `session/metrics` mux frames carry monotonic log/projection revisions, so clients reject stale frames and preserve the counters while prepending older pages. Cache reads and writes remain disjoint buckets; the cache-hit denominator is uncached input plus cache reads.
The `command.*` and `skill.*` domains expose the host command registry and skill catalog to clients. Every method addresses one session's agent by `sessionId` (a served session always has an Agent; `command.*` resumes cold sessions through the same path as `session.*`, while `skill.list` resolves the project root from the session header without touching the Agent registry). `command.execute` runs a slash-command line host-side and returns a detached result; the carrier's request signal cancels the running handler. `host/commands-changed` is the catalog invalidation frame: clients refetch `command.list` instead of diffing.
+1 -1
View File
@@ -18,7 +18,7 @@ Workspace 列表与 Session 列表是相互独立的重连基线。`workspace.cr
`host.pickDirectory` 会打开一个原生目录选择器并返回选中的路径;用户取消时返回 `null`。宿主实现不经 shell 调用平台工具:macOS 使用 `osascript`Windows 使用以 STA 模式运行的 PowerShell `FolderBrowserDialog`Linux 使用 Zenity,并以 KDialog 作为回退。选择器函数可在测试中注入。该方法需等待用户完成操作,是唯一不受默认 30 秒超时限制的一元调用;调用方发出的中止信号和连接中止仍会传播至原生进程。浏览器载体另行将这一特权方法限制为仅接受来自回环地址的同源请求。
`session.history` 按消息边界分页,其尾页(不带 `beforeSeq`额外携带两项页窗口本身无法提供的会话级数据:进行中局部消息的 chunk 事件,以及 `todos`——整份日志上最后一次 `todo/write` 的整表投影。较早的页面不带 `todos`,因为该投影是会话级而非分页级的;尾页响应缺少该字段意味着整份日志中没有任何 `todo/write`,因此客户端要把缺失字段读作空计划,而不是读作「状态未变」
`session.history` 按消息边界分页,其尾页(不带 `beforeSeq`携带页窗口本身无法提供的会话级投影:进行中局部消息的分片事件;`todos`,即最后一次 `todo/write` 的整表投影;以及 `metrics`,即按 `(turn, step)` 去重的完整日志用量,并在可用时包含当前 token 计量压力和所选精确路由的容量。较早的页面省略会话级投影。实时 `session/metrics` mux 帧携带单调递增的日志修订号与投影修订号,因此客户端会拒绝陈旧帧,并在向前加载较早页面时保留计数器。缓存读取与缓存写入保持为彼此独立的计数项;缓存命中率的分母是未缓存输入加缓存读取
`command.*``skill.*` 领域向客户端暴露宿主命令注册表和技能目录。每个方法都通过 `sessionId` 寻址一个会话的 Agent(被服务的会话必有 Agent;`command.*` 经由与 `session.*` 相同的路径恢复冷会话,而 `skill.list` 从会话头解析项目根目录,不触碰 Agent 注册表)。`command.execute` 在宿主侧运行一条斜杠命令行并返回脱耦结果;载体的请求信号可取消正在运行的处理器。`host/commands-changed` 是目录失效帧:客户端重新拉取 `command.list` 而不是做差分。
+73 -1
View File
@@ -40,6 +40,7 @@ import type {
} from '@deepseek-ai/dsh-user-interaction'
import { UserInteractionError } from '@deepseek-ai/dsh-user-interaction'
import { pickNativeDirectory } from './native-directory-picker.ts'
import { affectsSessionMetrics, SessionMetricsProjector } from './session-metrics.ts'
/** Page size when history is called without maxMessages. */
const DEFAULT_MAX_MESSAGES = 50
@@ -417,6 +418,54 @@ export function createApiProxy(ctx: Context, defaults: ApiProxyDefaults): ApiPro
for (const queue of muxQueues) queue.push(envelope)
}
const pendingMetricSessions = new Set<Session>()
let metricFlushScheduled = false
let metricsDisposed = false
const metricsProjector = new SessionMetricsProjector(
ctx,
agent => targetFor(agent).current,
(agent) => { scheduleMetrics(agent.session) },
)
/** Queue one full-log metrics publication after synchronous session listeners drain. */
function scheduleMetrics(session: Session): void {
if (metricsDisposed || muxQueues.size === 0) return
pendingMetricSessions.add(session)
if (metricFlushScheduled) return
metricFlushScheduled = true
queueMicrotask(() => {
metricFlushScheduled = false
if (metricsDisposed) {
pendingMetricSessions.clear()
return
}
const sessions = [...pendingMetricSessions]
pendingMetricSessions.clear()
for (const current of sessions) {
broadcast({
type: 'session/metrics',
sessionId: current.id,
metrics: metricsProjector.snapshot(current, ctx.agents.get(current.id)),
})
}
})
}
ctx.effect(() => {
const disposers = [
ctx.on('session/event', (session: Session, event: SessionEvent) => {
if (affectsSessionMetrics(event)) scheduleMetrics(session)
}),
ctx.on('agent/created', (agent: Agent) => { scheduleMetrics(agent.session) }),
ctx.on('session/disposed', (session: Session) => { pendingMetricSessions.delete(session) }),
]
return () => {
metricsDisposed = true
pendingMetricSessions.clear()
for (const dispose of disposers) dispose()
}
}, 'api-proxy: session metrics')
/**
* Per-session inbox mirror serving the mux-open queue snapshot (the same
* refresh-recovery baseline as pending questions). Keyed by the stable
@@ -723,7 +772,15 @@ export function createApiProxy(ctx: Context, defaults: ApiProxyDefaults): ApiPro
// log (the page window may not contain the last todo/write; a paged
// client cannot reconstruct session-level state from it).
const todos = beforeSeq === undefined ? backscanTodos(found.agent.session.events) : undefined
return ok(request, { events: entries, hasMore: page.hasMore, ...todos === undefined ? {} : { todos } })
const metrics = beforeSeq === undefined
? metricsProjector.snapshot(found.agent.session, found.agent)
: undefined
return ok(request, {
events: entries,
hasMore: page.hasMore,
...todos === undefined ? {} : { todos },
...metrics === undefined ? {} : { metrics },
})
},
async models(request) {
@@ -817,6 +874,11 @@ export function createApiProxy(ctx: Context, defaults: ApiProxyDefaults): ApiPro
: { reasoningEffort: resolved.reasoningEffort },
}
targetFor(found.agent).current = selected
broadcast({
type: 'session/metrics',
sessionId: found.agent.session.id,
metrics: metricsProjector.snapshot(found.agent.session, found.agent),
})
return ok(request, { selected: { ...selected } })
} catch (error: unknown) {
return err(request, {
@@ -1102,6 +1164,11 @@ export function createApiProxy(ctx: Context, defaults: ApiProxyDefaults): ApiPro
muxQueues.add(queue)
for (const session of ctx.sessions.list()) {
subscribeSession(queue, session)
queue.push(frame({
type: 'session/metrics',
sessionId: session.id,
metrics: metricsProjector.snapshot(session, ctx.agents.get(session.id)),
}))
}
for (const pending of pendingQuestions.values()) {
queue.push({
@@ -1154,6 +1221,11 @@ export function createApiProxy(ctx: Context, defaults: ApiProxyDefaults): ApiPro
}),
ctx.on('session/created', (session: Session) => {
subscribeSession(queue, session)
queue.push(frame({
type: 'session/metrics',
sessionId: session.id,
metrics: metricsProjector.snapshot(session, ctx.agents.get(session.id)),
}))
}),
ctx.on('session/disposed', (session: Session) => {
openCalls.delete(session.id)
@@ -10,7 +10,9 @@ import type { HostFrame, MuxFrame } from './events.ts'
import type { Wire } from './rpc.schema.ts'
import { rpcErrorSchema, rpcIdSchema } from './rpc.schema.ts'
import { approvalRequestIdSchema } from './approvals.schema.ts'
import { contentBlockSchema, sessionEventSchema, sessionIdSchema, toolEventViewSchema } from './sessions.schema.ts'
import {
contentBlockSchema, sessionEventSchema, sessionIdSchema, sessionMetricsSchema, toolEventViewSchema,
} from './sessions.schema.ts'
import { workspaceIdSchema, workspaceViewSchema } from './workspace.schema.ts'
/** Question shape validated strictly against core dsh-user-interaction. */
@@ -27,6 +29,7 @@ export const askUserQuestionItemSchema = z.object({
export const muxFrameSchema = z.discriminatedUnion('type', [
z.object({ type: z.literal('session/event'), sessionId: sessionIdSchema, event: sessionEventSchema, view: toolEventViewSchema.optional() }),
z.object({ type: z.literal('session/subscribed'), sessionId: sessionIdSchema, lastSeq: z.number().int() }),
z.object({ type: z.literal('session/metrics'), sessionId: sessionIdSchema, metrics: sessionMetricsSchema }),
z.object({ type: z.literal('session/title'), sessionId: sessionIdSchema, title: z.string().min(1), eventSeq: z.number().int().nonnegative(), updatedAt: z.number() }),
z.object({ type: z.literal('approval/requested'), sessionId: sessionIdSchema, approvalId: approvalRequestIdSchema, toolName: z.string(), callId: z.string().optional(), reason: z.string().optional() }),
z.object({ type: z.literal('approval/resolved'), sessionId: sessionIdSchema, approvalId: approvalRequestIdSchema, outcome: z.union([z.literal('allowed-once'), z.literal('rejected'), z.literal('cancelled'), z.literal('unavailable')]) }),
+2
View File
@@ -13,6 +13,7 @@ import type { CallId } from '@deepseek-ai/dsh-llm/brand'
import type { SessionEvent, SessionId } from '@deepseek-ai/dsh-session/types'
import type { ToolCallView, ToolResultView } from '@deepseek-ai/dsh-tools/presentation'
import type { RpcError, RpcId, RpcRequest } from './rpc.ts'
import type { SessionMetrics } from './sessions.ts'
import type { WorkspaceView } from './workspace.ts'
// Client-side consumers take the render-intent vocabulary from the contract;
@@ -57,6 +58,7 @@ export interface EventsApi {
export type MuxFrame =
| { type: 'session/event'; sessionId: SessionId; event: SessionEvent; view?: ToolEventView }
| { type: 'session/subscribed'; sessionId: SessionId; lastSeq: number }
| { type: 'session/metrics'; sessionId: SessionId; metrics: SessionMetrics }
| { type: 'session/title'; sessionId: SessionId; title: string; eventSeq: number; updatedAt: number }
| { type: 'approval/requested'; sessionId: SessionId; approvalId: ApprovalRequestId; toolName: string; callId?: CallId; reason?: string }
| { type: 'approval/resolved'; sessionId: SessionId; approvalId: ApprovalRequestId; outcome: ApprovalOutcome }
+1 -1
View File
@@ -27,7 +27,7 @@ export interface ApiProxy {
// ---- Domain interfaces and payload entities ----
export type {
HistoryEntry, ModelCatalogFailure, ModelCatalogModel, ModelProviderGroup, ModelReasoning,
ModelReasoningEffort, ModelTarget, SessionModels, SessionsApi, SessionSummary,
ModelReasoningEffort, ModelTarget, SessionMetrics, SessionModels, SessionsApi, SessionSummary,
} from './sessions.ts'
export type { HostApi } from './host.ts'
export type { WorkspaceApi, WorkspaceId, WorkspaceView } from './workspace.ts'
@@ -11,7 +11,7 @@ import type { RequestPayload, ResponseValue } from './rpc-map.ts'
import type { Wire } from './rpc.schema.ts'
import type {
HistoryEntry, ModelCatalogFailure, ModelCatalogModel, ModelProviderGroup, ModelReasoning,
ModelReasoningEffort, ModelTarget, SessionSummary,
ModelReasoningEffort, ModelTarget, SessionMetrics, SessionSummary,
} from './sessions.ts'
import type { ToolEventView } from './events.ts'
import type { WorkspaceId } from './workspace.ts'
@@ -145,11 +145,24 @@ export const todoItemSchema = z.object({
status: z.union([z.literal('pending'), z.literal('in_progress'), z.literal('completed')]),
})
/** Host-owned durable usage and current-context projection. */
export const sessionMetricsSchema = z.object({
logRevision: z.number().int().nonnegative(),
projectionRevision: z.number().int().nonnegative(),
uncachedInputTokens: z.number().nonnegative(),
outputTokens: z.number().nonnegative(),
cacheReadTokens: z.number().nonnegative(),
cacheWriteTokens: z.number().nonnegative(),
contextTokens: z.number().nonnegative().optional(),
contextWindow: z.number().int().positive().optional(),
}) satisfies z.ZodType<Wire<SessionMetrics>>
/** session.history response value. */
export const sessionHistoryValueSchema = z.object({
events: z.array(historyEntrySchema),
hasMore: z.boolean(),
todos: z.array(todoItemSchema).optional(),
metrics: sessionMetricsSchema.optional(),
}) satisfies z.ZodType<Wire<ResponseValue<'session.history'>>>
/** session.models request payload. */
+28 -1
View File
@@ -32,6 +32,31 @@ export interface HistoryEntry {
view?: ToolEventView
}
/**
* Host-owned token metrics for one durable session revision. Provider usage
* buckets are cumulative across the full log; current context fields describe
* the replayed request surface at this revision and are absent when the Host
* cannot measure pressure or resolve exact-route capacity.
*/
export interface SessionMetrics {
/** Number of durable events included in this projection. */
logRevision: number
/** Monotone ordering within one Host process and mux subscription generation. */
projectionRevision: number
/** Cumulative uncached provider input. */
uncachedInputTokens: number
/** Cumulative provider output. */
outputTokens: number
/** Cumulative provider cache reads. */
cacheReadTokens: number
/** Cumulative provider cache writes; excluded from the Web cache-hit formula. */
cacheWriteTokens: number
/** Current request pressure from `ctx.tokenMeter.measure(session).totalTokens`. */
contextTokens?: number
/** Exact selected-route capacity from `ctx.llm.resolveModelInfo()`. */
contextWindow?: number
}
/** Complete model target selected for one session. */
export interface ModelTarget {
/** Registered provider route. */
@@ -153,9 +178,11 @@ export interface SessionsApi {
* projection (latest `todo/write` over the FULL log, independent of the page window) —
* so a paged client restores the plan without walking history; absent when the session
* never wrote one. Older pages omit it (the projection is session-level, not per-page).
* The same tail-only rule carries `metrics`, whose cumulative usage and current context
* are Host projections over the full log rather than products of the returned page.
*/
history(request: RpcRequest<{ sessionId: SessionId; beforeSeq?: number; maxMessages?: number }>):
Promise<RpcResponse<{ events: HistoryEntry[]; hasMore: boolean; todos?: TodoItem[] }>>
Promise<RpcResponse<{ events: HistoryEntry[]; hasMore: boolean; todos?: TodoItem[]; metrics?: SessionMetrics }>>
/** Reads a fresh advisory model directory for this session. Provider lookups run independently. */
models(request: RpcRequest<{ sessionId: SessionId }>): Promise<RpcResponse<SessionModels>>
@@ -0,0 +1,194 @@
/**
* Full-log usage and current-context projection for Web clients.
*
* @module @deepseek-ai/dsh-host-apiproxy/session-metrics
*/
import type { Context } from 'cordis'
import type { Agent, AgentLlmTarget } from '@deepseek-ai/dsh-agent'
import type { TokenUsage } from '@deepseek-ai/dsh-llm'
import type { Session, SessionEvent } from '@deepseek-ai/dsh-session'
import type { SessionMetrics } from './api/sessions.ts'
interface UsageState {
logRevision: number
projectionRevision: number
uncachedInputTokens: number
outputTokens: number
cacheReadTokens: number
cacheWriteTokens: number
byStep: Map<string, TokenUsage>
}
interface CapacityState {
routeKey: string
generation: number
status: 'pending' | 'ready'
contextWindow?: number
}
interface TokenMeterLike {
measure(session: Session): { totalTokens: number }
}
interface LlmLike {
resolveModelInfo(provider: string, model: string): Promise<{
context?: { contextWindow: number }
}>
}
function usageFrom(event: SessionEvent): { turn: number; step: number; usage: TokenUsage } | undefined {
if (event.type === 'assistant/chunk' && event.data.chunk.type === 'usage') {
return { turn: event.data.turn, step: event.data.step, usage: event.data.chunk.usage }
}
if (event.type === 'assistant/message' && event.data.usage !== undefined) {
return { turn: event.data.turn, step: event.data.step, usage: event.data.usage }
}
return undefined
}
/**
* Whether an appended event can change cumulative usage or token-meter
* pressure. Text/reasoning stream deltas remain outside both projections.
* @param event - appended durable event.
* @returns true when the Host must publish a fresh metrics snapshot.
*/
export function affectsSessionMetrics(event: SessionEvent): boolean {
if (event.type === 'assistant/chunk') return event.data.chunk.type === 'usage'
if (event.type === 'request/header') return true
return 'surfaceOp' in event
}
function recordUsage(state: UsageState, turn: number, step: number, usage: TokenUsage): void {
const key = `${turn}:${step}`
const previous = state.byStep.get(key)
if (previous !== undefined) {
state.uncachedInputTokens -= previous.inputTokens
state.outputTokens -= previous.outputTokens
state.cacheReadTokens -= previous.cacheReadTokens ?? 0
state.cacheWriteTokens -= previous.cacheWriteTokens ?? 0
}
state.byStep.set(key, usage)
state.uncachedInputTokens += usage.inputTokens
state.outputTokens += usage.outputTokens
state.cacheReadTokens += usage.cacheReadTokens ?? 0
state.cacheWriteTokens += usage.cacheWriteTokens ?? 0
}
/**
* Projects durable cumulative usage and route-aware current context without
* awaiting model metadata on the session append path.
*/
export class SessionMetricsProjector {
private readonly usage = new WeakMap<Session, UsageState>()
private readonly capacities = new WeakMap<Agent, CapacityState>()
/**
* @param ctx - Host context providing optional token-meter and LLM services.
* @param targetFor - selected route owner for one attached Web agent.
* @param onCapacityResolved - schedules a fresh live projection after exact-route metadata resolves.
*/
constructor(
private readonly ctx: Context,
private readonly targetFor: (agent: Agent) => Pick<AgentLlmTarget, 'provider' | 'model'>,
private readonly onCapacityResolved: (agent: Agent) => void,
) {}
/**
* Read a fresh detached projection through the session's durable tail.
* @param session - authoritative durable log owner.
* @param agent - attached route owner, when available.
* @returns cumulative usage and any currently available pressure/capacity.
*/
snapshot(session: Session, agent?: Agent): SessionMetrics {
const state = this.syncUsage(session)
const tokenMeter = this.ctx.get('tokenMeter') as TokenMeterLike | undefined
let contextTokens: number | undefined
if (tokenMeter !== undefined) {
try {
contextTokens = tokenMeter.measure(session).totalTokens
} catch {
// A malformed or temporarily unmeasurable replay has no honest pressure value.
}
}
const contextWindow = agent === undefined ? undefined : this.capacityFor(agent)
return {
logRevision: state.logRevision,
projectionRevision: state.projectionRevision++,
uncachedInputTokens: state.uncachedInputTokens,
outputTokens: state.outputTokens,
cacheReadTokens: state.cacheReadTokens,
cacheWriteTokens: state.cacheWriteTokens,
...contextTokens === undefined ? {} : { contextTokens },
...contextWindow === undefined ? {} : { contextWindow },
}
}
private syncUsage(session: Session): UsageState {
let state = this.usage.get(session)
if (state === undefined) {
state = {
logRevision: 0,
projectionRevision: 0,
uncachedInputTokens: 0,
outputTokens: 0,
cacheReadTokens: 0,
cacheWriteTokens: 0,
byStep: new Map(),
}
this.usage.set(session, state)
}
while (state.logRevision < session.events.length) {
const event = session.events[state.logRevision]
/* v8 ignore next -- Session events are append-only and dense; logRevision is bounded by length. */
if (event === undefined) break
const usage = usageFrom(event)
if (usage !== undefined) recordUsage(state, usage.turn, usage.step, usage.usage)
state.logRevision++
}
return state
}
private capacityFor(agent: Agent): number | undefined {
const target = this.targetFor(agent)
const routeKey = `${target.provider}\u0000${target.model}`
let state = this.capacities.get(agent)
if (state === undefined || state.routeKey !== routeKey) {
state = {
routeKey,
generation: (state?.generation ?? 0) + 1,
status: 'pending',
}
this.capacities.set(agent, state)
this.resolveCapacity(agent, target, state)
}
return state.status === 'ready' ? state.contextWindow : undefined
}
private resolveCapacity(
agent: Agent,
target: Pick<AgentLlmTarget, 'provider' | 'model'>,
pending: CapacityState,
): void {
const llm = this.ctx.get('llm') as LlmLike | undefined
if (llm === undefined) {
pending.status = 'ready'
return
}
void Promise.resolve()
.then(() => llm.resolveModelInfo(target.provider, target.model))
.then(
(resolved) => {
if (this.capacities.get(agent)?.generation !== pending.generation) return
const current = this.targetFor(agent)
if (`${current.provider}\u0000${current.model}` !== pending.routeKey) return
pending.status = 'ready'
if (resolved.context !== undefined) pending.contextWindow = resolved.context.contextWindow
this.onCapacityResolved(agent)
},
() => {
if (this.capacities.get(agent)?.generation === pending.generation) pending.status = 'ready'
},
)
}
}
@@ -19,6 +19,7 @@ import SystemPrompt from '@deepseek-ai/dsh-system-prompt'
import UserInteractionService from '@deepseek-ai/dsh-user-interaction'
import type { RpcRequest } from '@deepseek-ai/dsh-host-apiproxy/api/rpc'
import { RpcId } from '@deepseek-ai/dsh-host-apiproxy/api/rpc'
import type { MuxFrame } from '@deepseek-ai/dsh-host-apiproxy/api'
import { createApiProxy } from '../src/api-proxy.ts'
let nextRpc = 1
@@ -52,6 +53,7 @@ class CatalogAdapter extends LlmAdapter {
provider,
id: model,
name: model,
context: { contextWindow: model === 'private-preview' ? 128_000 : 64_000 },
...this.reasoning === undefined ? {} : { reasoning: this.reasoning },
})
}
@@ -117,6 +119,16 @@ function expectValue<T>(response: { result: { ok: true; value: T } | { ok: false
return response.result.value
}
async function nextMetrics(
iterator: AsyncIterator<RpcRequest<MuxFrame>>,
): Promise<Extract<MuxFrame, { type: 'session/metrics' }>['metrics']> {
for (;;) {
const next = await iterator.next()
if (next.done) throw new Error('mux ended before a metrics frame')
if (next.value.payload.type === 'session/metrics') return next.value.payload.metrics
}
}
describe('Web session model selection', () => {
it('groups successful providers, isolates failures, and preserves an unlisted current model', async () => {
const { ctx, sessionId } = await harness({
@@ -230,4 +242,26 @@ describe('Web session model selection', () => {
.toEqual({ provider: 'deepseek', model: 'private-preview', reasoningEffort: 'max' })
await ctx.fiber.dispose()
})
it('publishes unknown capacity immediately on selection, then the exact selected route capacity', async () => {
const { ctx, sessionId } = await harness()
const api = createApiProxy(ctx, { provider: 'deepseek', model: 'deepseek-chat', cwd: '/tmp', workspaceRoot: '/tmp' })
const controller = new AbortController()
const iterator = api.events.mux(request({}), controller.signal)[Symbol.asyncIterator]()
expect((await nextMetrics(iterator)).contextWindow).toBeUndefined()
expect((await nextMetrics(iterator)).contextWindow).toBe(64_000)
expectValue(await api.sessions.selectModel(request({
sessionId,
provider: 'deepseek',
model: 'private-preview',
})))
expect((await nextMetrics(iterator)).contextWindow).toBeUndefined()
expect((await nextMetrics(iterator)).contextWindow).toBe(128_000)
controller.abort()
await iterator.return?.()
await ctx.fiber.dispose()
})
})
@@ -10,7 +10,7 @@ import {
sessionCreateValueSchema, sessionEventSchema, sessionHistoryRequestSchema, sessionHistoryValueSchema,
sessionIdSchema, sessionListRequestSchema, sessionListValueSchema, sessionModelsRequestSchema,
sessionModelsValueSchema, sessionPromptRequestSchema, sessionPromptValueSchema,
sessionSelectModelRequestSchema, sessionSelectModelValueSchema, sessionSummarySchema,
sessionSelectModelRequestSchema, sessionSelectModelValueSchema, sessionSummarySchema, sessionMetricsSchema,
} from '../src/api/sessions.schema.ts'
import { hostDescribeRequestSchema, hostDescribeValueSchema } from '../src/api/host.schema.ts'
import {
@@ -138,8 +138,35 @@ describe('sessions domain schemas', () => {
expect(sessionHistoryValueSchema.parse({
events: [],
hasMore: false,
metrics: {
logRevision: 12,
projectionRevision: 4,
uncachedInputTokens: 1_000,
outputTokens: 200,
cacheReadTokens: 4_000,
cacheWriteTokens: 500,
contextTokens: 8_000,
contextWindow: 128_000,
},
modelTarget: { provider: 'deepseek', model: 'deepseek-v4-flash' },
}).hasMore).toBe(false)
}).metrics?.contextWindow).toBe(128_000)
expect(() => sessionMetricsSchema.parse({
logRevision: 1,
projectionRevision: 0,
uncachedInputTokens: -1,
outputTokens: 0,
cacheReadTokens: 0,
cacheWriteTokens: 0,
})).toThrow()
expect(() => sessionMetricsSchema.parse({
logRevision: 1,
projectionRevision: 0,
uncachedInputTokens: 0,
outputTokens: 0,
cacheReadTokens: 0,
cacheWriteTokens: 0,
contextWindow: 0,
})).toThrow()
expect(sessionModelsRequestSchema.parse({ sessionId: 's1' }).sessionId).toBe('s1')
expect(sessionModelsValueSchema.parse({
current: { provider: 'deepseek', model: 'deepseek-v4-flash', reasoningEffort: 'max' },
@@ -305,6 +332,18 @@ describe('events frame schemas', () => {
{ type: 'session/event', sessionId: 's', event: { type: 't', seq: 0, time: 1, data: null } },
{ type: 'session/subscribed', sessionId: 's', lastSeq: -1 },
{ type: 'session/title', sessionId: 's', title: 'Durable title', eventSeq: 2, updatedAt: 3 },
{
type: 'session/metrics',
sessionId: 's',
metrics: {
logRevision: 3,
projectionRevision: 1,
uncachedInputTokens: 100,
outputTokens: 20,
cacheReadTokens: 300,
cacheWriteTokens: 40,
},
},
{ type: 'approval/requested', sessionId: 's', approvalId: 'a', toolName: 'bash', callId: 'c', reason: 'r' },
{ type: 'approval/resolved', sessionId: 's', approvalId: 'a', outcome: 'allowed-once' },
{ type: 'question/requested', sessionId: 's', questions: [{ id: 'q', question: 'Q?', options: [{ label: 'L' }], multiSelect: true }] },
@@ -0,0 +1,277 @@
import { describe, expect, it, vi } from 'vitest'
import { Context } from 'cordis'
import type { Agent, AgentLlmTarget } from '@deepseek-ai/dsh-agent'
import { Session, SessionId } from '@deepseek-ai/dsh-session'
import { affectsSessionMetrics, SessionMetricsProjector } from '../src/session-metrics.ts'
function assistant(
session: Session,
turn: number,
step: number,
usage: {
inputTokens: number
outputTokens: number
cacheReadTokens?: number
cacheWriteTokens?: number
},
): void {
session.append('assistant/chunk', {
turn,
step,
chunk: { type: 'usage', usage },
})
session.append('assistant/message', {
turn,
step,
content: [{ type: 'text', text: `answer-${turn}-${step}` }],
provenance: { provider: 'test', model: 'alpha' },
usage,
}, { surfaceOp: 'append' })
}
function agent(session: Session): Agent {
return { id: session.id, session } as Agent
}
describe('SessionMetricsProjector', () => {
it('filters text/reasoning stream deltas while retaining usage, headers, and surface mutations', () => {
const session = new Session(SessionId('metrics-filter'))
const text = session.append('assistant/chunk', {
turn: 1,
step: 1,
chunk: { type: 'text-delta', index: 0, text: 'x' },
})
const usage = session.append('assistant/chunk', {
turn: 1,
step: 1,
chunk: { type: 'usage', usage: { inputTokens: 1, outputTokens: 1 } },
})
const header = session.append('request/header', {
header: { config: { provider: 'test', model: 'alpha' } },
reason: 'initial',
})
const surface = session.append('user/message', {
content: [{ type: 'text', text: 'question' }],
source: { kind: 'user' },
}, { surfaceOp: 'append' })
expect(affectsSessionMetrics(text)).toBe(false)
expect(affectsSessionMetrics(usage)).toBe(true)
expect(affectsSessionMetrics(header)).toBe(true)
expect(affectsSessionMetrics(surface)).toBe(true)
})
it('reconciles usage by turn:step, keeps cache writes disjoint, and survives a surface replacement', () => {
const ctx = new Context()
ctx.provide('tokenMeter', {
measure(session: Session) {
return { totalTokens: session.surface.nodes.length * 100 }
},
})
const session = new Session(SessionId('metrics-fold'))
const first = session.append('user/message', {
content: [{ type: 'text', text: 'large old surface' }],
source: { kind: 'user' },
}, { surfaceOp: 'append' })
assistant(session, 1, 1, {
inputTokens: 11,
outputTokens: 3,
cacheReadTokens: 89,
cacheWriteTokens: 8,
})
const current: AgentLlmTarget = { provider: 'test', model: 'alpha' }
const projector = new SessionMetricsProjector(ctx, () => current, () => {})
const attached = agent(session)
const before = projector.snapshot(session, attached)
expect(before).toMatchObject({
uncachedInputTokens: 11,
outputTokens: 3,
cacheReadTokens: 89,
cacheWriteTokens: 8,
contextTokens: 200,
})
const assistantSeq = session.surface.nodes.at(-1)
if (assistantSeq === undefined) throw new Error('assistant surface missing')
session.append('user/message', {
content: [{ type: 'text', text: 'compact summary' }],
source: { kind: 'plugin', plugin: 'test' },
}, {
surfaceOp: { op: 'replace', start: first.seq, end: assistantSeq },
sourceEventSeqs: [first.seq, assistantSeq],
})
// A replayed usage event for the same step replaces the settled value.
session.append('assistant/chunk', {
turn: 1,
step: 1,
chunk: {
type: 'usage',
usage: {
inputTokens: 12,
outputTokens: 4,
cacheReadTokens: 88,
cacheWriteTokens: 9,
},
},
})
const compacted = projector.snapshot(session, attached)
expect(compacted).toMatchObject({
uncachedInputTokens: 12,
outputTokens: 4,
cacheReadTokens: 88,
cacheWriteTokens: 9,
contextTokens: 100,
})
assistant(session, 1, 2, {
inputTokens: 1_000,
outputTokens: 500,
cacheReadTokens: 2_000,
cacheWriteTokens: 3_000,
})
const after = projector.snapshot(session, attached)
expect(after).toMatchObject({
logRevision: session.events.length,
projectionRevision: 2,
uncachedInputTokens: 1_012,
outputTokens: 504,
cacheReadTokens: 2_088,
cacheWriteTokens: 3_009,
contextTokens: 200,
})
expect(after.uncachedInputTokens).not.toBe(
after.uncachedInputTokens + after.cacheReadTokens + after.cacheWriteTokens,
)
})
it('publishes only the selected route capacity when asynchronous resolutions race', async () => {
const ctx = new Context()
const resolutions = new Map<string, (contextWindow: number) => void>()
ctx.provide('tokenMeter', { measure: () => ({ totalTokens: 35_000 }) })
ctx.provide('llm', {
resolveModelInfo(_provider: string, model: string) {
return new Promise<{ context: { contextWindow: number } }>((resolve) => {
resolutions.set(model, (contextWindow) => { resolve({ context: { contextWindow } }) })
})
},
})
const session = new Session(SessionId('capacity-race'))
const attached = agent(session)
let current: AgentLlmTarget = { provider: 'test', model: 'alpha' }
const resolved = vi.fn()
const projector = new SessionMetricsProjector(ctx, () => current, resolved)
expect(projector.snapshot(session, attached).contextWindow).toBeUndefined()
await vi.waitFor(() => { expect(resolutions.has('alpha')).toBe(true) })
current = { provider: 'test', model: 'beta' }
expect(projector.snapshot(session, attached).contextWindow).toBeUndefined()
await vi.waitFor(() => { expect(resolutions.has('beta')).toBe(true) })
resolutions.get('alpha')?.(64_000)
await Promise.resolve()
expect(resolved).not.toHaveBeenCalled()
resolutions.get('beta')?.(128_000)
await vi.waitFor(() => { expect(resolved).toHaveBeenCalledOnce() })
expect(projector.snapshot(session, attached)).toMatchObject({
contextTokens: 35_000,
contextWindow: 128_000,
})
})
it('omits current context fields when measurement or model metadata is unavailable', async () => {
const ctx = new Context()
ctx.provide('tokenMeter', { measure: () => { throw new Error('unmeasurable') } })
ctx.provide('llm', { resolveModelInfo: () => Promise.reject(new Error('metadata unavailable')) })
const session = new Session(SessionId('missing-metrics'))
const attached = agent(session)
const projector = new SessionMetricsProjector(
ctx,
() => ({ provider: 'test', model: 'missing' }),
() => {},
)
const metrics = projector.snapshot(session, attached)
expect(metrics.contextTokens).toBeUndefined()
expect(metrics.contextWindow).toBeUndefined()
await vi.waitFor(() => {
expect(projector.snapshot(session, attached).contextWindow).toBeUndefined()
})
})
it('keeps optional usage buckets at zero and tolerates absent host services or detached agents', async () => {
const ctx = new Context()
const session = new Session(SessionId('optional-metrics'))
assistant(session, 1, 0, { inputTokens: 7, outputTokens: 2 })
const attached = agent(session)
const projector = new SessionMetricsProjector(
ctx,
() => ({ provider: 'test', model: 'no-service' }),
() => {},
)
expect(projector.snapshot(session)).toMatchObject({
uncachedInputTokens: 7,
outputTokens: 2,
cacheReadTokens: 0,
cacheWriteTokens: 0,
})
expect(projector.snapshot(session, attached).contextWindow).toBeUndefined()
expect(projector.snapshot(session, attached).contextWindow).toBeUndefined()
await Promise.resolve()
})
it('publishes a resolved route with no advertised capacity as unknown', async () => {
const ctx = new Context()
ctx.provide('llm', { resolveModelInfo: () => Promise.resolve({}) })
const session = new Session(SessionId('no-capacity'))
const attached = agent(session)
const resolved = vi.fn()
const projector = new SessionMetricsProjector(
ctx,
() => ({ provider: 'test', model: 'metadata-without-context' }),
resolved,
)
expect(projector.snapshot(session, attached).contextWindow).toBeUndefined()
await vi.waitFor(() => { expect(resolved).toHaveBeenCalledOnce() })
expect(projector.snapshot(session, attached).contextWindow).toBeUndefined()
})
it('ignores stale resolution failures and route metadata after the target moves', async () => {
const ctx = new Context()
const resolutions = new Map<string, {
resolve(value: { context: { contextWindow: number } }): void
reject(error: Error): void
}>()
ctx.provide('llm', {
resolveModelInfo(_provider: string, model: string) {
return new Promise<{ context: { contextWindow: number } }>((resolve, reject) => {
resolutions.set(model, { resolve, reject })
})
},
})
const session = new Session(SessionId('stale-capacity'))
const attached = agent(session)
let current: AgentLlmTarget = { provider: 'test', model: 'alpha' }
const resolved = vi.fn()
const targetFor = vi.fn(() => current)
const projector = new SessionMetricsProjector(ctx, targetFor, resolved)
projector.snapshot(session, attached)
await vi.waitFor(() => { expect(resolutions.has('alpha')).toBe(true) })
current = { provider: 'test', model: 'route-moved-before-snapshot' }
resolutions.get('alpha')?.resolve({ context: { contextWindow: 64_000 } })
await vi.waitFor(() => { expect(targetFor).toHaveBeenCalledTimes(2) })
expect(resolved).not.toHaveBeenCalled()
projector.snapshot(session, attached)
await vi.waitFor(() => { expect(resolutions.has('route-moved-before-snapshot')).toBe(true) })
current = { provider: 'test', model: 'beta' }
projector.snapshot(session, attached)
await vi.waitFor(() => { expect(resolutions.has('beta')).toBe(true) })
resolutions.get('route-moved-before-snapshot')?.reject(new Error('stale failure'))
await Promise.resolve()
resolutions.get('beta')?.resolve({ context: { contextWindow: 128_000 } })
await vi.waitFor(() => { expect(resolved).toHaveBeenCalledOnce() })
expect(projector.snapshot(session, attached).contextWindow).toBe(128_000)
})
})