fix(llm-pi-ai): let a model declare the request modalities it accepts
A model the installed pi-ai catalog does not describe was reported as text-only with no way to say otherwise, so a vision model added through the custom-provider form was refused at every image admission point. The justification in the source described the DeepSeek chat-completions serializer, which does reject image blocks; the pi-ai request converter and every wire protocol it speaks carry images. Modalities now resolve entry `input` -> installed catalog entry -> route `defaultInput`, the chain the two capacity fallbacks already use, so the route value is a fallback and never narrows a catalog model. Its default stays `[text]`: nothing can interrogate a gateway for its modalities, and over-claiming admits an image the provider rejects mid-turn, after prompt admission has already committed the message.
This commit is contained in:
@@ -2,5 +2,5 @@
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write docs/config-catalog.md
|
||||
config-catalog.md: b2a5a0469bad63297d0c3509f8d76f387beb58cb
|
||||
config-catalog.zh.md: 5bf93b4b3dcbcfc4e55ee4ed5fb1c77d17c86ab2
|
||||
config-catalog.md: 5442e18d57de9991f53283e6623c2dbe4788b32b
|
||||
config-catalog.zh.md: 7ba2449a00494963cb014e1be95b60e143efe4b4
|
||||
+28
-2
@@ -858,6 +858,17 @@ export interface PiAiProviderProfile {
|
||||
* never becomes a per-request cap on its own.
|
||||
*/
|
||||
defaultMaxTokens?: number
|
||||
/**
|
||||
* Request modalities for a model this route lists that neither its entry's
|
||||
* {@link PiAiModelProfile.input} nor the installed catalog declares (default
|
||||
* `[text]`). A fallback like the capacities above, not an override: a
|
||||
* catalog model keeps the modalities the catalog records for it, and this
|
||||
* value never narrows one. A gateway serving vision models the catalog does
|
||||
* not describe declares `[text, image]` once here instead of on every entry.
|
||||
* Unlike an entry's list, this one may not be empty — nothing sits below it
|
||||
* to answer instead.
|
||||
*/
|
||||
defaultInput?: PiAiModality[]
|
||||
/** Provider request headers; Harness attribution wins reserved names. */
|
||||
headers?: Record<string, string>
|
||||
/** Provider-neutral pi-ai reasoning level. */
|
||||
@@ -893,6 +904,18 @@ export interface PiAiModelProfile {
|
||||
* default on its own.
|
||||
*/
|
||||
maxTokens?: number
|
||||
/**
|
||||
* Request modalities this model accepts. Absent — or empty, which describes
|
||||
* a model that accepts nothing and so states no answer either — keeps the
|
||||
* installed catalog entry's modalities, then the route's `defaultInput`.
|
||||
* Declaring images is what makes a hand-declared vision model usable, and
|
||||
* declaring text alone corrects a catalog model whose gateway does not serve
|
||||
* what the catalog records. This is a claim about the endpoint, not a check
|
||||
* of it: nothing interrogates a gateway for what it accepts, so a model
|
||||
* claiming images its endpoint refuses is refused by the provider instead,
|
||||
* mid-turn.
|
||||
*/
|
||||
input?: PiAiModality[]
|
||||
/**
|
||||
* Selectable reasoning efforts. Absent inherits the installed catalog
|
||||
* entry's capability (a hand-declared model has none and does not reason);
|
||||
@@ -930,6 +953,9 @@ export interface PiAiCompatProfile {
|
||||
supportsReasoningEffort?: boolean
|
||||
}
|
||||
|
||||
/** One request modality a pi-ai model may accept. */
|
||||
export type PiAiModality = Model<Api>['input'][number]
|
||||
|
||||
/**
|
||||
* Selectable reasoning efforts for one model: each key is a level the model
|
||||
* offers (and selectors show), and its value is the wire spelling dispatch
|
||||
@@ -953,9 +979,9 @@ type PiThinkingFormat = NonNullable<OpenAICompletionsCompat['thinkingFormat']>
|
||||
type WithheldThinkingFormat = 'chat-template' | 'qwen-chat-template'
|
||||
```
|
||||
|
||||
Depends on: `CacheRetention` (`@earendil-works/pi-ai`) · `ModelThinkingLevel` (`@earendil-works/pi-ai`) · `OpenAICompletionsCompat` (`@earendil-works/pi-ai`) · [`RetryPolicyConfig`](../packages/llm/llm/src/index.ts) · `ThinkingBudgets` (`@earendil-works/pi-ai`) · `Transport` (`@earendil-works/pi-ai`)
|
||||
Depends on: `Api` (`@earendil-works/pi-ai`) · `CacheRetention` (`@earendil-works/pi-ai`) · `Model` (`@earendil-works/pi-ai`) · `ModelThinkingLevel` (`@earendil-works/pi-ai`) · `OpenAICompletionsCompat` (`@earendil-works/pi-ai`) · [`RetryPolicyConfig`](../packages/llm/llm/src/index.ts) · `ThinkingBudgets` (`@earendil-works/pi-ai`) · `Transport` (`@earendil-works/pi-ai`)
|
||||
|
||||
Source: [`packages/llm/llm-pi-ai/src/config.ts:142`](../packages/llm/llm-pi-ai/src/config.ts)
|
||||
Source: [`packages/llm/llm-pi-ai/src/config.ts:172`](../packages/llm/llm-pi-ai/src/config.ts)
|
||||
|
||||
## `@deepseek-ai/dsh-llm-replay`
|
||||
|
||||
|
||||
@@ -860,6 +860,17 @@ export interface PiAiProviderProfile {
|
||||
* never becomes a per-request cap on its own.
|
||||
*/
|
||||
defaultMaxTokens?: number
|
||||
/**
|
||||
* Request modalities for a model this route lists that neither its entry's
|
||||
* {@link PiAiModelProfile.input} nor the installed catalog declares (default
|
||||
* `[text]`). A fallback like the capacities above, not an override: a
|
||||
* catalog model keeps the modalities the catalog records for it, and this
|
||||
* value never narrows one. A gateway serving vision models the catalog does
|
||||
* not describe declares `[text, image]` once here instead of on every entry.
|
||||
* Unlike an entry's list, this one may not be empty — nothing sits below it
|
||||
* to answer instead.
|
||||
*/
|
||||
defaultInput?: PiAiModality[]
|
||||
/** Provider request headers; Harness attribution wins reserved names. */
|
||||
headers?: Record<string, string>
|
||||
/** Provider-neutral pi-ai reasoning level. */
|
||||
@@ -895,6 +906,18 @@ export interface PiAiModelProfile {
|
||||
* default on its own.
|
||||
*/
|
||||
maxTokens?: number
|
||||
/**
|
||||
* Request modalities this model accepts. Absent — or empty, which describes
|
||||
* a model that accepts nothing and so states no answer either — keeps the
|
||||
* installed catalog entry's modalities, then the route's `defaultInput`.
|
||||
* Declaring images is what makes a hand-declared vision model usable, and
|
||||
* declaring text alone corrects a catalog model whose gateway does not serve
|
||||
* what the catalog records. This is a claim about the endpoint, not a check
|
||||
* of it: nothing interrogates a gateway for what it accepts, so a model
|
||||
* claiming images its endpoint refuses is refused by the provider instead,
|
||||
* mid-turn.
|
||||
*/
|
||||
input?: PiAiModality[]
|
||||
/**
|
||||
* Selectable reasoning efforts. Absent inherits the installed catalog
|
||||
* entry's capability (a hand-declared model has none and does not reason);
|
||||
@@ -932,6 +955,9 @@ export interface PiAiCompatProfile {
|
||||
supportsReasoningEffort?: boolean
|
||||
}
|
||||
|
||||
/** One request modality a pi-ai model may accept. */
|
||||
export type PiAiModality = Model<Api>['input'][number]
|
||||
|
||||
/**
|
||||
* Selectable reasoning efforts for one model: each key is a level the model
|
||||
* offers (and selectors show), and its value is the wire spelling dispatch
|
||||
@@ -955,9 +981,9 @@ type PiThinkingFormat = NonNullable<OpenAICompletionsCompat['thinkingFormat']>
|
||||
type WithheldThinkingFormat = 'chat-template' | 'qwen-chat-template'
|
||||
```
|
||||
|
||||
依赖:`CacheRetention`(`@earendil-works/pi-ai`)· `ModelThinkingLevel`(`@earendil-works/pi-ai`)· `OpenAICompletionsCompat`(`@earendil-works/pi-ai`)· [`RetryPolicyConfig`](../packages/llm/llm/src/index.ts) · `ThinkingBudgets`(`@earendil-works/pi-ai`)· `Transport`(`@earendil-works/pi-ai`)
|
||||
依赖:`Api`(`@earendil-works/pi-ai`)· `CacheRetention`(`@earendil-works/pi-ai`)· `Model`(`@earendil-works/pi-ai`)· `ModelThinkingLevel`(`@earendil-works/pi-ai`)· `OpenAICompletionsCompat`(`@earendil-works/pi-ai`)· [`RetryPolicyConfig`](../packages/llm/llm/src/index.ts) · `ThinkingBudgets`(`@earendil-works/pi-ai`)· `Transport`(`@earendil-works/pi-ai`)
|
||||
|
||||
来源:[`packages/llm/llm-pi-ai/src/config.ts:142`](../packages/llm/llm-pi-ai/src/config.ts)
|
||||
来源:[`packages/llm/llm-pi-ai/src/config.ts:172`](../packages/llm/llm-pi-ai/src/config.ts)
|
||||
|
||||
## `@deepseek-ai/dsh-llm-replay`
|
||||
|
||||
|
||||
@@ -2,5 +2,5 @@
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write docs/user/guide/providers.md
|
||||
providers.md: a3f94f0cc86401c0f9e5b94cfd823bf9f08e6bfc
|
||||
providers.zh.md: 7d74e0086e62d8a0c2fb39085207125b4b4354e7
|
||||
providers.md: 099f434ec4602aa402239e83c708d81fcadd7732
|
||||
providers.zh.md: 367c90b525ad628b3cd86b2d22045c25064e88a1
|
||||
@@ -28,6 +28,57 @@ The Provider ID is permanent because requests, saved sessions, model defaults, a
|
||||
|
||||
Under **Model catalog**, choose **Fetch available models** to query the base URL and credential currently shown in the form. Selecting candidates updates the draft; the provider is not stored until you save. Catalog providers use their installed catalog without a network request.
|
||||
|
||||
### Image input
|
||||
|
||||
A model you enter by hand is treated as text-only until it says otherwise, because nothing can ask an endpoint which modalities it accepts. Attaching an image to such a model is refused before it is sent, naming the model.
|
||||
|
||||
A vision model on a custom provider therefore needs one line. The form has no field for it; add `input` to the model in `$DSH_HOME/settings.yaml`:
|
||||
|
||||
```yaml
|
||||
llm-pi-ai:
|
||||
providers:
|
||||
my-gateway:
|
||||
apiKeyEnv: GATEWAY_API_KEY
|
||||
api: openai-completions
|
||||
baseURL: https://gateway.example/v1
|
||||
models:
|
||||
- id: legacy-chat
|
||||
- id: vision-preview
|
||||
input: [text, image]
|
||||
```
|
||||
|
||||
`input` accepts `text` and `image`, and applies to that model alone, so one route can serve both kinds. Omitting it — or writing an empty list, which means the same thing — keeps whatever the installed catalog records for that model, and falls back to the route's `defaultInput` for a model the catalog does not describe.
|
||||
|
||||
If every model you entered by hand takes images, set the fallback once on the route instead of on each of them:
|
||||
|
||||
```yaml
|
||||
llm-pi-ai:
|
||||
providers:
|
||||
vision-gateway:
|
||||
apiKeyEnv: GATEWAY_API_KEY
|
||||
api: openai-completions
|
||||
baseURL: https://vision.example/v1
|
||||
defaultInput: [text, image]
|
||||
models:
|
||||
- id: first-model
|
||||
- id: second-model
|
||||
```
|
||||
|
||||
`defaultInput` is a fallback, not an override, and defaults to `[text]`: on a catalog provider it answers only for models the catalog does not describe, so it never removes images from a catalog model that has them. Narrow one of those with that model's own `input`. A catalog provider has no `models` list to put it in, so write it under `modelOverrides`, keyed by model id:
|
||||
|
||||
```yaml
|
||||
llm-pi-ai:
|
||||
providers:
|
||||
anthropic:
|
||||
modelOverrides:
|
||||
claude-sonnet-4-5:
|
||||
input: [text]
|
||||
```
|
||||
|
||||
Every list must name at least one modality except a model's own, where an empty list means the same as omitting it. An unknown modality is refused wherever it is written.
|
||||
|
||||
Both fields state a claim about your endpoint rather than checking it. A model that declares images its endpoint does not serve is not caught here; the provider rejects the request instead.
|
||||
|
||||
## Select a model
|
||||
|
||||
Configured providers appear in the model picker. Selecting a model also makes it the default for new sessions. A session that has already sent a request retains the model recorded in its own log.
|
||||
@@ -39,6 +90,8 @@ If a saved default names a provider that was deleted, the composer displays **Se
|
||||
- **`MISSING_CREDENTIAL`** — Store the provider key through the Models page or supply the referenced environment variable.
|
||||
- **`UNKNOWN_MODEL`** — Select a configured model or add the missing model to the custom provider.
|
||||
- **Fetching available models returns 401** — Check the key. Model discovery calls the OpenAI-compatible `GET /models` endpoint; enter models manually for endpoints that do not provide it.
|
||||
- **An image is refused before sending** — The model declares no image modality. Give a custom provider's model `input: [text, image]`; DeepSeek's own chat-completions route is text-only and cannot be configured otherwise.
|
||||
- **The provider rejects a request carrying an image** — The model declares images its endpoint does not actually serve. Remove `image` from whichever list granted it — the model's `input`, or the route's `defaultInput` — then start a new session: the attached image stays in the session log, so the same request repeats until the session moves off it.
|
||||
|
||||
## Advanced configuration
|
||||
|
||||
|
||||
@@ -28,6 +28,57 @@ Provider ID 是永久的,因为请求、已保存会话、模型默认值和
|
||||
|
||||
在**模型目录**中选择**获取可用模型**,可查询表单当前显示的基础 URL 和凭据。选择候选项只会更新草稿;保存前不会存储提供方。目录提供方使用已安装目录,不发起网络请求。
|
||||
|
||||
### 图片输入
|
||||
|
||||
手动输入的模型在自己声明之前一律按纯文本对待,因为没有任何环节能去询问端点接受哪些模态。给这类模型附加图片,会在发送前就被拒绝,并点名该模型。
|
||||
|
||||
因此自定义提供方下的视觉模型需要加一行。表单没有对应字段;请在 `$DSH_HOME/settings.yaml` 中给该模型加上 `input`:
|
||||
|
||||
```yaml
|
||||
llm-pi-ai:
|
||||
providers:
|
||||
my-gateway:
|
||||
apiKeyEnv: GATEWAY_API_KEY
|
||||
api: openai-completions
|
||||
baseURL: https://gateway.example/v1
|
||||
models:
|
||||
- id: legacy-chat
|
||||
- id: vision-preview
|
||||
input: [text, image]
|
||||
```
|
||||
|
||||
`input` 接受 `text` 和 `image`,且只作用于该模型,因此一条路由可以同时服务两类模型。省略它——或写成空列表,两者同义——则保留已安装目录为该模型记录的模态;目录未描述的模型则回退到该路由的 `defaultInput`。
|
||||
|
||||
如果你手动录入的模型全都接受图片,可以在路由上设置一次回退值,不必逐个模型写:
|
||||
|
||||
```yaml
|
||||
llm-pi-ai:
|
||||
providers:
|
||||
vision-gateway:
|
||||
apiKeyEnv: GATEWAY_API_KEY
|
||||
api: openai-completions
|
||||
baseURL: https://vision.example/v1
|
||||
defaultInput: [text, image]
|
||||
models:
|
||||
- id: first-model
|
||||
- id: second-model
|
||||
```
|
||||
|
||||
`defaultInput` 是回退值而不是覆盖值,默认为 `[text]`:在目录提供方上,它只为目录未描述的模型作答,因此绝不会把目录中本就具备图片能力的模型的该能力去掉。要收窄这类模型,请用它自己的 `input`。目录提供方没有可供填写的 `models` 列表,因此写在 `modelOverrides` 下,以模型 id 为键:
|
||||
|
||||
```yaml
|
||||
llm-pi-ai:
|
||||
providers:
|
||||
anthropic:
|
||||
modelOverrides:
|
||||
claude-sonnet-4-5:
|
||||
input: [text]
|
||||
```
|
||||
|
||||
除模型自身的列表外,每个列表都至少要写一项模态;模型自身的空列表与省略它同义。未知模态在任何位置写入都会被拒绝。
|
||||
|
||||
这两个字段都是对你端点的断言,而不是对它的检查。声明了端点并不提供的图片能力的模型不会在这里被拦下,改由提供方拒绝该请求。
|
||||
|
||||
## 选择模型
|
||||
|
||||
已配置的提供方会出现在模型选择器中。选择模型也会将其设为新会话的默认值。已发送过请求的会话会保留自身日志中记录的模型。
|
||||
@@ -39,6 +90,8 @@ Provider ID 是永久的,因为请求、已保存会话、模型默认值和
|
||||
- **`MISSING_CREDENTIAL`**:通过模型页存储提供方密钥,或提供被引用的环境变量。
|
||||
- **`UNKNOWN_MODEL`**:选择已配置的模型,或向自定义提供方添加缺失的模型。
|
||||
- **获取可用模型返回 401**:检查密钥。模型发现会调用 OpenAI 兼容的 `GET /models` 端点;对于不提供该端点的服务,请手动输入模型。
|
||||
- **图片在发送前被拒绝**:该模型未声明图片模态。请给自定义提供方的模型加上 `input: [text, image]`;DeepSeek 自身的 chat-completions 路由是纯文本的,且无法通过配置改变。
|
||||
- **提供方拒绝了带图片的请求**:该模型声明了其端点实际并不提供的图片能力。请从授予它图片能力的那个列表中移除 `image`——可能是模型的 `input`,也可能是路由的 `defaultInput`——然后开启新会话:附加的图片会留在会话日志里,因此在会话离开它之前,同一个请求会不断重复。
|
||||
|
||||
## 进阶配置
|
||||
|
||||
|
||||
Reference in New Issue
Block a user