fix(llm-pi-ai): let a model declare the request modalities it accepts
A model the installed pi-ai catalog does not describe was reported as text-only with no way to say otherwise, so a vision model added through the custom-provider form was refused at every image admission point. The justification in the source described the DeepSeek chat-completions serializer, which does reject image blocks; the pi-ai request converter and every wire protocol it speaks carry images. Modalities now resolve entry `input` -> installed catalog entry -> route `defaultInput`, the chain the two capacity fallbacks already use, so the route value is a fallback and never narrows a catalog model. Its default stays `[text]`: nothing can interrogate a gateway for its modalities, and over-claiming admits an image the provider rejects mid-turn, after prompt admission has already committed the message.
This commit is contained in:
@@ -28,6 +28,57 @@ The Provider ID is permanent because requests, saved sessions, model defaults, a
|
||||
|
||||
Under **Model catalog**, choose **Fetch available models** to query the base URL and credential currently shown in the form. Selecting candidates updates the draft; the provider is not stored until you save. Catalog providers use their installed catalog without a network request.
|
||||
|
||||
### Image input
|
||||
|
||||
A model you enter by hand is treated as text-only until it says otherwise, because nothing can ask an endpoint which modalities it accepts. Attaching an image to such a model is refused before it is sent, naming the model.
|
||||
|
||||
A vision model on a custom provider therefore needs one line. The form has no field for it; add `input` to the model in `$DSH_HOME/settings.yaml`:
|
||||
|
||||
```yaml
|
||||
llm-pi-ai:
|
||||
providers:
|
||||
my-gateway:
|
||||
apiKeyEnv: GATEWAY_API_KEY
|
||||
api: openai-completions
|
||||
baseURL: https://gateway.example/v1
|
||||
models:
|
||||
- id: legacy-chat
|
||||
- id: vision-preview
|
||||
input: [text, image]
|
||||
```
|
||||
|
||||
`input` accepts `text` and `image`, and applies to that model alone, so one route can serve both kinds. Omitting it — or writing an empty list, which means the same thing — keeps whatever the installed catalog records for that model, and falls back to the route's `defaultInput` for a model the catalog does not describe.
|
||||
|
||||
If every model you entered by hand takes images, set the fallback once on the route instead of on each of them:
|
||||
|
||||
```yaml
|
||||
llm-pi-ai:
|
||||
providers:
|
||||
vision-gateway:
|
||||
apiKeyEnv: GATEWAY_API_KEY
|
||||
api: openai-completions
|
||||
baseURL: https://vision.example/v1
|
||||
defaultInput: [text, image]
|
||||
models:
|
||||
- id: first-model
|
||||
- id: second-model
|
||||
```
|
||||
|
||||
`defaultInput` is a fallback, not an override, and defaults to `[text]`: on a catalog provider it answers only for models the catalog does not describe, so it never removes images from a catalog model that has them. Narrow one of those with that model's own `input`. A catalog provider has no `models` list to put it in, so write it under `modelOverrides`, keyed by model id:
|
||||
|
||||
```yaml
|
||||
llm-pi-ai:
|
||||
providers:
|
||||
anthropic:
|
||||
modelOverrides:
|
||||
claude-sonnet-4-5:
|
||||
input: [text]
|
||||
```
|
||||
|
||||
Every list must name at least one modality except a model's own, where an empty list means the same as omitting it. An unknown modality is refused wherever it is written.
|
||||
|
||||
Both fields state a claim about your endpoint rather than checking it. A model that declares images its endpoint does not serve is not caught here; the provider rejects the request instead.
|
||||
|
||||
## Select a model
|
||||
|
||||
Configured providers appear in the model picker. Selecting a model also makes it the default for new sessions. A session that has already sent a request retains the model recorded in its own log.
|
||||
@@ -39,6 +90,8 @@ If a saved default names a provider that was deleted, the composer displays **Se
|
||||
- **`MISSING_CREDENTIAL`** — Store the provider key through the Models page or supply the referenced environment variable.
|
||||
- **`UNKNOWN_MODEL`** — Select a configured model or add the missing model to the custom provider.
|
||||
- **Fetching available models returns 401** — Check the key. Model discovery calls the OpenAI-compatible `GET /models` endpoint; enter models manually for endpoints that do not provide it.
|
||||
- **An image is refused before sending** — The model declares no image modality. Give a custom provider's model `input: [text, image]`; DeepSeek's own chat-completions route is text-only and cannot be configured otherwise.
|
||||
- **The provider rejects a request carrying an image** — The model declares images its endpoint does not actually serve. Remove `image` from whichever list granted it — the model's `input`, or the route's `defaultInput` — then start a new session: the attached image stays in the session log, so the same request repeats until the session moves off it.
|
||||
|
||||
## Advanced configuration
|
||||
|
||||
|
||||
Reference in New Issue
Block a user