DSH plugin: custom relay reasoning 400 on every tool call

TroubleshootingPublished 2026-10-03Author: DeepSeek Plugin Market
DeepSeek HarnessDSHcustom relay400reasoning_contentopenai-completionscompattool calling
On a custom OpenAI-compatible relay, a thinking DeepSeek V4 returns 400 INVALID_REQUEST on every tool call — a missing compat key or the reasoning replay key.

Same config, same relay: gpt-5.6-terra calls tools without a problem, but switch to deepseek-v4-flash / deepseek-v4.1-flash and any tool-calling turn returns 400 INVALID_REQUEST — with a body that says only {"message":"Upstream error: 400","type":"invalid_request_error"} and names no field. There are actually two independent causes stacked on one symptom. First, the suppressReasoningContentReplay you wrote into compat does not exist anywhere in pi-ai 0.85.1, and an unknown compat key is not ignored but rejected at dispatch with INVALID_CONFIG while being named — meaning the other four switches in that compat block never took effect (#7134). Second, the cause that is genuinely unique to reasoning models lives in the "replay shape": whether thinking is replayed as reasoning_content or as structured reasoning_details depends on the thinking block's thinkingSignature, and a relay that requires non-empty reasoning_content on tool-calling turns will reject every turn in the structured shape (#7134 reply). The fixes are concrete: delete the nonexistent key and retest, confirm the replay shape from one session log field, then decide whether to add reasoningEfforts, rename the route, or ask the relay to relax its validation.

Triage first: is this 400 "config did not take effect" or "the config itself is illegal"

Before touching YAML, make one minimal determination: did the request on that route even leave DSH? The two paths look very similar (both are in the 400 family) but the fixes are completely different.

CriterionA: illegal compat keyB: replay shape does not match the relay's requirement
Error codeINVALID_CONFIG (produced on the DSH side)INVALID_REQUEST (from the relay upstream)
Does it reach the relayNo, rejected before dispatchYes, the relay returns 400
Blast radiusEvery model call on that routeOnly turns with thinking + tool calls
Are non-reasoning models affectedYes (model-independent)No (there is no thinking block)
Error text characteristicNames the key and lists the configurable switchesOnly Upstream error: 400, no field name

Step one: confirm whether the compat keys you wrote actually exist. The cheapest way is to look for whether the error names it — DSH writes undeclared keys straight into the message:

text
llm-pi-ai: provider "<route>" route sets compat "suppressReasoningContentReplay", which no wire protocol declares;
the configurable switches are supportsStore, supportsDeveloperRole, supportsReasoningEffort, supportsUsageInStreaming,
supportsFinishReason, maxTokensField, requiresToolResultName, requiresAssistantAfterToolResult, requiresThinkingAsText,
requiresReasoningContentOnAssistantMessages, thinkingFormat, chatTemplateKwargs, chatTemplateArgs,
supportsThinkingTokenBudget, thinkingTokenBudgetField, vllmPriority, supportsStrictMode, cacheControlFormat,
supportsLongCacheRetention, supportsMaxOutputTokens, supportsEagerToolInputStreaming, supportsCacheControlOnTools,
supportsTemperature, forceAdaptiveThinking, allowEmptySignature, supportsStrictTools
code: INVALID_CONFIG

Step two: confirm which level that line actually sits at. This decides whether all your earlier conclusions have to be thrown out:

  • Validation runs at both route level and model level (catalog.ts:879, catalog.ts:883);
  • compat is legal on a route (config.ts:329) and legal on each model (config.ts:311 via modelFields), and model overrides route field by field (catalog.ts:342-343);
  • The section name in settings.yaml is llm-pi-ai (packages/llm/llm-pi-ai/src/index.ts:5-7), and providers.<your route name> lives underneath it.
yaml
llm-pi-ai:
  providers:
    <your route name>:
      api: openai-completions
      compat:                       # ← route level
        supportsDeveloperRole: false
        # ...
      models:
        - id: deepseek-v4-flash
          compat:                   # ← model level, overrides route field by field
            thinkingFormat: deepseek

Step three: run a single-variable control. Comment out the whole compat block, keep only api: openai-completions and the model list, and repeat "Chinese instruction → thinking → call Pwsh". If the error now changes from INVALID_CONFIG to a real INVALID_REQUEST (or simply passes), the class-A problem is confirmed; if the behaviour is completely unchanged, that line never entered llm-pi-ai's compat at all and was swallowed by some other schema, and your earlier conclusions have to start over.

Cause one: an unknown compat key is not ignored, it is flagged at dispatch

The key takeaway of this section: DSH lets you write an unknown compat key and then rejects it at dispatch — so "the config loads" does not mean "the config is in effect", let alone "the other config is in effect too".

1. The key does not exist in the repository

All three checks for suppressReasoningContentReplay come up empty:

  • Not in the schema: packages/llm/llm-pi-ai/src/config.ts:254-281;
  • Not in COMPLETIONS_COMPAT_GATE: packages/llm/llm-pi-ai/src/catalog.ts:231-258;
  • Not findable in the installed pi-ai 0.85.1 either.

Compare the other four keys you listed — they all exist and are configurable for openai-completions:

keyExistsEvidence
supportsDeveloperRoleYesconfig.ts:256, catalog.ts:233 (offer)
requiresReasoningContentOnAssistantMessagesYesconfig.ts:264, catalog.ts:241 (offer)
thinkingFormatYes, and deepseek is a legal valueconfig.ts:265, catalog.ts:242; allowed values in catalog.ts:99-114, deepseek at catalog.ts:101
supportsStrictModeYesconfig.ts:272, catalog.ts:248 (offer)
suppressReasoningContentReplayNoZero hits across the whole repository

2. Validation happens at dispatch, and it names the key

assertOfferedCompatFields in catalog.ts:537-557 flags any key "no protocol has declared", and the message also lists the switches that really are configurable; keys with empty values are rejected too (catalog.ts:566-569). So the correct mental model is:

text
write the config → accepted (no unknown-key validation)
   ↓
first dispatch → assertOfferedCompatFields → unknown key → INVALID_CONFIG
   ↓
every model call on that route fails before it reaches the relay

The corollary is hard: as long as that line really sits under compat:, not one of your subsequent tests of those four switches means anything — the request never left.

3. Why this is more than "a missing feature"

What surfaces here is a common misreading: taking "the config loads" for "the config is validated". The flow above is the exact opposite — the more permissively you accept, the later you fail. The recommended fix (also present in the report) is to move unknown-compat-key validation forward to write time. A switch that cannot be applied should not be silently accepted and then end some later request with INVALID_CONFIG.

Cause two: requiresReasoningContentOnAssistantMessages goes silently inactive on custom routes

The key takeaway of this section: the switch is double-gated by model.reasoning, and on a custom route model.reasoning defaults to false.

In pi-ai 0.85.1 the condition is:

js
if (compat.requiresReasoningContentOnAssistantMessages && model.reasoning && assistantMsg.reasoning_content === undefined)
    assistantMsg.reasoning_content = ""

All three conditions must hold. And dsh resolves models under a custom provider key as base?.reasoning ?? false — the installed catalog has no description of your route, so model.reasoning is simply false, unless your model entry explicitly declares reasoningEfforts. Two shapes measured:

compat switchreasoningEfforts declaredField in the outbound request
trueNoNo reasoning_content key at all
trueYes"reasoning_content": ""

This reproduces your observation that "the compat setting has no effect". There is also a more fundamental mismatch: the switch's semantics are only to pad an empty string (:1045-1049 literally writes ""), not to restore the real thinking content. So even if it did take effect on a custom route, it would not fix a relay demanding non-empty real thinking.

While we are here, correct one shape: which reasoning field replay writes depends on whether the thinking block's thinkingSignature is reasoning / reasoning_content / reasoning_text (pi-ai 0.85.1, dist/api/openai-completions.js:155 and :1001-1007), and the whole block is skipped when preservedReasoningDetails is present (:999). The tool_calls mapping is at :1019-1041, which is positioned after the reasoning write and does not participate in that decision. So "DSH dropped reasoning_content because there are tool_calls" is not what that code does.

Measured: does the thinking content actually go out

The conclusion first: the real thinking text does go over the wire; what changes is which key it hangs under.

Using a local OpenAI-compatible relay plus a route declaration in exactly your shape (api: openai-completions, its own models), and running two replay rounds through the real @deepseek-ai/dsh-llm-pi-ai (0.1.5-rc.2, pi-ai 0.85.1) — round 1 the relay returns thinking + a tool call, round 2 is the round you say returns 400. The request bodies recorded:

jsonc
// when the relay reports thinking as reasoning_content, round 2 sends:
{"role":"assistant","content":null,"reasoning_content":"I should list the files.","tool_calls":[…]}
jsonc
// when the relay reports structured reasoning_details, round 2 sends:
{ …,"reasoning_details":[{"type":"reasoning.text","index":0,"text":"I should list the files."}] }

The second shape has no reasoning_content — because to pi-ai details are replay metadata, and it "replays in the shape it received". So the two shapes look completely different to the relay:

  • Relay only wants the plain reasoning field → both shapes pass (the first carries the field directly);
  • Relay requires non-empty reasoning_content → the second shape always fails.

This explains why non-reasoning models are fine on the same relay (no thinking block at all, so this validation never triggers) while the DeepSeek V4 family (with thinking) always fails.

One field settles half of it: look at thinkingSignature

You do not need a proxy to halve the classification — the session log of the failing session already holds the answer.

Open the first assistant message of the failing session and read:

text
source.replayState.blocks[0].thinkingSignature

Read it by value:

ValueMeaningNext step
Starts with [{"type":"reasoningStructured-details shape; replay will not carry reasoning_contentGo for "ask the relay to relax" or "change the route shape"
Exactly "reasoning_content"Plain-field shape; the field was sentThe 400 is elsewhere — go look at the failing response body and the full request body

If it is the second shape, the report also records several differences from a catalog route worth checking at the same time: the request carries store, uses max_completion_tokens instead of max_tokens, and has no thinking object when thinkingFormat is unset. Any of these can become a rejection reason for some relays.

Fixes: four paths, from safest to most thorough

Fix one (mandatory): delete the nonexistent compat key

diff
   compat:
     supportsDeveloperRole: false
     requiresReasoningContentOnAssistantMessages: true
     thinkingFormat: deepseek
     supportsStrictMode: true
-    suppressReasoningContentReplay: false     # ← does not exist anywhere; makes every call on this route INVALID_CONFIG

Delete it and retest. Without this step, every later experiment is invalid — the other four switches have never actually executed.

Fix two: declare reasoningEfforts on the model so the switch has a chance to fire

yaml
llm-pi-ai:
  providers:
    <your route name>:
      api: openai-completions
      compat:
        supportsDeveloperRole: false
        requiresReasoningContentOnAssistantMessages: true
        thinkingFormat: deepseek
        supportsStrictMode: true
      models:
        - id: deepseek-v4-flash
          reasoningEfforts: ["low", "medium", "high"]   # ← makes model.reasoning resolve to true

But be clear about its ceiling: once it works, all it produces is "reasoning_content": "" (empty string). If your relay wants non-empty real thinking, this path does not lead anywhere — it fixes "the field is missing", not "the field is empty".

Fix three: rename to a catalog route instead of hand-writing one

Per the repository's own notes, without explicit config pi-ai infers these compatibility switches from the provider id and baseURL; and:

A private gateway URL says nothing: for an unrecognised endpoint the inference amounts to it being OpenAI itself, which is wrong for most OpenAI-compatible gateways. (catalog.ts:345-350)

A hand-written route lands exactly in that pit; a catalog route carries the switches that vendor needs (the same trade-off is written at catalog.ts:218-221 and :553-554). And baseURL can be overridden:

ts
// catalog.ts:893
request.baseURL ?? base?.baseUrl ?? providerBaseUrl

So a workable shape is — name the route deepseek directly, then point baseURL at your relay:

yaml
llm-pi-ai:
  providers:
    deepseek:                        # ← use the vendor name so the catalog brings the right switches
      api: openai-completions
      baseURL: https://<your relay>/v1   # ← override the default address
      models:
        - id: deepseek-v4-flash

One precondition you have to confirm yourself: whether your relay really forwards per the official DeepSeek protocol (especially the send/receive semantics of reasoning_content). The reporter marks this as unverified, and this article keeps it that way.

Fix four: capture the outbound request body instead of guessing

The 400 body you hold is the relay's generic wrapper, with no field names:

json
{"message":"Upstream error: 400","type":"invalid_request_error"}

DSH only classifies it by text (packages/llm/llm-pi-ai/src/stream.ts:49 matches 400 / invalid request), and this code is not in the default retry set (packages/llm/llm/src/retry-policy.ts:18-24), so "always" is the terminal state and it will not self-heal. The only thing that can point at which field was rejected is the relay's upstream error text. Two ways to get it:

  1. Run a forwarding proxy locally that prints the request body and forwards it verbatim, and temporarily point the route's baseURL at it;
  2. Ask the relay directly for its upstream request log.

(Within the scope of what the report checked, no built-in "print the outbound request body" switch was found in DSH — DEBUG, DSH_DEBUG, DSH_LOG, logLevel and similar all came up empty; the reporter is explicit that this is "not found", not proof of absence. So the two routes above are the executable paths today.)

Troubleshooting checklist

  1. Separate INVALID_CONFIG from INVALID_REQUEST first. The former is produced on the DSH side, is model-independent, and affects every call on that route; the latter comes from the relay. Mixing them sends you the long way round.
  2. One illegal compat key poisons your entire experiment log. It makes the whole compat block fail at dispatch and the other switches never execute — no conclusion drawn before deleting it is trustworthy.
  3. compat at route level or model level is validated either way (catalog.ts:879, :883), and model overrides route field by field (catalog.ts:342-343). First confirm which level the line sits at and whether another schema swallowed it.
  4. requiresReasoningContentOnAssistantMessages has two hidden preconditions: model.reasoning must be true (false by default on custom routes, so declare reasoningEfforts), and it only writes an empty string.
  5. Do not use "non-reasoning models work" to prove "the relay is fine". Non-reasoning models produce no thinking blocks and never trigger this validation; that control only proves the problem is related to reasoning replay.
  6. The replay shape is decided by thinkingSignature. Asserting "DSH dropped reasoning_content" without reading that field is likely to point the wrong way.
  7. A 400 does not mean it will retry. That code is not in the default retry set, so it is a deterministic terminal state — do not hope a few more attempts will help.
  8. The relay's generic wrapper hides field names. Without the raw request body or the upstream error, every conclusion is an inference — and this article marks every inference as unverified where it applies.

Source: Discussion #7134 — [Bug] DSH desktop 0.1.0-rc.6: custom relay returns 400 when a DeepSeek V4 model calls tools. The source locations and the wire measurements quoted here (the request bodies recorded across two replay rounds, the switch vs reasoningEfforts comparison) come from replies in that thread; the reasoningEfforts conclusions are the reporter's own measurements, and "whether the relay forwards per the official DeepSeek protocol" remains unverified, as marked above.

With problems like custom relays, compat switches and protocol differences, the expensive part is rarely the fix — it is confirming which layer of config you actually changed and whether the change took effect. DSH Plugin Hub offers five views: market, installed list, custom install, settings and system logs. The settings page centralises update checks, npm registry mirrors and proxies, while the system logs page records install, uninstall, settings-change and diagnostic runs by category and level, with full-text export and a jump to the log file — so on config problems you can align environment, version and operation history before touching YAML.

DSH Plugin Hub system logs view: execution traces by category and level, with full-text export and a jump to the log file

FAQ

Same relay, same config: `gpt-5.6-terra` calls tools fine, but `deepseek-v4-flash` always returns 400. Is the relay unsupported?

More likely the *replay shape of a reasoning model* does not match what the relay validates. A non-reasoning model has no thinking blocks at all, so there is no reasoning content in the assistant history and replay is never validated; a reasoning model carries its thinking back into the history, and which key carries it (reasoning_content or structured reasoning_details) depends on the shape the relay originally returned. If the relay returned structured reasoning_details, replay carries only reasoning_details and no reasoning_content — and a relay that **requires** non-empty reasoning_content on tool-calling turns will reject it. This matches your control test exactly.

I added `suppressReasoningContentReplay: false` to compat, but it had no effect at all. Why?

Because that key does not exist anywhere in pi-ai 0.85.1 — zero hits across the whole repository. More importantly, an unknown compat key is **not silently ignored**: assertOfferedCompatFields flags it at dispatch, names it, reports INVALID_CONFIG and lists the switches that really are configurable. So as long as that line really sits under compat:, **every model call** on that route fails before it ever reaches the relay — which also means the other four switches you tuned never took effect either. Step one is to delete the key and retest.

`requiresReasoningContentOnAssistantMessages: true` still does nothing. Is that a DSH bug?

On a custom route it is blocked by an extra condition. pi-ai decides it as compat.requiresReasoningContentOnAssistantMessages && model.reasoning && assistantMsg.reasoning_content === undefined. A custom route has no entry in the installed catalog description, so model.reasoning resolves as base?.reasoning ?? false → false — unless your model entry explicitly declares reasoningEfforts. Measured: with the switch true but no reasoningEfforts, the outbound request does not even contain a reasoning_content key; declare reasoningEfforts and the same switch produces "reasoning_content": "". And its semantics are only **to pad an empty string**, not to put the real thinking content back — so even when it does take effect it cannot fix a relay that requires non-empty content.

I suspect DSH drops reasoning_content from assistant messages that carry tool_calls. How do I verify that?

That shape is wrong. Which reasoning field replay writes depends on which of reasoning / reasoning_content / reasoning_text the thinking block's thinkingSignature falls into; the tool_calls mapping happens **after** the reasoning write and does not participate in that decision, so "no reasoning_content because there are tool_calls" is not what that code does. The fastest check is to open the first assistant message of the failing session and look at source.replayState.blocks[0].thinkingSignature: starting with [{"type":"reasoning means the structured-details shape; exactly "reasoning_content" means the plain-field shape, in which case the field really was sent and the 400 has to be found elsewhere.

How do I see the request body DSH actually sends? Is there a built-in switch?

Within the scope of what the report checked, **no** built-in "print the outbound request body" switch was found (searching DEBUG, DSH_DEBUG, DSH_LOG, logLevel and similar turned up nothing — though the reporter notes this is "not found", not proven absent). What works is to run a forwarding proxy on your machine that prints the body, point the route's baseURL at it and have it forward verbatim to the relay; or simply ask the relay for its upstream request log. The reason is practical: the 400 body you hold is the relay's generic wrapper and **contains no field name**, so without the raw request body you can only guess.

Related Terms

compat (compatibility switch block)
The config block in `llm-pi-ai` that describes how an endpoint differs from the standard OpenAI protocol; it can be written at route level and model level, with model overriding route field by field. It only accepts keys that some wire protocol has declared — undeclared keys are flagged at dispatch and named, and keys with empty values are rejected too.— https://github.com/deepseek-ai/deepseek-harness/discussions/7134
requiresReasoningContentOnAssistantMessages
One of the compat switches; it means "this endpoint requires assistant messages that carry tool calls to include a `reasoning_content` field". In practice it only fires when `model.reasoning` is true, and it writes the **empty string** `""` rather than real thinking content. On a custom route the catalog does not describe, `model.reasoning` defaults to `false`, so the switch silently goes inactive unless the model entry explicitly declares `reasoningEfforts`.— https://github.com/deepseek-ai/deepseek-harness/discussions/7134
thinkingSignature (thinking block signature)
A marker persisted on the thinking block of an assistant message, recording the shape in which that thinking arrived over the wire. Replay uses it to decide whether to write `reasoning_content` or structured `reasoning_details`. When a provider only reports thinking as structured details, the durable message keeps just an empty `reasoning` block and the text exists only inside this opaque signature — if the signature is lost, the thinking is neither visible nor recoverable.— https://github.com/deepseek-ai/deepseek-harness/discussions/7134
catalog route vs hand-written route
A catalog route is described by the built-in catalog and carries the compatibility switches that vendor actually needs; a hand-written route (custom provider id plus your own `models` list) has no catalog description, so pi-ai can only infer switches from the provider id and baseURL — and "a private gateway URL says nothing". For an unrecognised endpoint the inference amounts to "it is OpenAI itself", which is wrong for most OpenAI-compatible gateways. `baseURL` can be overridden, so naming the route after the vendor (e.g. `deepseek`) and pointing `baseURL` at your own relay is a common workaround.— https://github.com/deepseek-ai/deepseek-harness/discussions/7134

Sources