DSH plugin: dsh_session_log 413 self-lock and how to recover

TroubleshootingPublished 2026-10-03Author: DeepSeek Plugin Market
DeepSeek HarnessDSH413dsh_session_logsession logrequest body too largestuck session
On 0.1.7, a long session fails with 413: dsh_session_log packs the full log into the request body and watermark only advances after 2xx — permanent self-lock.

When a DSH session suddenly cannot send any message and every turn reports DeepSeek Messages request failed (413), even for a two-character message — it is usually not a context overflow, but the telemetry field dsh_session_log packing the whole session log into the request body (205.87 MB measured, 99.7% of the body) and hitting the front gateway's size limit; worse, the callback that advances the "watermark" only runs after a successful request, so "too large → fail → watermark stays → next time even larger → still too large" becomes a permanent self-lock. This article splits into "how the self-lock forms → the two trigger paths and their version attribution → multiple recovery options", each with paste-ready commands and configuration.

How the 413 self-lock forms: the watermark only advances after success

The most counterintuitive part of this error is that retries never work, because the failure itself blocks the only repair path.

Scene signature: three signals visible at a glance

  1. Even ping fails: message content and size do not matter; sending one word or two characters (one user measured sending "继续") immediately returns 413, and there is no llm/retry event between turns (#7658, #7699).
  2. One request turn looks like this in the local log:
compaction/start {turn:946}
compaction/end   {turn:946, error:"DeepSeek Messages request failed (413)"}
turn/end         {turn:946, reason:{kind:"error", code:"CONTEXT_WINDOW_EXCEEDED", status:413}}
  1. The token gauge on screen stalls without any alarm: one user saw surface tokens stuck at 380,109 against a 450,000 window, with "nothing unusual" on the UI, while the real request body was already 206 MB (#7699).

Measured evidence: 0.58 MB of real conversation vs a 205.87 MB field

The trick of "change the message count → see how the body changes" isolates the constant. One real payload captured on the adapter side:

[DSH-BYTEGUARD] {"purpose":"compaction","ceiling":8388608,
  "baseBytes":579474,
  "fieldBytes":{"dsh_session_log":205872373,"dsh_plugin_packages":5916},
  "dropped":["dsh_session_log"],"payloadBytes":585413}
QuantityValue
The dsh_session_log field205,872,373 bytes (205.87 MB)
Real conversation surface baseBytes579,474 bytes (0.58 MB)
Rejected request bodies (3 samples)206,429,335 / 207,740,314 / 206,456,506 bytes
Main request message count863; compaction request 208

The last row is the key: message counts differ by 4x (208 vs 863) but the body differs by only 1.3 MB — so a ~206 MB constant travels with every request, unrelated to conversation content (#7699). Another user's session had 29,290 events and an 84.9 MB decompressed log, with breakdown dsh_session_log 94.99 MB, messages 1.48 MB, tools 28 KB (#7658).

The self-lock cycle: four steps, locking itself

  1. The dsh-session-log-deepseek plugin injects a top-level dsh_session_log field into every DeepSeek request, containing "all session events after the last accepted watermark".
  2. The watermark is advanced by session-log-deepseek/delivery-accepted, which accept() writes, and dsh-llm-deepseek only calls it after HTTP 2xx.
  3. The field crosses the gateway limit → the gateway returns 413 → accept() does not run → the watermark does not move.
  4. The next request resends the same payload (or a larger one as events accumulate) → still 413 → permanently stuck (#7699).

Source-level confirmation nails these three lines down: session-log-deepseek:111/113's if (acceptedFormatVersion !== session.header.version) continue decides the watermark set; the migration path at :118-150 has no re-anchoring; and :185's accept is gated on 2xx (community line-by-line review, #7658).

How to tell "413 is a size problem" from "it really is a context overflow"

Two models, two numbers — don't read the wrong gauge. The server-side limit can be measured with a "token-free probe": request https://api.deepseek.com/anthropic/v1/messages with an invalid body; over the limit returns 413, under returns 422, and neither consumes tokens. Measured: 8 / 16 / 24 / 32 MB pass, 48 MB → 413 (#7658).

When troubleshooting, walk steps 1/2/3/4:

  1. Check whether context metrics are healthy: if messages is only 1–2 MB and contextPressure.surfaceTokens is far below contextWindow (e.g. 279,257 / 1,000,000), a "context overflow" is basically ruled out (#7658).
  2. Capture one real payload on the request path: get the per-field byte counts and confirm whether one top-level field dominates.
  3. Run the count contrast: compare message counts of the main and compaction requests; if the counts differ several-fold while size barely moves, a constant field is at work (#7699).
  4. Look at the gateway response body: if it is openresty's HTML error page 413 Request Entity Too Large, you hit the front gateway, not the model API:
HTTP/1.1 413
server: openresty
content-type: text/html
eo-cache-status: MISS

<html><head><title>413 Request Entity Too Large</title></head>
<body><center><h1>413 Request Entity Too Large</h1></center></body></html>

This is also why the adapter classifier is blind: when the response is not JSON, the fallback text DeepSeek Messages request failed (413) matches no pattern in isContextWindowExceededError(), so it is classified as INVALID_REQUEST and automatic compaction recovery never triggers (#7699).

The two trigger paths and their version attribution: why it "suddenly hit en masse"

Both are afterSeq = -1, but two different states: "the watermark was filtered out" and "the watermark was never established". Distinguish them to know which half the fix belongs in.

Path A: after V3→V4 migration, the old generation's watermark is filtered by a read-side equality

After the V3→V4 session-format migration, the old delivery-accepted records are still in the log, field for field, only skipped by an equality on the read side. That equality is:

ts
const acceptedFormatVersion = event.data.sessionFormatVersion ?? 0;   // session-log-deepseek:111
if (acceptedFormatVersion !== session.header.version) continue;       // :113

Scoping the watermark by generation is deliberate (only "the highest confirmed sequence within this session-format generation" counts), but the migration path lacks re-anchoring for "resending is undeliverable", and the provider-side size limit directly conflicts with the assumption that "resending is always deliverable" (#7658).

A contributor further notes: invalidation is not deletion but read-side filtering; and the migration stage reassigns sequence numbers (nextSeq = 0, mapping[sourceSeq] = targetSeq), while remapV3References's transform table covers only six types and delivery-accepted is not among them — so a pre-migration marker's throughSeq is a number in the old coordinate space: "the number exists, what is missing is its meaning". This explains why simply "treating the old watermark as the new lower bound" is not safe (a one-sided community conclusion, #7658).

Path B: the watermark was never established (the default flip is the trigger)

The other path needs no migration at all: an old session simply never wrote a single delivery-accepted. One user measured a session that "lived 12 days and had 40,717 events, with not one watermark"; its 14 accepted events were all sessionFormatVersion: 4 (the same as the header, so it is not path A) (#7658). The root cause is a default-value flip:

0.1.5-rc.2 / rc.30.1.7-rc.1 / alpha.2
dsh-session-log-deepseek's enabled defaultfalsetrue
README wording"the default config does not register the request field""the default config registers the request field"

A machine diff of the two bundle trees shows this plugin changed only one line: enabled: false -> true; another machine measured the same false → true at lib/index.js:16. Users who never enabled it are hit after upgrading too — these old sessions have no watermark to anchor, so the very first injection is afterSeq = -1, sending the entire log (#7658).

Blast radius: not one session, a batch

Existing sessions under the same home trip together:

  • One user counted: among 29 session generations, 16 had "zero watermark and >500 events", sized 2.0–89.9 MB, each waiting for its first official DeepSeek request to deadlock (#7658).
  • Another example: sessions with 32,130 / 23,071 / 21,799 / 11,700 events under the same home all had 0 acceptable records (#7658).

So "I changed nothing, how did it suddenly break" is normal — the upgrade itself swapped the default behavior (#7699).

Why "roll back to the old version" does not save this class

Only "telemetry on by default" was introduced by 0.1.7-alpha.2; the other three are long-standing architectural problems. A side-by-side comparison of two installs on one machine (#7699):

0.1.7-alpha.20.1.5-rc.3
Telemetry field on by defaultYesNo
Recovery path with retainTokens = 0YesYes
Empty-body 413 classification defectYesYes
Total request-body byte guardNoNo

Three long-standing defects are worth remembering separately:

  1. The recovery action replays the failure verbatim: both overflow recovery and manual /compact call selectCompactableRange(session, measurement, 0); retainTokens = 0 means "keep nothing, summarize the whole span", so the recovery request is as large as the failed request — fixing a failure with the failing method (#7699).
  2. The error classifier does not recognize non-JSON provider errors (see the previous section).
  3. There is no "total byte" guard on the request path: the pressure model rests entirely on tokens, and while the adapter has a 20 MB limit for images (maxInlineRequestImageBytes), there is none for the body as a whole; DEFAULT_CONTEXT_WINDOW = 1e6 / DEFAULT_MAX_TOKENS = 256e3 push the compaction trigger to about 678k tokens — a guard behind the wall (#7699).

Multiple recovery options: from "rescue the session" to "fix the root"

The priority is: first make the session able to send messages, then decide whether to keep the telemetry feature. The four options below are ordered by increasing intrusiveness; Option 1 stops the pain immediately on nearly every platform.

Option 1: disable the injecting field (fastest, all platforms)

This is the community-verified immediate fix; the request body drops from 96.5 MB to 1.53 MB at once and the same session answers normally. Key detail: the patch must go at the home layer ($DSH_HOME/cordis.patch.yml); editing only profiles/web/cordis.patch.yml does not cover tui / headless (#7658).

  1. Shut down all DSH processes.
  2. Edit the home-layer patch file and append the following (when config.enabled !== true the plugin returns before registering the field, so nothing is injected):
yaml
- id: session-log-deepseek
  config:
    enabled: false
  1. Restart DSH and open the previously stuck session.
  2. Verify: send a minimal message (e.g. "continue"); expect a normal answer, and if logging is on, no new delivery-accepted should be produced.

Cost: the official API no longer receives the session-log suffix, on-demand reads of the raw session log stop working, and /feedback delivery for that profile fails. Roll back by deleting the block and restarting (#7658).

If you don't want to give up server-side raw-log retrieval, you can put a byte budget on this field without touching the core. The community plugin @argszero/cordis-plugin-session-log-budget wraps the registry's prepare and, when over budget, sends only the longest loadable prefix (afterSeq + 1 … throughSeq remains true; a tail window cannot say that) and writes its own accepted record covering just that prefix (#7658).

  1. Install:
bash
npm install @argszero/cordis-plugin-session-log-budget
  1. Mount it in the profile's cordis.patch.yml and tune the budget:
yaml
- insert:
    - id: session-log-budget
      name: '@argszero/cordis-plugin-session-log-budget'
      config:
        mode: enforce        # or report: measure and record only, change nothing
        maxFieldBytes: 8000000
  1. Restart, open the stuck session, and send several turns: the pending backlog drains one budget's worth per turn, then returns to the normal incremental path.
  2. To observe before acting, set mode to report and look at the measured field bytes first.

Know the boundaries: it cannot save the first request's construction cost (by the time the plugin sees a 90 MB value it has already been built); during draining the server first receives the oldest segment; a shortened request replaces the registry's joint-accept transaction (today dsh_session_log is the product's only such field); and an envelope larger than the whole budget is undeliverable — the field is dropped and named in the host log (#7658).

Option 3: upgrade to a build with maxBytes

The 0.1.7-rc.2 artifacts already contain maxBytes (default 8 MiB) and "take the longest pending prefix that fits" behavior. The community judges that upgrading both stops new occurrences and lets stuck sessions drain prefix by prefix (#7658). But know the other half at the same time: the equality that breaks the watermark (:113) is unchanged; the "cap" and the "invalidation" are two halves, and only one moved (#7658).

  1. Upgrade to 0.1.7-rc.2 or newer.
  2. Open the stuck session and send several turns to drain the backlog; if a single turn still fails, even the first prefix is over the limit, and you need Option 1 or 2 as well.
  3. After upgrading, still check existing sessions' watermarks: using Option 1 to disable the field to rescue old sessions is the more robust order.

Option 4: build your own "total request-body byte gate" (for developers)

The root fix direction is "any optional contribution injected into a request must have a byte cap; over the cap, truncate or skip, never throw". In the community's reference implementation (3 files changed), the dsh-llm-deepseek side adds four things: an extension-field byte gate, a whole-envelope byte ladder, a per-request inline-image limit, and the classifier's isBodylessStatus; the dsh-compaction-basic side replaces retainTokens = 0 with a budget that "guarantees something to summarize while capping the summarize input at 131,072 tokens" (#7699). Measured effect:

bodyBytes=585413  -> allowed, watermark advances normally
dsh_session_log: 205.87 MB -> 93 KB

A four-step self-check (implementation-agnostic, so you can align your own patch):

  1. Add a byte gate: cap every extension field and the whole body at request assembly; over the cap, degrade (drop the field / downscale images / truncate long text) and never throw.
  2. Ensure recovery requests are strictly smaller: audit every retainTokens = 0 recovery path; it necessarily produces a body equal to the failed one.
  3. Fix the classifier: classify empty-body / HTML-body 400/413 as a size or context overflow and expose the raw response body.
  4. Put bodyBytes into health metrics: a token-only gauge is actively misleading when a huge extension field exists.

A three-step self-rescue for non-programmers

You can recover without writing code; do these in order (#7658):

  1. Fully quit DSH (including background processes).
  2. Find cordis.patch.yml under the DSH home directory and append Option 1's session-log-deepseek / enabled: false block at the end; copy the file as a backup first.
  3. Restart DSH, open the stuck session and send a message to verify. Once confirmed, if you later want logging back, delete the appended block to roll back.

Troubleshooting notes

This failure is deceptive: the UI shows "plenty of tokens", the error is sometimes transport failed, and reinstalling or switching models does nothing. Nine points:

  1. First tell "context overflow" from "request-body overflow": look at the actual bytes on the messages surface and at surfaceTokens; they are different models (#7658).
  2. 413 does not equal a context overflow: in this scenario 413 is the front gateway rejecting body size, with an HTML response body (#7699).
  3. A screen full of transport failed may be this: 413 is classified as INVALID_REQUEST, the retry allowlist excludes it, and uploading a large body approaches the watchdog, so it eventually surfaces as TRANSPORT (#7658).
  4. Don't do useless things: reinstalling, switching models, and switching networks do not help; the pollution source is in the session records and in a plugin whose default is true (#7699).
  5. /compact not only fails but replays the failure verbatim: the recovery path uses retainTokens = 0, so the recovery request is as large as the failed one (#7699).
  6. Patch at the home layer: only $DSH_HOME/cordis.patch.yml covers tui / headless (#7658).
  7. Upgrading is not a cure-all: rc.2 only added the "cap", not "post-migration watermark invalidation", so existing old sessions can still lock up (#7658).
  8. Small sessions can be hit too: one user reported sessions with 132/58 events and a field of ≈0 MB also reporting TRANSPORT, suggesting an independent cause that needs separate investigation (#7658).
  9. --dump-config won't show it: it prints only explicit config and does not expand schema defaults, so a feature that is on by default and can produce 205 MB of behavior does not appear in any config export (#7699).

Failures like these usually come with a string of install/update/log actions. If you normally manage plugin installs, update confirmations, system logs, and diagnostics together in DSH Plugin Hub, you can at least first confirm on the system-logs page whether it is a plugin-layer or request-layer problem before deciding to touch config files.

DSH Plugin Hub · System logs

Source: Discussion #7658, Discussion #7699. Also see the supplementary thread Discussion #7753 and the summary thread Discussion #7737 cited in those discussions.

FAQ

DSH clearly has not exceeded context, so why does it keep reporting DeepSeek Messages request failed (413)?

Here 413 is not a context overflow but the request body exceeding the server-side limit. In measurements the real conversation messages were only 0.58 MB, while the dsh_session_log telemetry extension field took 205.87 MB, so the whole ~206 MB body was rejected by the gateway with 413. Context pressure and request-body byte count are two different models; the former can look roomy (380k/1M) while the latter has already hit the wall.

After upgrading DSH, an old session fails on its first message and retries do not help. Why?

dsh_session_log injection degrades to 'send the entire log' when the watermark is missing or invalid. There are two routes: after a V3->V4 session-format migration, the old generation's delivery-accepted records are filtered out by a read-side equality; or the field's default flipped from false to true in 0.1.7-alpha.2/rc.1, and the old session never established a watermark. Both make afterSeq fall back to -1.

Does disabling dsh_session_log affect normal conversation, and what is the cost?

It does not affect conversation; it is an optional telemetry extension field. Disabling it drops the request body from 96.5 MB to about 1.53 MB immediately and the same session answers normally. The cost is that the official API no longer receives the session-log suffix, on-demand reads of the raw session log stop working, and server-side /feedback delivery for that profile fails. A new session that keeps it on usually pays one larger upload only on the first request.

Does upgrading to 0.1.7-rc.2 automatically unwedge an already stuck session?

Per the community's review of 0.1.7-rc.2 artifacts, that version adds maxBytes (default 8 MiB) and 'take the longest pending prefix that fits' behavior, so after upgrading new occurrences are capped and a stuck session can drain prefix by prefix across turns. But the equality that breaks the watermark was not changed: the cap and the invalidation are two halves, and only one half moved.

Why do I see DeepSeek Messages transport failed instead of a clean 413?

Because 413/400 is classified as INVALID_REQUEST, and the retry allowlist does not include it; meanwhile uploading a ~48 MB body takes long enough to approach the idle watchdog, so the fallback branch surfaces it as TRANSPORT. The community therefore judges that a screen full of transport failed is very likely this problem's surface, so do not only investigate in the network direction.

Related Terms

watermark (delivery-accepted watermark)
The watermark is what session-log-deepseek uses to record 'the highest sequence number the server has accepted for the session log', stored in the session-log-deepseek/delivery-accepted event. Each request sends only the increment after the watermark; the watermark is advanced by accept() after a successful (HTTP 2xx) request, which is exactly why a 413 can become a self-lock.— https://github.com/deepseek-ai/deepseek-harness/discussions/7658
dsh_session_log
dsh_session_log is a top-level extension field on dsh-llm-deepseek requests, injected by the dsh-session-log-deepseek plugin, containing 'all session events after the last accepted watermark'. It is for telemetry/logging, not model input, so context-compaction machinery cannot see it and cannot bound it.— https://github.com/deepseek-ai/deepseek-harness/discussions/7699
413 Request Entity Too Large
413 is the HTTP status code meaning the server/front gateway refuses a request body that exceeds its size limit. In this DSH scenario it is returned by the openresty front gateway with an HTML error page, not by the DeepSeek model API as JSON, which is why the adapter's error classifier fails to recognize it and misclassifies it as INVALID_REQUEST.— https://github.com/deepseek-ai/deepseek-harness/discussions/7699
request extension field
A request extension field is a top-level field a plugin attaches to a model request through the registry (such as dsh_session_log, dsh_plugin_packages). They are merged with the body at request assembly time (JSON.stringify({...body, ...extensions.fields})), so they are not governed by the token-based context pressure model and need a separate byte cap.— https://github.com/deepseek-ai/deepseek-harness/discussions/7658

Sources