DSH plugin: dsh_session_log 413 self-lock and how to recover
When a DSH session suddenly cannot send any message and every turn reports DeepSeek Messages request failed (413), even for a two-character message — it is usually not a context overflow, but the telemetry field dsh_session_log packing the whole session log into the request body (205.87 MB measured, 99.7% of the body) and hitting the front gateway's size limit; worse, the callback that advances the "watermark" only runs after a successful request, so "too large → fail → watermark stays → next time even larger → still too large" becomes a permanent self-lock. This article splits into "how the self-lock forms → the two trigger paths and their version attribution → multiple recovery options", each with paste-ready commands and configuration.
How the 413 self-lock forms: the watermark only advances after success
The most counterintuitive part of this error is that retries never work, because the failure itself blocks the only repair path.
Scene signature: three signals visible at a glance
- Even
pingfails: message content and size do not matter; sending one word or two characters (one user measured sending "继续") immediately returns 413, and there is nollm/retryevent between turns (#7658, #7699). - One request turn looks like this in the local log:
compaction/start {turn:946}
compaction/end {turn:946, error:"DeepSeek Messages request failed (413)"}
turn/end {turn:946, reason:{kind:"error", code:"CONTEXT_WINDOW_EXCEEDED", status:413}}
- The token gauge on screen stalls without any alarm: one user saw surface tokens stuck at 380,109 against a 450,000 window, with "nothing unusual" on the UI, while the real request body was already 206 MB (#7699).
Measured evidence: 0.58 MB of real conversation vs a 205.87 MB field
The trick of "change the message count → see how the body changes" isolates the constant. One real payload captured on the adapter side:
[DSH-BYTEGUARD] {"purpose":"compaction","ceiling":8388608,
"baseBytes":579474,
"fieldBytes":{"dsh_session_log":205872373,"dsh_plugin_packages":5916},
"dropped":["dsh_session_log"],"payloadBytes":585413}
| Quantity | Value |
|---|---|
The dsh_session_log field | 205,872,373 bytes (205.87 MB) |
Real conversation surface baseBytes | 579,474 bytes (0.58 MB) |
| Rejected request bodies (3 samples) | 206,429,335 / 207,740,314 / 206,456,506 bytes |
| Main request message count | 863; compaction request 208 |
The last row is the key: message counts differ by 4x (208 vs 863) but the body differs by only 1.3 MB — so a ~206 MB constant travels with every request, unrelated to conversation content (#7699). Another user's session had 29,290 events and an 84.9 MB decompressed log, with breakdown dsh_session_log 94.99 MB, messages 1.48 MB, tools 28 KB (#7658).
The self-lock cycle: four steps, locking itself
- The
dsh-session-log-deepseekplugin injects a top-leveldsh_session_logfield into every DeepSeek request, containing "all session events after the last accepted watermark". - The watermark is advanced by
session-log-deepseek/delivery-accepted, whichaccept()writes, anddsh-llm-deepseekonly calls it after HTTP 2xx. - The field crosses the gateway limit → the gateway returns 413 →
accept()does not run → the watermark does not move. - The next request resends the same payload (or a larger one as events accumulate) → still 413 → permanently stuck (#7699).
Source-level confirmation nails these three lines down: session-log-deepseek:111/113's if (acceptedFormatVersion !== session.header.version) continue decides the watermark set; the migration path at :118-150 has no re-anchoring; and :185's accept is gated on 2xx (community line-by-line review, #7658).
How to tell "413 is a size problem" from "it really is a context overflow"
Two models, two numbers — don't read the wrong gauge. The server-side limit can be measured with a "token-free probe": request https://api.deepseek.com/anthropic/v1/messages with an invalid body; over the limit returns 413, under returns 422, and neither consumes tokens. Measured: 8 / 16 / 24 / 32 MB pass, 48 MB → 413 (#7658).
When troubleshooting, walk steps 1/2/3/4:
- Check whether context metrics are healthy: if
messagesis only 1–2 MB andcontextPressure.surfaceTokensis far belowcontextWindow(e.g. 279,257 / 1,000,000), a "context overflow" is basically ruled out (#7658). - Capture one real payload on the request path: get the per-field byte counts and confirm whether one top-level field dominates.
- Run the count contrast: compare message counts of the main and compaction requests; if the counts differ several-fold while size barely moves, a constant field is at work (#7699).
- Look at the gateway response body: if it is openresty's HTML error page
413 Request Entity Too Large, you hit the front gateway, not the model API:
HTTP/1.1 413
server: openresty
content-type: text/html
eo-cache-status: MISS
<html><head><title>413 Request Entity Too Large</title></head>
<body><center><h1>413 Request Entity Too Large</h1></center></body></html>
This is also why the adapter classifier is blind: when the response is not JSON, the fallback text DeepSeek Messages request failed (413) matches no pattern in isContextWindowExceededError(), so it is classified as INVALID_REQUEST and automatic compaction recovery never triggers (#7699).
The two trigger paths and their version attribution: why it "suddenly hit en masse"
Both are afterSeq = -1, but two different states: "the watermark was filtered out" and "the watermark was never established". Distinguish them to know which half the fix belongs in.
Path A: after V3→V4 migration, the old generation's watermark is filtered by a read-side equality
After the V3→V4 session-format migration, the old delivery-accepted records are still in the log, field for field, only skipped by an equality on the read side. That equality is:
const acceptedFormatVersion = event.data.sessionFormatVersion ?? 0; // session-log-deepseek:111
if (acceptedFormatVersion !== session.header.version) continue; // :113
Scoping the watermark by generation is deliberate (only "the highest confirmed sequence within this session-format generation" counts), but the migration path lacks re-anchoring for "resending is undeliverable", and the provider-side size limit directly conflicts with the assumption that "resending is always deliverable" (#7658).
A contributor further notes: invalidation is not deletion but read-side filtering; and the migration stage reassigns sequence numbers (nextSeq = 0, mapping[sourceSeq] = targetSeq), while remapV3References's transform table covers only six types and delivery-accepted is not among them — so a pre-migration marker's throughSeq is a number in the old coordinate space: "the number exists, what is missing is its meaning". This explains why simply "treating the old watermark as the new lower bound" is not safe (a one-sided community conclusion, #7658).
Path B: the watermark was never established (the default flip is the trigger)
The other path needs no migration at all: an old session simply never wrote a single delivery-accepted. One user measured a session that "lived 12 days and had 40,717 events, with not one watermark"; its 14 accepted events were all sessionFormatVersion: 4 (the same as the header, so it is not path A) (#7658). The root cause is a default-value flip:
| 0.1.5-rc.2 / rc.3 | 0.1.7-rc.1 / alpha.2 | |
|---|---|---|
dsh-session-log-deepseek's enabled default | false | true |
| README wording | "the default config does not register the request field" | "the default config registers the request field" |
A machine diff of the two bundle trees shows this plugin changed only one line: enabled: false -> true; another machine measured the same false → true at lib/index.js:16. Users who never enabled it are hit after upgrading too — these old sessions have no watermark to anchor, so the very first injection is afterSeq = -1, sending the entire log (#7658).
Blast radius: not one session, a batch
Existing sessions under the same home trip together:
- One user counted: among 29 session generations, 16 had "zero watermark and >500 events", sized 2.0–89.9 MB, each waiting for its first official DeepSeek request to deadlock (#7658).
- Another example: sessions with 32,130 / 23,071 / 21,799 / 11,700 events under the same home all had 0 acceptable records (#7658).
So "I changed nothing, how did it suddenly break" is normal — the upgrade itself swapped the default behavior (#7699).
Why "roll back to the old version" does not save this class
Only "telemetry on by default" was introduced by 0.1.7-alpha.2; the other three are long-standing architectural problems. A side-by-side comparison of two installs on one machine (#7699):
| 0.1.7-alpha.2 | 0.1.5-rc.3 | |
|---|---|---|
| Telemetry field on by default | Yes | No |
Recovery path with retainTokens = 0 | Yes | Yes |
| Empty-body 413 classification defect | Yes | Yes |
| Total request-body byte guard | No | No |
Three long-standing defects are worth remembering separately:
- The recovery action replays the failure verbatim: both overflow recovery and manual
/compactcallselectCompactableRange(session, measurement, 0);retainTokens = 0means "keep nothing, summarize the whole span", so the recovery request is as large as the failed request — fixing a failure with the failing method (#7699). - The error classifier does not recognize non-JSON provider errors (see the previous section).
- There is no "total byte" guard on the request path: the pressure model rests entirely on tokens, and while the adapter has a 20 MB limit for images (
maxInlineRequestImageBytes), there is none for the body as a whole;DEFAULT_CONTEXT_WINDOW = 1e6/DEFAULT_MAX_TOKENS = 256e3push the compaction trigger to about 678k tokens — a guard behind the wall (#7699).
Multiple recovery options: from "rescue the session" to "fix the root"
The priority is: first make the session able to send messages, then decide whether to keep the telemetry feature. The four options below are ordered by increasing intrusiveness; Option 1 stops the pain immediately on nearly every platform.
Option 1: disable the injecting field (fastest, all platforms)
This is the community-verified immediate fix; the request body drops from 96.5 MB to 1.53 MB at once and the same session answers normally. Key detail: the patch must go at the home layer ($DSH_HOME/cordis.patch.yml); editing only profiles/web/cordis.patch.yml does not cover tui / headless (#7658).
- Shut down all DSH processes.
- Edit the home-layer patch file and append the following (when
config.enabled !== truethe plugin returns before registering the field, so nothing is injected):
- id: session-log-deepseek
config:
enabled: false
- Restart DSH and open the previously stuck session.
- Verify: send a minimal message (e.g. "continue"); expect a normal answer, and if logging is on, no new
delivery-acceptedshould be produced.
Cost: the official API no longer receives the session-log suffix, on-demand reads of the raw session log stop working, and /feedback delivery for that profile fails. Roll back by deleting the block and restarting (#7658).
Option 2: keep the feature with a byte-budget plugin (recommended if you want the logs)
If you don't want to give up server-side raw-log retrieval, you can put a byte budget on this field without touching the core. The community plugin @argszero/cordis-plugin-session-log-budget wraps the registry's prepare and, when over budget, sends only the longest loadable prefix (afterSeq + 1 … throughSeq remains true; a tail window cannot say that) and writes its own accepted record covering just that prefix (#7658).
- Install:
npm install @argszero/cordis-plugin-session-log-budget
- Mount it in the profile's
cordis.patch.ymland tune the budget:
- insert:
- id: session-log-budget
name: '@argszero/cordis-plugin-session-log-budget'
config:
mode: enforce # or report: measure and record only, change nothing
maxFieldBytes: 8000000
- Restart, open the stuck session, and send several turns: the pending backlog drains one budget's worth per turn, then returns to the normal incremental path.
- To observe before acting, set
modetoreportand look at the measured field bytes first.
Know the boundaries: it cannot save the first request's construction cost (by the time the plugin sees a 90 MB value it has already been built); during draining the server first receives the oldest segment; a shortened request replaces the registry's joint-accept transaction (today dsh_session_log is the product's only such field); and an envelope larger than the whole budget is undeliverable — the field is dropped and named in the host log (#7658).
Option 3: upgrade to a build with maxBytes
The 0.1.7-rc.2 artifacts already contain maxBytes (default 8 MiB) and "take the longest pending prefix that fits" behavior. The community judges that upgrading both stops new occurrences and lets stuck sessions drain prefix by prefix (#7658). But know the other half at the same time: the equality that breaks the watermark (:113) is unchanged; the "cap" and the "invalidation" are two halves, and only one moved (#7658).
- Upgrade to
0.1.7-rc.2or newer. - Open the stuck session and send several turns to drain the backlog; if a single turn still fails, even the first prefix is over the limit, and you need Option 1 or 2 as well.
- After upgrading, still check existing sessions' watermarks: using Option 1 to disable the field to rescue old sessions is the more robust order.
Option 4: build your own "total request-body byte gate" (for developers)
The root fix direction is "any optional contribution injected into a request must have a byte cap; over the cap, truncate or skip, never throw". In the community's reference implementation (3 files changed), the dsh-llm-deepseek side adds four things: an extension-field byte gate, a whole-envelope byte ladder, a per-request inline-image limit, and the classifier's isBodylessStatus; the dsh-compaction-basic side replaces retainTokens = 0 with a budget that "guarantees something to summarize while capping the summarize input at 131,072 tokens" (#7699). Measured effect:
bodyBytes=585413 -> allowed, watermark advances normally
dsh_session_log: 205.87 MB -> 93 KB
A four-step self-check (implementation-agnostic, so you can align your own patch):
- Add a byte gate: cap every extension field and the whole body at request assembly; over the cap, degrade (drop the field / downscale images / truncate long text) and never throw.
- Ensure recovery requests are strictly smaller: audit every
retainTokens = 0recovery path; it necessarily produces a body equal to the failed one. - Fix the classifier: classify empty-body / HTML-body 400/413 as a size or context overflow and expose the raw response body.
- Put
bodyBytesinto health metrics: a token-only gauge is actively misleading when a huge extension field exists.
A three-step self-rescue for non-programmers
You can recover without writing code; do these in order (#7658):
- Fully quit DSH (including background processes).
- Find
cordis.patch.ymlunder the DSH home directory and append Option 1'ssession-log-deepseek / enabled: falseblock at the end; copy the file as a backup first. - Restart DSH, open the stuck session and send a message to verify. Once confirmed, if you later want logging back, delete the appended block to roll back.
Troubleshooting notes
This failure is deceptive: the UI shows "plenty of tokens", the error is sometimes transport failed, and reinstalling or switching models does nothing. Nine points:
- First tell "context overflow" from "request-body overflow": look at the actual bytes on the
messagessurface and atsurfaceTokens; they are different models (#7658). - 413 does not equal a context overflow: in this scenario 413 is the front gateway rejecting body size, with an HTML response body (#7699).
- A screen full of
transport failedmay be this: 413 is classified asINVALID_REQUEST, the retry allowlist excludes it, and uploading a large body approaches the watchdog, so it eventually surfaces as TRANSPORT (#7658). - Don't do useless things: reinstalling, switching models, and switching networks do not help; the pollution source is in the session records and in a plugin whose default is
true(#7699). /compactnot only fails but replays the failure verbatim: the recovery path usesretainTokens = 0, so the recovery request is as large as the failed one (#7699).- Patch at the home layer: only
$DSH_HOME/cordis.patch.ymlcoverstui/headless(#7658). - Upgrading is not a cure-all: rc.2 only added the "cap", not "post-migration watermark invalidation", so existing old sessions can still lock up (#7658).
- Small sessions can be hit too: one user reported sessions with 132/58 events and a field of ≈0 MB also reporting TRANSPORT, suggesting an independent cause that needs separate investigation (#7658).
--dump-configwon't show it: it prints only explicit config and does not expand schema defaults, so a feature that is on by default and can produce 205 MB of behavior does not appear in any config export (#7699).
Failures like these usually come with a string of install/update/log actions. If you normally manage plugin installs, update confirmations, system logs, and diagnostics together in DSH Plugin Hub, you can at least first confirm on the system-logs page whether it is a plugin-layer or request-layer problem before deciding to touch config files.

Source: Discussion #7658, Discussion #7699. Also see the supplementary thread Discussion #7753 and the summary thread Discussion #7737 cited in those discussions.
FAQ
Here 413 is not a context overflow but the request body exceeding the server-side limit. In measurements the real conversation messages were only 0.58 MB, while the dsh_session_log telemetry extension field took 205.87 MB, so the whole ~206 MB body was rejected by the gateway with 413. Context pressure and request-body byte count are two different models; the former can look roomy (380k/1M) while the latter has already hit the wall.
dsh_session_log injection degrades to 'send the entire log' when the watermark is missing or invalid. There are two routes: after a V3->V4 session-format migration, the old generation's delivery-accepted records are filtered out by a read-side equality; or the field's default flipped from false to true in 0.1.7-alpha.2/rc.1, and the old session never established a watermark. Both make afterSeq fall back to -1.
It does not affect conversation; it is an optional telemetry extension field. Disabling it drops the request body from 96.5 MB to about 1.53 MB immediately and the same session answers normally. The cost is that the official API no longer receives the session-log suffix, on-demand reads of the raw session log stop working, and server-side /feedback delivery for that profile fails. A new session that keeps it on usually pays one larger upload only on the first request.
Per the community's review of 0.1.7-rc.2 artifacts, that version adds maxBytes (default 8 MiB) and 'take the longest pending prefix that fits' behavior, so after upgrading new occurrences are capped and a stuck session can drain prefix by prefix across turns. But the equality that breaks the watermark was not changed: the cap and the invalidation are two halves, and only one half moved.
Because 413/400 is classified as INVALID_REQUEST, and the retry allowlist does not include it; meanwhile uploading a ~48 MB body takes long enough to approach the idle watchdog, so the fallback branch surfaces it as TRANSPORT. The community therefore judges that a screen full of transport failed is very likely this problem's surface, so do not only investigate in the network direction.
Related Terms
- watermark (delivery-accepted watermark)
- The watermark is what session-log-deepseek uses to record 'the highest sequence number the server has accepted for the session log', stored in the session-log-deepseek/delivery-accepted event. Each request sends only the increment after the watermark; the watermark is advanced by accept() after a successful (HTTP 2xx) request, which is exactly why a 413 can become a self-lock.— https://github.com/deepseek-ai/deepseek-harness/discussions/7658
- dsh_session_log
- dsh_session_log is a top-level extension field on dsh-llm-deepseek requests, injected by the dsh-session-log-deepseek plugin, containing 'all session events after the last accepted watermark'. It is for telemetry/logging, not model input, so context-compaction machinery cannot see it and cannot bound it.— https://github.com/deepseek-ai/deepseek-harness/discussions/7699
- 413 Request Entity Too Large
- 413 is the HTTP status code meaning the server/front gateway refuses a request body that exceeds its size limit. In this DSH scenario it is returned by the openresty front gateway with an HTML error page, not by the DeepSeek model API as JSON, which is why the adapter's error classifier fails to recognize it and misclassifies it as INVALID_REQUEST.— https://github.com/deepseek-ai/deepseek-harness/discussions/7699
- request extension field
- A request extension field is a top-level field a plugin attaches to a model request through the registry (such as dsh_session_log, dsh_plugin_packages). They are merged with the body at request assembly time (JSON.stringify({...body, ...extensions.fields})), so they are not governed by the token-based context pressure model and need a separate byte cap.— https://github.com/deepseek-ai/deepseek-harness/discussions/7658
Sources
- #7658 — dsh_session_log watermark breaks after session-format migration: every request resends the whole log -> 413 and a permanent lock· deepseek-ai (GitHub Discussions)
- #7699 — The session-log telemetry field is on by default, inflates the request body to 205.87 MB, and permanently locks the session· deepseek-ai (GitHub Discussions)