DSH plugin resume fails: session already owned (write lock)
A plugin triggers resume and throws session "dsh-xxxxxx-032b77e23bc41872" is already owned by an active write handle, yet you only opened that session in the UI for a look and did nothing else. This is not a false positive — "opening" itself activates the Agent, and activation takes the write lock: after the client opens the live event stream, history.follow() first returns a snapshot, then calls promote(...) for a source in the prepared state, resolving and activating the Agent in the background; the first thing the Agent does once in the agent loop is take the write handle (#7156). As long as that handle lives, it holds the lock. There are actually two lock layers — the in-process tracker.claimWrite(id) and the cross-process session.lock kernel lease — and they throw the same exception with the same message, so you tell them apart by checking whether ctx.agents.get(id) is undefined at the moment of failure (#7156 reply). The correct posture on the plugin side is resume-or-deliver: ask the registry first, deliver if one is live, resume only if nobody holds it, and adopt the winner when you lose the race. But be clear — this only protects your plugin; "making a read-only open genuinely free" is still a core change.
Triage first: two lock layers, one error code, distinguished by one line
Key conclusion: first determine which layer you are colliding with, because "fixing the plugin works" is only true for one of them.
| Criterion | Layer one: in-process claim | Layer two: cross-process lease |
|---|---|---|
| Implementation | tracker.claimWrite(id) (storage.ts:429) | kernel lease on <session dir>/session.lock (lease.ts:71-115) |
| Mechanism | in-process writer tracking table | POSIX exclusive lock; on Windows it surfaces as a sharing conflict |
ctx.agents.get(id) at failure | not undefined (a live Agent in this process) | undefined |
| Fixable by the plugin alone | yes (resume-or-deliver) | no — you must find and end that process first |
| Common cause | you opened that session in the UI | another dsh process still holds that session |
How to use the criterion concretely:
const live = ctx.agents.get(id) // public API: AgentRegistry.get
if (live !== undefined) {
// Layer one: a live Agent in this process holds the claim
} else {
// Layer two: another dsh process holds the exclusive session.lock
}
A pitfall you must know: under POSIX the session.lock file is not deleted after the lease is released (lease.ts). So "the file exists" is not evidence that "someone holds it" — the real holder is determined by the exclusive lock on that file, not by the file itself. Using file existence as your criterion will lead you to the exact opposite conclusion.
Mechanism: why "just opening" takes the write lock
The conclusion of this section: the occupation happens in the background activation after the snapshot is displayed, and the user cannot perceive it.
1. The full chain for opening a session (5 steps)
- The client opens the live event stream;
history.follow()first returns a snapshot;- For a session whose observed source is
prepared, after returning the snapshot it immediately callspromote(...)in the background (packages/api/session-controller/src/history.ts:203-206); promote()triesagents.resume(); if there is no existing Agent, it enters the Agent resume flow (packages/api/session-controller/src/index.ts:173-184);- The Agent loop eventually performs
persistence.open(id, 'write'), atomically seizing the write lock (packages/core/agent-loop/src/index.ts:891-895,packages/session/session-persistence-jsonl/src/index.ts).
Hence three counterintuitive facts:
- "Opened but did not run the model" does not mean "does not hold the write lock" — the mere activation of the Agent takes it;
- Switching to another session does not necessarily release it — the client design keeps the session resident, and off-screen Remote sources keep running;
- The latest code merely normalized the conflict error into
session/writer-held; it did not remove the automatic promote / resume behavior.
Read-only paths such as page() genuinely do not activate the Agent, but a normal UI open goes through the live follow path, so it does.
2. Two lock layers, one exception
- In-process:
tracker.claimWrite(id)—packages/session/session-persistence-jsonl/src/storage.ts:429; - Cross-process: the kernel lease on
<session dir>/session.lock—packages/session/session-persistence-jsonl/src/lease.ts:71-115.
Both throw the same SessionAlreadyOwnedError with the message you saw (packages/session/session-persistence/src/errors.ts:34), and the API path normalizes it into the stable code session/writer-held (packages/api/session-controller/src/agent.ts:218-219).
3. Only the write open claims, and it is deliberately placed first
open()'s read branch never touches ownership (session-persistence-jsonl/src/index.ts:344-364); the write branch claims in its very first step (:365). The agent loop is even more deliberate about taking the lock before doing anything else, and the source comment says it plainly:
// Taking write ownership FIRST excludes a concurrent resume of the
// same id (in this process, a live agent's handle holds the claim).
handle = await raceAbortCall(() => persistence.open(id, 'write', { signal: fused }), /* ... */)
So the exact meaning of that error is: this session already has an activation, and it is holding the handle. Your sequence — "DSH opened the session → your plugin then resumes" — is precisely the situation this comment names.
4. An underrated knock-on: the error code is handled on only one client path
session/writer-held is specially handled in the client in only one place (packages/client/ui-model-selection/src/client/index.ts:169); everywhere else it is either thrown as-is or swallowed. That is why the experience is always "it just does not work" — no explanation, no guidance.
The asymmetry: why a UI open is fine but a plugin resume blows up
The key conclusion of this section: the Host API reuses a live Agent first; ctx.agents.resume() does not.
The Host API's resolution path returns the existing live Agent before attempting any action (liveAgent(sessionId), packages/api/session-controller/src/agent.ts:187-188), so the same session works fine through the UI. ctx.agents.resume() lacks this precondition check — it goes straight to the factory and performs the write open, so it collides with the existing claim.
This explains "the same operation, fine when done manually, an error when done by a plugin" — it is not a permissions problem, it is a different entry point.
Fixes: resume-or-deliver and a ready-made implementation
Fix one (plugin side, preferred): resume-or-deliver
Core idea: do not resume blindly. Ask the registry first; deliver if one is live; resume only once you have confirmed nobody holds it.
const live = ctx.agents.get(id) // public API: AgentRegistry.get
if (live !== undefined) live.followup(message) // deliver to the live Agent — takes no write lock
else await ctx.agents.resume({ resumeSessionId: id }) // safe: nobody holds it right now
This is the shape first-party code in the tree already uses; it is not a trick:
packages/api/session-controller/src/commands.ts:359-361does exactly this for user prompts: first guard that "the Agent I hold is still the one in the registry", thenagent.steer(message)/agent.followup(message);packages/subagent/subagent/src/inbox.ts:54andpackages/schedule/schedule/src/index.ts:50+runtime.ts:161,273do the same: attach to the live Agent, re-check liveness before acting (ctx.agents.get(agent.id) === agent && ctx.agents.roots().includes(agent)), deliver withfollowup, and never resume.
That is, the semantics of "take the lock only when you actually need to run" already exist at the Agent object level; what is missing is only a convenience like resumeIfLive() on the service.
The argument form must be written correctly (this was specifically corrected)
await ctx.agents.resume({ // packages/core/agent/src/index.ts:407
resumeSessionId: id, // the only required field
agentOptions: ctx.agentDefaultModel.currentSelection(), // optional; this is what the Host API passes
setup: composition.setup, // optional
})
The one taking two arguments, (ownerCtx, options), is AgentFactory.resume; a plugin never calls it directly, and the two-argument form does not compile. Verification basis: the canonical call sites packages/api/session-controller/src/agent.ts:437 and packages/bundle/headless/src/index.ts:279, plus the ResumeAgentOptions field table.
One more step: when you lose the race, adopt the winner instead of erroring
There is still a window between "ask first, resume if nobody" — an activation might register in exactly that moment. The robust approach is to accept defeat: if the claim is taken by that just-registered activation, deliver the message to it rather than throwing the failure at the user.
Fix two (ready-made implementation): @argszero/cordis-plugin-session-trigger
If you do not want to write this logic yourself, you can just use this plugin — it was built for exactly this gap:
import { deliverToSession } from '@argszero/cordis-plugin-session-trigger'
const { path } = await deliverToSession({ agents: ctx.agents }, { sessionId, message })
// path: 'live' | 'live-after-race' | 'resumed'
Its behavior is defined clearly:
- Ask the registry first; in the
livecase it never attempts resume; - Resume only when nobody holds it;
- Accept a lost race and adopt the winner — if the claim is taken by an activation registered at the same moment, the message goes to that Agent;
- When it really is held, it gives a diagnosis naming both layers (a write claim in this process / another dsh process holding
session.lock), and explains what to do for each.
Mounting (a cordis.patch.yml insert, with config mode: followup|steer, retain: true|false): it registers ctx.sessionTrigger and retains every Agent it resumes — because the queued turn needs the handle alive; on unload it disposes them all, which is also the action that releases the claim.
There is one detail worth learning on the compatibility side: this package imports no @deepseek-ai/dsh-* module at runtime — it programs only against the documented ctx.agents (get, resume) and the optional ctx.agentDefaultModel (currentSelection), matching rejections by error name with the message text as a fallback. As a result, one build serves the 0.1.2 / 0.1.3 / 0.1.5 / 0.1.6 lines at once.
Its testing is worth noting too: it runs against a real @deepseek-ai/cordis context and the real SessionAlreadyOwnedError exported by the published dsh-session-persistence (not a look-alike stand-in), and it carries a control arm — the naive ctx.agents.resume() it replaces must fail on a live session under the same fixture. 15 tests, plus 6 mutations to the plugin's own logic, each confirmed to turn the suite red.
Boundary and a related lesson: not getting hurt is not the same as "opening" being free
Boundary: this only keeps your plugin from getting hurt, it does not make "opening" free
This is the most important section for expectations; do not skip it.
The core-side design is: every activation takes write ownership up front, and today there is no mechanism for a plugin to make the UI release it. Therefore:
- The plugin-side pattern only protects the plugin's own resume;
- It cannot make "just taking a look" free;
- Your original request — "do not take the lock when opening, only when sending a prompt" — is a core change: either make write lazy, or provide a genuinely read-only follow path.
If you are going to report this, that is the request to state explicitly: not "fix some plugin's error", but "give the UI a read-only open semantic, or make the write claim lazy".
A related lesson: forgetting to declare a dependency at packaging time gives ERR_MODULE_NOT_FOUND on install
The same discussion contains a piece of experience worth calling out separately, because it is a release-pipeline pitfall rather than a logic bug.
@argszero/cordis-plugin-session-trigger's 0.1.0 would not install:
Error [ERR_MODULE_NOT_FOUND]: Cannot find package '@deepseek-ai/schemastery'
imported from .../node_modules/@argszero/cordis-plugin-session-trigger/lib/index.js
The cause is classic: the build artifact lib/index.js imports @deepseek-ai/schemastery (the schema layer behind Config), but 0.1.0 declared it only in devDependencies. In the author's own repository that specifier resolves — pretest builds it, and the package is present in the workspace tree — so the test suite could not see the problem at all. The only thing that exposes it is "installing the published tarball by name in a clean project".
The interesting part: the author had verified by tarball install, and tarball install happens to make the dependency available in the tree; only switching to a by-name install from the registry exposed it. The same bytes, different resolution results.
The fix and hardening (0.1.1, npm latest, repo 936559d):
- Declare
@deepseek-ai/schemasteryindependencies— which is also what in-repo consumers do (packages/core/toolsandpackages/llm/llm-deepseekboth declare it underdependencies, not peer); - Add a
test/packaging.spec.mjs: scan the builtlib/andsrc/for bare specifiers and assert each one appears in the manifest; also assert the reverse — none is declared but un-imported. This guard is not decoration: remove thedependenciesblock and it fails with a precise diagnosis:
lib/ imports @deepseek-ai/schemastery but package.json declares only @deepseek-ai/cordis —
a clean install of the published tarball will fail with ERR_MODULE_NOT_FOUND
Verification on 0.1.1: 18/18 tests (15 behavioral + 3 guard); the guard fails on the real 0.1.0 defect and passes after the fix; the packaged tarball installs and imports; the 15 behavioral tests run against the installed artifact, with imports resolved through node_modules to @argszero/cordis-plugin-session-trigger rather than the local lib/.
The final sentence is the most valuable conclusion: installing by name (npm i <pkg> in an empty directory, then import) should now be the primary check, not tarball install — nothing else can witness an undeclared dependency.
Troubleshooting notes
- Layer first with
ctx.agents.get(id)(not undefined = a live Agent in this process; undefined = a cross-process lease), then talk about fixes. - Do not judge by whether the
session.lockfile exists. Under POSIX the file is not deleted after a lease is released, so its presence or absence says nothing. - "Just opening" is also activation.
history.follow()→promote()→ agent loop → write claim, the whole chain is in the background and the user feels nothing. - Switching sessions away is not release. Off-screen Remote sources keep running.
- Watch the UI / plugin asymmetry: the Host API reuses the live Agent first (
liveAgent(sessionId));ctx.agents.resume()has no such precondition check. ctx.agents.resume()takes only one argument (resumeSessionIdis required); the two-argument form belongs toAgentFactory.resume, do not mix them up.- Do not just catch without handling.
session/writer-heldis specially handled in only one client place; elsewhere it is thrown as-is or swallowed — that is the source of "it just does not work"; when you catch it yourself, give a diagnosis that names both layers. - Understand the
retainsemantics: the queued turn needs the handle alive, so the plugin must retain the Agent it resumes and dispose it on unload — that is when the claim is released. - To be version-portable, do not import
dsh-*packages; match by error name with a message fallback, or you need one build per version line. - Before publishing an npm package, do a final check with "install by name in an empty directory, then import", and add a guard test that "every bare specifier in the build output is in the manifest".
Writing plugins around sessions, locks, and concurrency, the easiest thing to underestimate is "who holds what, and when" — especially when the occupation happens in the background and the user only sees a failure message. DSH Plugin Hub provides five screens: plugin market, installed list, custom install, settings, and system logs. The installed list labels each plugin's source and version and can locate its directory directly; the system log page keeps install, uninstall, and diagnostic trails by category and level and supports exporting the full text; the notification center aggregates history records and the live progress of in-flight tasks. When investigating "which plugin did what, and when", aligning the scene first saves a great deal of guesswork.

Source: Discussion #7156, @argszero/cordis-plugin-session-trigger (npm / source).
FAQ
Because "opening" itself activates the Agent. The chain is: the client opens the live event stream → history.follow() first returns a snapshot → when the observed source is prepared it calls promote(...) → the Agent is resolved and activated in the background → activation starts the agent loop → and the agent loop's first act is to take the write handle. As long as that handle lives, it holds the write lock, so you "just looked" and the lock is already gone.
At the moment of failure, look at ctx.agents.get(id). **Not undefined** means a live Agent in the same process holds it (the tracker.claimWrite(id) layer); **undefined** means another process holds the kernel lease (the session.lock layer). Note: do not decide based on whether the session.lock file exists — under POSIX the file is **not deleted** after the lease is released, so the file's existence proves nothing; the real holder is identified by the exclusive lock on that file.
Because the Host API's resolution path **returns an already-live Agent first** before considering anything else (liveAgent(sessionId)), so the UI path never touches the write open; ctx.agents.resume() has no such precondition check — it goes straight to the factory and performs the write open, so it inevitably collides with the existing claim. This asymmetry is the single most confusing point for plugin authors.
One. The registry wrapper is ctx.agents.resume({ resumeSessionId: id, agentOptions?, setup? }) (packages/core/agent/src/index.ts:407), and resumeSessionId is the only required field. The one taking (ownerCtx, options) is AgentFactory.resume, which a plugin never calls directly. This was specifically corrected in the discussion — the two-argument form does not compile.
Not today. By the core-side design, **every activation takes write ownership up front**, and there is no mechanism for a plugin to make the UI release it. So the request "do not take the lock when opening, only when sending a prompt" is a **core change** (either make write lazy, or provide a genuinely read-only follow path) — not something a plugin can solve on its own. All the plugin side can do is "not become a victim of this behavior".
Related Terms
- SessionAlreadyOwnedError / session/writer-held
- The error thrown when a session already has a writer. The in-process `tracker.claimWrite(id)` and the cross-process `session.lock` lease throw the **same** exception type with the same message (`packages/session/session-persistence/src/errors.ts:34`), and the API path normalizes it into the stable error code `session/writer-held` (`packages/api/session-controller/src/agent.ts:218-219`). Because the two layers share one code, you must find another criterion to tell them apart.— https://github.com/deepseek-ai/deepseek-harness/discussions/7156
- promote (background activation after the snapshot)
- What `history.follow()` calls after returning the snapshot, for sessions whose observed source is `prepared` (`packages/api/session-controller/src/history.ts:203-206`); it resolves/activates the Agent in the background. This is the direct source of "just opening also takes the lock": the moment something is displayed, the write claim is already on its way.— https://github.com/deepseek-ai/deepseek-harness/discussions/7156
- write-first
- The agent loop's deliberate ordering: take the write handle **before** reading or repairing anything, to exclude a concurrent resume of the same id (the comment at `packages/core/agent-loop/src/index.ts:891-895` states this outright). It contrasts with `open()`'s read branch — the read branch never touches ownership (`session-persistence-jsonl/src/index.ts:344-364`), while the write branch claims in its first step (`:365`).— https://github.com/deepseek-ai/deepseek-harness/discussions/7156
- adopt a losing race
- A robustness technique on the plugin side: if, after "ask the registry first, resume only if nobody holds it", you still lose the claim to an activation that registered **at the same moment**, do not error out — deliver the message to **that** Agent instead. It turns "a guaranteed failure" into "at worst one extra step", and is a key part of the resume-or-deliver pattern.— https://github.com/deepseek-ai/deepseek-harness/discussions/7156
Sources
- #7156 — session-persistence seizes session.lock and blocks plugins from resuming· deepseek-ai (GitHub Discussions)