DSH plugin resume fails: session already owned (write lock)

TroubleshootingPublished 2026-10-03Author: DeepSeek Plugin Market
DeepSeek HarnessDSHsession-persistencesession.lockwriter-heldagents.resumeplugin developmentconcurrency locks
A plugin resume says session already owned by an active write handle, although you only opened it. Activation takes the write lock; two lock layers, one code.

A plugin triggers resume and throws session "dsh-xxxxxx-032b77e23bc41872" is already owned by an active write handle, yet you only opened that session in the UI for a look and did nothing else. This is not a false positive — "opening" itself activates the Agent, and activation takes the write lock: after the client opens the live event stream, history.follow() first returns a snapshot, then calls promote(...) for a source in the prepared state, resolving and activating the Agent in the background; the first thing the Agent does once in the agent loop is take the write handle (#7156). As long as that handle lives, it holds the lock. There are actually two lock layers — the in-process tracker.claimWrite(id) and the cross-process session.lock kernel lease — and they throw the same exception with the same message, so you tell them apart by checking whether ctx.agents.get(id) is undefined at the moment of failure (#7156 reply). The correct posture on the plugin side is resume-or-deliver: ask the registry first, deliver if one is live, resume only if nobody holds it, and adopt the winner when you lose the race. But be clear — this only protects your plugin; "making a read-only open genuinely free" is still a core change.

Triage first: two lock layers, one error code, distinguished by one line

Key conclusion: first determine which layer you are colliding with, because "fixing the plugin works" is only true for one of them.

CriterionLayer one: in-process claimLayer two: cross-process lease
Implementationtracker.claimWrite(id) (storage.ts:429)kernel lease on <session dir>/session.lock (lease.ts:71-115)
Mechanismin-process writer tracking tablePOSIX exclusive lock; on Windows it surfaces as a sharing conflict
ctx.agents.get(id) at failurenot undefined (a live Agent in this process)undefined
Fixable by the plugin aloneyes (resume-or-deliver)no — you must find and end that process first
Common causeyou opened that session in the UIanother dsh process still holds that session

How to use the criterion concretely:

ts
const live = ctx.agents.get(id)     // public API: AgentRegistry.get
if (live !== undefined) {
  // Layer one: a live Agent in this process holds the claim
} else {
  // Layer two: another dsh process holds the exclusive session.lock
}

A pitfall you must know: under POSIX the session.lock file is not deleted after the lease is released (lease.ts). So "the file exists" is not evidence that "someone holds it" — the real holder is determined by the exclusive lock on that file, not by the file itself. Using file existence as your criterion will lead you to the exact opposite conclusion.

Mechanism: why "just opening" takes the write lock

The conclusion of this section: the occupation happens in the background activation after the snapshot is displayed, and the user cannot perceive it.

1. The full chain for opening a session (5 steps)

  1. The client opens the live event stream;
  2. history.follow() first returns a snapshot;
  3. For a session whose observed source is prepared, after returning the snapshot it immediately calls promote(...) in the background (packages/api/session-controller/src/history.ts:203-206);
  4. promote() tries agents.resume(); if there is no existing Agent, it enters the Agent resume flow (packages/api/session-controller/src/index.ts:173-184);
  5. The Agent loop eventually performs persistence.open(id, 'write'), atomically seizing the write lock (packages/core/agent-loop/src/index.ts:891-895, packages/session/session-persistence-jsonl/src/index.ts).

Hence three counterintuitive facts:

  • "Opened but did not run the model" does not mean "does not hold the write lock" — the mere activation of the Agent takes it;
  • Switching to another session does not necessarily release it — the client design keeps the session resident, and off-screen Remote sources keep running;
  • The latest code merely normalized the conflict error into session/writer-held; it did not remove the automatic promote / resume behavior.

Read-only paths such as page() genuinely do not activate the Agent, but a normal UI open goes through the live follow path, so it does.

2. Two lock layers, one exception

  • In-process: tracker.claimWrite(id) — packages/session/session-persistence-jsonl/src/storage.ts:429;
  • Cross-process: the kernel lease on <session dir>/session.lock — packages/session/session-persistence-jsonl/src/lease.ts:71-115.

Both throw the same SessionAlreadyOwnedError with the message you saw (packages/session/session-persistence/src/errors.ts:34), and the API path normalizes it into the stable code session/writer-held (packages/api/session-controller/src/agent.ts:218-219).

3. Only the write open claims, and it is deliberately placed first

open()'s read branch never touches ownership (session-persistence-jsonl/src/index.ts:344-364); the write branch claims in its very first step (:365). The agent loop is even more deliberate about taking the lock before doing anything else, and the source comment says it plainly:

ts
// Taking write ownership FIRST excludes a concurrent resume of the
// same id (in this process, a live agent's handle holds the claim).
handle = await raceAbortCall(() => persistence.open(id, 'write', { signal: fused }), /* ... */)

So the exact meaning of that error is: this session already has an activation, and it is holding the handle. Your sequence — "DSH opened the session → your plugin then resumes" — is precisely the situation this comment names.

4. An underrated knock-on: the error code is handled on only one client path

session/writer-held is specially handled in the client in only one place (packages/client/ui-model-selection/src/client/index.ts:169); everywhere else it is either thrown as-is or swallowed. That is why the experience is always "it just does not work" — no explanation, no guidance.

The asymmetry: why a UI open is fine but a plugin resume blows up

The key conclusion of this section: the Host API reuses a live Agent first; ctx.agents.resume() does not.

The Host API's resolution path returns the existing live Agent before attempting any action (liveAgent(sessionId), packages/api/session-controller/src/agent.ts:187-188), so the same session works fine through the UI. ctx.agents.resume() lacks this precondition check — it goes straight to the factory and performs the write open, so it collides with the existing claim.

This explains "the same operation, fine when done manually, an error when done by a plugin" — it is not a permissions problem, it is a different entry point.

Fixes: resume-or-deliver and a ready-made implementation

Fix one (plugin side, preferred): resume-or-deliver

Core idea: do not resume blindly. Ask the registry first; deliver if one is live; resume only once you have confirmed nobody holds it.

ts
const live = ctx.agents.get(id)                  // public API: AgentRegistry.get
if (live !== undefined) live.followup(message)   // deliver to the live Agent — takes no write lock
else await ctx.agents.resume({ resumeSessionId: id })   // safe: nobody holds it right now

This is the shape first-party code in the tree already uses; it is not a trick:

  • packages/api/session-controller/src/commands.ts:359-361 does exactly this for user prompts: first guard that "the Agent I hold is still the one in the registry", then agent.steer(message) / agent.followup(message);
  • packages/subagent/subagent/src/inbox.ts:54 and packages/schedule/schedule/src/index.ts:50 + runtime.ts:161,273 do the same: attach to the live Agent, re-check liveness before acting (ctx.agents.get(agent.id) === agent && ctx.agents.roots().includes(agent)), deliver with followup, and never resume.

That is, the semantics of "take the lock only when you actually need to run" already exist at the Agent object level; what is missing is only a convenience like resumeIfLive() on the service.

The argument form must be written correctly (this was specifically corrected)

ts
await ctx.agents.resume({                                   // packages/core/agent/src/index.ts:407
  resumeSessionId: id,                                      // the only required field
  agentOptions: ctx.agentDefaultModel.currentSelection(),    // optional; this is what the Host API passes
  setup: composition.setup,                                  // optional
})

The one taking two arguments, (ownerCtx, options), is AgentFactory.resume; a plugin never calls it directly, and the two-argument form does not compile. Verification basis: the canonical call sites packages/api/session-controller/src/agent.ts:437 and packages/bundle/headless/src/index.ts:279, plus the ResumeAgentOptions field table.

One more step: when you lose the race, adopt the winner instead of erroring

There is still a window between "ask first, resume if nobody" — an activation might register in exactly that moment. The robust approach is to accept defeat: if the claim is taken by that just-registered activation, deliver the message to it rather than throwing the failure at the user.

Fix two (ready-made implementation): @argszero/cordis-plugin-session-trigger

If you do not want to write this logic yourself, you can just use this plugin — it was built for exactly this gap:

ts
import { deliverToSession } from '@argszero/cordis-plugin-session-trigger'

const { path } = await deliverToSession({ agents: ctx.agents }, { sessionId, message })
// path: 'live' | 'live-after-race' | 'resumed'

Its behavior is defined clearly:

  1. Ask the registry first; in the live case it never attempts resume;
  2. Resume only when nobody holds it;
  3. Accept a lost race and adopt the winner — if the claim is taken by an activation registered at the same moment, the message goes to that Agent;
  4. When it really is held, it gives a diagnosis naming both layers (a write claim in this process / another dsh process holding session.lock), and explains what to do for each.

Mounting (a cordis.patch.yml insert, with config mode: followup|steer, retain: true|false): it registers ctx.sessionTrigger and retains every Agent it resumes — because the queued turn needs the handle alive; on unload it disposes them all, which is also the action that releases the claim.

There is one detail worth learning on the compatibility side: this package imports no @deepseek-ai/dsh-* module at runtime — it programs only against the documented ctx.agents (get, resume) and the optional ctx.agentDefaultModel (currentSelection), matching rejections by error name with the message text as a fallback. As a result, one build serves the 0.1.2 / 0.1.3 / 0.1.5 / 0.1.6 lines at once.

Its testing is worth noting too: it runs against a real @deepseek-ai/cordis context and the real SessionAlreadyOwnedError exported by the published dsh-session-persistence (not a look-alike stand-in), and it carries a control arm — the naive ctx.agents.resume() it replaces must fail on a live session under the same fixture. 15 tests, plus 6 mutations to the plugin's own logic, each confirmed to turn the suite red.

Boundary: this only keeps your plugin from getting hurt, it does not make "opening" free

This is the most important section for expectations; do not skip it.

The core-side design is: every activation takes write ownership up front, and today there is no mechanism for a plugin to make the UI release it. Therefore:

  • The plugin-side pattern only protects the plugin's own resume;
  • It cannot make "just taking a look" free;
  • Your original request — "do not take the lock when opening, only when sending a prompt" — is a core change: either make write lazy, or provide a genuinely read-only follow path.

If you are going to report this, that is the request to state explicitly: not "fix some plugin's error", but "give the UI a read-only open semantic, or make the write claim lazy".

The same discussion contains a piece of experience worth calling out separately, because it is a release-pipeline pitfall rather than a logic bug.

@argszero/cordis-plugin-session-trigger's 0.1.0 would not install:

text
Error [ERR_MODULE_NOT_FOUND]: Cannot find package '@deepseek-ai/schemastery'
imported from .../node_modules/@argszero/cordis-plugin-session-trigger/lib/index.js

The cause is classic: the build artifact lib/index.js imports @deepseek-ai/schemastery (the schema layer behind Config), but 0.1.0 declared it only in devDependencies. In the author's own repository that specifier resolves — pretest builds it, and the package is present in the workspace tree — so the test suite could not see the problem at all. The only thing that exposes it is "installing the published tarball by name in a clean project".

The interesting part: the author had verified by tarball install, and tarball install happens to make the dependency available in the tree; only switching to a by-name install from the registry exposed it. The same bytes, different resolution results.

The fix and hardening (0.1.1, npm latest, repo 936559d):

  1. Declare @deepseek-ai/schemastery in dependencies — which is also what in-repo consumers do (packages/core/tools and packages/llm/llm-deepseek both declare it under dependencies, not peer);
  2. Add a test/packaging.spec.mjs: scan the built lib/ and src/ for bare specifiers and assert each one appears in the manifest; also assert the reverse — none is declared but un-imported. This guard is not decoration: remove the dependencies block and it fails with a precise diagnosis:
text
lib/ imports @deepseek-ai/schemastery but package.json declares only @deepseek-ai/cordis —
a clean install of the published tarball will fail with ERR_MODULE_NOT_FOUND

Verification on 0.1.1: 18/18 tests (15 behavioral + 3 guard); the guard fails on the real 0.1.0 defect and passes after the fix; the packaged tarball installs and imports; the 15 behavioral tests run against the installed artifact, with imports resolved through node_modules to @argszero/cordis-plugin-session-trigger rather than the local lib/.

The final sentence is the most valuable conclusion: installing by name (npm i <pkg> in an empty directory, then import) should now be the primary check, not tarball install — nothing else can witness an undeclared dependency.

Troubleshooting notes

  1. Layer first with ctx.agents.get(id) (not undefined = a live Agent in this process; undefined = a cross-process lease), then talk about fixes.
  2. Do not judge by whether the session.lock file exists. Under POSIX the file is not deleted after a lease is released, so its presence or absence says nothing.
  3. "Just opening" is also activation. history.follow() → promote() → agent loop → write claim, the whole chain is in the background and the user feels nothing.
  4. Switching sessions away is not release. Off-screen Remote sources keep running.
  5. Watch the UI / plugin asymmetry: the Host API reuses the live Agent first (liveAgent(sessionId)); ctx.agents.resume() has no such precondition check.
  6. ctx.agents.resume() takes only one argument (resumeSessionId is required); the two-argument form belongs to AgentFactory.resume, do not mix them up.
  7. Do not just catch without handling. session/writer-held is specially handled in only one client place; elsewhere it is thrown as-is or swallowed — that is the source of "it just does not work"; when you catch it yourself, give a diagnosis that names both layers.
  8. Understand the retain semantics: the queued turn needs the handle alive, so the plugin must retain the Agent it resumes and dispose it on unload — that is when the claim is released.
  9. To be version-portable, do not import dsh-* packages; match by error name with a message fallback, or you need one build per version line.
  10. Before publishing an npm package, do a final check with "install by name in an empty directory, then import", and add a guard test that "every bare specifier in the build output is in the manifest".

Writing plugins around sessions, locks, and concurrency, the easiest thing to underestimate is "who holds what, and when" — especially when the occupation happens in the background and the user only sees a failure message. DSH Plugin Hub provides five screens: plugin market, installed list, custom install, settings, and system logs. The installed list labels each plugin's source and version and can locate its directory directly; the system log page keeps install, uninstall, and diagnostic trails by category and level and supports exporting the full text; the notification center aggregates history records and the live progress of in-flight tasks. When investigating "which plugin did what, and when", aligning the scene first saves a great deal of guesswork.

DSH Plugin Hub · Notification center: history of install, uninstall, and update records plus live progress of in-flight tasks

Source: Discussion #7156, @argszero/cordis-plugin-session-trigger (npm / source).

FAQ

I just opened a session in DSH and did nothing else — why does a plugin resume report already owned?

Because "opening" itself activates the Agent. The chain is: the client opens the live event stream → history.follow() first returns a snapshot → when the observed source is prepared it calls promote(...) → the Agent is resolved and activated in the background → activation starts the agent loop → and the agent loop's first act is to take the write handle. As long as that handle lives, it holds the write lock, so you "just looked" and the lock is already gone.

It is the same error — how do I tell whether I am blocked by an Agent in the same process or by another dsh process?

At the moment of failure, look at ctx.agents.get(id). **Not undefined** means a live Agent in the same process holds it (the tracker.claimWrite(id) layer); **undefined** means another process holds the kernel lease (the session.lock layer). Note: do not decide based on whether the session.lock file exists — under POSIX the file is **not deleted** after the lease is released, so the file's existence proves nothing; the real holder is identified by the exclusive lock on that file.

Why does the same session work fine when I operate it from the UI but blow up when a plugin resumes it?

Because the Host API's resolution path **returns an already-live Agent first** before considering anything else (liveAgent(sessionId)), so the UI path never touches the write open; ctx.agents.resume() has no such precondition check — it goes straight to the factory and performs the write open, so it inevitably collides with the existing claim. This asymmetry is the single most confusing point for plugin authors.

Exactly how many arguments does `ctx.agents.resume()` take? I have seen two forms.

One. The registry wrapper is ctx.agents.resume({ resumeSessionId: id, agentOptions?, setup? }) (packages/core/agent/src/index.ts:407), and resumeSessionId is the only required field. The one taking (ownerCtx, options) is AgentFactory.resume, which a plugin never calls directly. This was specifically corrected in the discussion — the two-argument form does not compile.

Is there any way to make a "read-only open" genuinely not take the lock?

Not today. By the core-side design, **every activation takes write ownership up front**, and there is no mechanism for a plugin to make the UI release it. So the request "do not take the lock when opening, only when sending a prompt" is a **core change** (either make write lazy, or provide a genuinely read-only follow path) — not something a plugin can solve on its own. All the plugin side can do is "not become a victim of this behavior".

Related Terms

SessionAlreadyOwnedError / session/writer-held
The error thrown when a session already has a writer. The in-process `tracker.claimWrite(id)` and the cross-process `session.lock` lease throw the **same** exception type with the same message (`packages/session/session-persistence/src/errors.ts:34`), and the API path normalizes it into the stable error code `session/writer-held` (`packages/api/session-controller/src/agent.ts:218-219`). Because the two layers share one code, you must find another criterion to tell them apart.— https://github.com/deepseek-ai/deepseek-harness/discussions/7156
promote (background activation after the snapshot)
What `history.follow()` calls after returning the snapshot, for sessions whose observed source is `prepared` (`packages/api/session-controller/src/history.ts:203-206`); it resolves/activates the Agent in the background. This is the direct source of "just opening also takes the lock": the moment something is displayed, the write claim is already on its way.— https://github.com/deepseek-ai/deepseek-harness/discussions/7156
write-first
The agent loop's deliberate ordering: take the write handle **before** reading or repairing anything, to exclude a concurrent resume of the same id (the comment at `packages/core/agent-loop/src/index.ts:891-895` states this outright). It contrasts with `open()`'s read branch — the read branch never touches ownership (`session-persistence-jsonl/src/index.ts:344-364`), while the write branch claims in its first step (`:365`).— https://github.com/deepseek-ai/deepseek-harness/discussions/7156
adopt a losing race
A robustness technique on the plugin side: if, after "ask the registry first, resume only if nobody holds it", you still lose the claim to an activation that registered **at the same moment**, do not error out — deliver the message to **that** Agent instead. It turns "a guaranteed failure" into "at worst one extra step", and is a key part of the resume-or-deliver pattern.— https://github.com/deepseek-ai/deepseek-harness/discussions/7156

Sources