DSH plugin: an optional component failure breaks everything
Two apparently unrelated failures share one root sentence: the failure of an optional component gets promoted into the failure of the entire tree. The first is dsh web, which prints its Web URL and then exits a few seconds later leaving a single line: dsh: fatal load failure: Error: EPERM: operation not permitted, realpath 'D:\codes\skills\audio-transcribe\.pytest_cache' — you assume "cannot read that directory" is what is fatal, but the provider caught that error long ago; what is actually fatal is a Promise inside chokidar that nobody catches escaping as an unhandledRejection after the watcher's close() (#7537). The second is that no new session can be created in the Web UI at all and session/create always returns 500 with mcp-client(playwright-mcp): initial connection or tool synchronization failed — an experimental, optional browser tool failing to start drags the whole session-creation flow down with it (#7865). Both share the same fix direction: confine "a tool is unavailable" to that tool's own scope, and degrade outward instead of dying. Below we go "triage first → mechanism one → mechanism two → fixes → notes", and every step gives a command you can paste straight into a terminal.
Triage first: the two symptoms look nothing alike, so do not rush to delete a directory
Look at where the symptom lands first: if the crash happens during startup and the terminal exits directly, take the first path; if the crash happens during Web interaction and the terminal is still alive, take the second.
| Criterion | Mechanism one: skills directory EPERM crash (#7537) | Mechanism two: MCP taking session creation down (#7865) |
|---|---|---|
| When it triggers | dsh web startup scan phase | Creating/opening a session at runtime |
| Terminal behavior | Prints URL, then fatal load failure, exit 1 | Host process stays alive, Web page shows an error |
| Key log | EPERM ... realpath '...\.pytest_cache' | mcp-client(playwright-mcp): initial connection or tool synchronization failed |
| Protocol-level behavior | None (the process already exited) | POST /api/session/create → 500 |
| Leftover processes | None | 12 orphan Edge processes (~1.5 GB) |
| One-line triage command | icacls <dir> / ls -ld <dir> | curl -s -X POST .../session/create -d '…' and check the status code |
Step one, confirm the first mechanism: does the tree pointed to by customSkillDirs contain at least two directories the current user cannot read? Note the original evidence in the report — the directory itself exists, but the ACL is unreadable:
Resolve-Path 'D:\codes\skills\audio-transcribe\.pytest_cache' # returns normally
Get-Item -Force 'D:\codes\skills\audio-transcribe\.pytest_cache' # returns normally
icacls 'D:\codes\skills\audio-transcribe\.pytest_cache'
# D:\codes\skills\audio-transcribe\.pytest_cache: Access is denied.
# Successfully processed 0 files; Failed processing 1 files
Step two, confirm the second mechanism: you do not need the UI, just look at the return code and timing of session/create. Failing after roughly 500 ms, and failing every time afterwards, is the typical fingerprint of this one:
curl -s -o /dev/null -w '%{http_code} %{time_total}s\n' \
-X POST http://127.0.0.1:3080/api/session/create \
-H 'content-type: application/json' \
-d '{"type":"client-request","rpcId":"'"$(uuidgen)"'","method":"session/create","payload":{"args":{"request":{"workspaceId":"<id>"}}}}'
# 500 0.541s
# 500 0.513s
# 500 0.500s ← stable reproduction, stable timing — not like a handshake timeout
Step three, run a controlled experiment that immediately distinguishes "the environment is broken" from "the state inside the process is broken". The most convincing step in the second report is running an equivalent MCP mount offline, on the same machine, with the same config, the same @playwright/mcp@0.0.80, the same executablePath, and the same cwd:
SUCCESS after 483 ms; registered tools: 24
The config is fine, the environment is fine, the executable is fine — the problem is inside the long-running Host process. That one comparison directly decides the fix direction afterwards: not editing YAML, but limiting the scope of the failure.
Mechanism one: the error is caught by the provider, then escapes from the "event channel" to the process level
The conclusion of this section: the fatal point is not the realpath failure, but that after the watcher is close()d, a fire-and-forget sibling-scan Promise is still flying with that same error object.
1. Look at the chokidar options block first: the provider does not pass ignorePermissionErrors
In packages/skill/skill-filesystem/src/index.ts:492-505, the options passed to chokidar.watch are:
{
persistent: /* … */,
ignoreInitial: /* … */,
depth: 1,
followSymlinks: /* … */,
atomic: /* … */,
awaitWriteFinish: /* … */,
usePolling: /* … */,
interval: /* … */,
}
There is no ignorePermissionErrors. So chokidar's EPERM on a watched directory takes the default path and re-emit('error')s. The reporter's half of the diagnosis is correct.
2. The full escape chain (6 steps)
Stitch it together in source order, each step marked with its location so you can check it against your own build:
- At
chokidar/handler.js:591,await fsrealpath(path)rejects with anEPERMerror. _addToNodeFshands the error to_handleError, which re-emit('error')s it (chokidar/index.js:552-562). OnlyENOENT/ENOTDIR, orEPERM/EACCESwithignorePermissionErrorsenabled, get suppressed — this case satisfies neither condition.- The provider receives the first error before
ready, so it doesreadiness.reject(error)(skill-filesystem/src/index.ts:520-525). - The outer catch (same file,
:536-540) closes the watcher and rethrows. close()internally callsremoveAllListeners()(chokidar/index.js:419) — after this moment, anyemit('error')on that same tree has no listener at all.- On another concurrent branch,
this._addToNodeFs(path, initialAdd, wh, depth + 1)(handler.js:488) is fire-and-forget: noawait, no.catch(). Its rejection lands after the listener has been removed, so Node can only rethrow it out ofemit, promoting it to anunhandledRejectionthat the launcher'sinstallFailLoudhandles uniformly and prints:
dsh: fatal load failure: Error: EPERM: operation not permitted, realpath 'D:\codes\skills\audio-transcribe\.pytest_cache'
at async realpath (node:internal/fs/promises:1177:10)
at async NodeFsHandler._addToNodeFs (.../chokidar/handler.js:591:45)
[ELIFECYCLE] Command failed with exit code 1.
Note this is only two frames — because app-boot/src/index.ts:660 prints ${err.stack}, and the error object is rethrown verbatim through the emit round-trip, the stack gained no context. That is also the source of the illusion that "realpath crashed directly".
3. A 6-arm measurement: quantify which step is actually the fatal one
This is the most valuable data set. Same code, only one variable changed, counting emitted errors, _handleError calls, and escaped unhandledRejections:
| Arm | emitted error | _handleError calls | unhandledRejection |
|---|---|---|---|
| provider as-is | 1 | 16 | 15 |
| watcher not closed | 16 | 16 | 0 |
ignorePermissionErrors: true | 0 | 0 | 0 |
| re-attach a swallow-all listener after close | 1 | 16 | 0 |
add .catch to handler.js:488 after close | 1 | 16 | 0 |
add .catch to the add() chain after close | 1 | 16 | 15 |
Three key readings:
- Watcher not closed → 0 escapes. So the fatal thing is not "the error itself", but the act of "close removing the listener".
.catchonhandler.js:488→ 0 escapes, whereas.catchon theadd()chain does nothing (still 15). So the escape point is on_addToNodeFs's internal call, not on the Promise returned by the outeradd()— you cannot catch it from the outside layer.ignorePermissionErrors: true→ all three columns 0. It produces no error event from the very top, so even the timing of close is irrelevant; see the cost below.
4. Race boundary: a single unreadable entry does not reproduce at all
Many people assume "any unreadable directory means a guaranteed crash"; measured, it is not. 10 runs per group:
| Unreadable entries | Reproduced in 10 | Watcher not closed |
|---|---|---|
| 1 | 0 / 10 | 0 / 10 |
| 2 | 8–10 / 10 | 0 / 10 |
| 16 | 10 / 10 | 0 / 10 |
This is the explanation for "I deleted the directory, switched skills, and it crashed again": the directory is not magic; you need a concurrent second rejection to hit the window. The window is not millisecond-scale either — the close path runs microtasks, while the sibling rejection comes from a later turn of the fs threadpool, and their ordering depends on threadpool scheduling.
5. Why this path gets reached at all: the browser is the trigger, not the root cause
FileSystemSkillProvider.list() is reached only through ctx.skills.list(), and there are exactly two production call sites:
packages/client/ui-skill/src/client/index.ts:121— warm-up at composer scope birth (packages/client/ui-input-trigger/src/types.ts:202);packages/skill/tool-skill/src/index.ts:134— the tool-side listing.
That is, it is the action of "opening/switching the input box scope" that triggers the provider's live listing, which in turn makes the watcher scan that tree. The browser is merely the person flipping the switch.
6. A note on why macOS cannot reproduce it
On macOS, chmod 000 a directory and readdir will fail, but realpath still resolves, and readdirp absorbs this class of error before entering _handleError — measured 0 _handleError calls, 0 emits. So this path is produced by the combination of Windows ACL semantics (a directory like .pytest_cache, left behind by other tools, where even reading its ACL as an ordinary user is denied). When you hit it, do not suspect "my system is broken".
Mechanism two: an optional tool's startup failure is upgraded to session-level fatality by failOnStartupError
The conclusion of this section: the combination of failOnStartupError: true and reconnect: { enabled: false } turns "one optional MCP cannot connect" into "this Host can never create a session again".
1. The mount happens inside the agent/created hook, and it is awaited
The final shape of packages/experimental/browser-use-runtime/src/mcp.ts (line numbers are for the finalized version):
ctx.on('agent/created', async ({ agent, signal }) => {
// :177 register the hook, { prepend: true }
const resources = /* … */
await resources.get(agent, signal) // :188
// open() // :129
await scope.ctx.plugin(McpClient, McpClient.Config({ // :144
transport: 'stdio',
serverName: options.name,
command: options.command,
args: options.args,
failOnStartupError: true, // :152 ← the only switch
reconnect: { enabled: false }, // :152
}))
}, { prepend: true })
There are two key points. First, this plugin() call is awaited, so its throw bubbles all the way up through the hook into session/create, becoming a 500. Second, the same function already has a graceful-degradation branch at :183-187 — when the resource is occupied it merely marks that tool as blocked and then returns without throwing. That is, the same file already has a precedent for "failure only degrades"; the MCP startup-failure path simply does not take it. This is the most natural landing spot for the fix (see fix two below).
2. The throw site and its cause
packages/mcp/mcp-client/src/index.ts:201 (corresponding to lib/index.js:832 in the published build):
const outcome = await connection.ready
if (outcome.error !== void 0 && config.failOnStartupError) {
throw new Error(
`mcp-client(${config.serverName}): initial connection or tool synchronization failed`,
{ cause: outcome.error },
)
}
It already carries a cause, it is just that neither the RPC error text nor the Host log shows the cause, so from the user side you only see that entirely uninformative sentence. This is where the report's second request comes from.
3. A/B: within a single Host process, the tool list flips between 24 and 0
Count "how many browser tools this session actually registered" by counting mcp__playwright-mcp__* in request headers:
| Host PID | First Agent | New sessions afterwards |
|---|---|---|
| 2464 | 24 | all fail |
| 24532 | 0 | 24 |
| 26644 | 24 | 24 / 0 |
The value flips between 0 and 24, consistent with the semantics of an "exclusive mount": this tool is allowed to be mounted only once per Host process, and whoever grabs it first wins. This explains "after a clean boot it works at first, then everything crashes as you keep using it" — the first mount occupies the exclusive slot, and subsequent mount failures punch straight through session creation.
4. rc.2 explicitly does not change the behavior (this matters — do not upgrade for nothing)
The reporter did two things to rule out "has it been fixed already":
git log --oneline dsh-v0.1.7-rc.1..HEAD -- packages/mcp/ packages/experimental/browser-use-runtime/— only two commits:787b746b80 release(dsh): 0.1.7-rc.2ande7def469e1 feat(i18n): … (#5036), and neither touches this path (162 commits from alpha.1→alpha.2, 346 from rc.1→rc.2, 508 total, untouched).npm pack, then a per-file hash comparison — inside artifacts such aslib/types/mcp.jsit is stillfailOnStartupError: true+reconnect: { enabled: false }, byte-for-byte identical to rc.1.
Conclusion: upgrading to rc.2 will not fix this. The only effective on-site stopgap is to change the config.
5. Resource leakage is the amplifier that makes it "worse the more you use it"
After the crash, 12 orphan Edge processes (~1.5 GB) remain, with a parent chain pointing at a browser tree no longer associated with any agent; they must be cleaned up manually:
# Windows: kill the whole process tree
taskkill /T /F /PID <orphan-edge-pid>
# macOS / Linux: clear hung MCP servers and isolated browsers
pkill -f playwright-mcp
Furthermore, each session spins up its own playwright-mcp process plus an isolated headless Edge, and some of those processes stay resident after the session ends. These leftovers keep holding memory and browser profiles, which raises the probability of the next MCP startup failure — producing the loop of "worse the more you use it, fine again after a restart".
6. One unproven candidate root cause
On the same machine, the host-installed @deepseek-ai/dsh-scope is 0.1.7-alpha.1, while the @deepseek-ai/dsh-scope resolved by browser-use-runtime inside the profile is 0.1.0-rc.8 (measured via the junction target), and dsh-mcp-client is another physical copy as well; the profile's pnpm-lock.yaml contains both a ^0.1.0-rc.8 and a ^0.1.7-alpha.1 requirement for dsh-scope. Both versions of dsh-scope use const kScope = Symbol("dsh.scope") (one per module instance, not Symbol.for), so the scope keys of the two instances are not equal. This is consistent with the phenomenon described in #4573, but the reporter explicitly could not reproduce it reliably, so it is recorded only as a candidate, not a conclusion.
Fix one (mechanism one): three layers, from the narrowest to the least effort
Priority suggestion: do 3 as a stopgap first, then pick 1 or 2 depending on the granularity you can accept.
Option A: least effort — one line on the provider side, coarsest cost
Add one switch to the options block of chokidar.watch:
chokidar.watch(roots, {
persistent: /* … */,
ignoreInitial: /* … */,
depth: 1,
followSymlinks: /* … */,
atomic: /* … */,
awaitWriteFinish: /* … */,
usePolling: /* … */,
interval: /* … */,
ignorePermissionErrors: true, // ← added
})
Measured all three columns 0 (0 emits / 0 _handleError / 0 rejections). Cost: an unreadable root directory is also silenced, the provider returns complete: true with an empty candidate list. Per the contract, only complete: false makes the layer above fall back to cache and log skill provider "filesystem" skipped. In other words, this option makes "the whole directory cannot be read" look like "the directory simply has no skills".
Option B: the narrowest — keep a swallow-only listener after close
In closeWatcher at skill-filesystem/src/index.ts:594-601, do not discard the watcher directly; first attach a listener that only swallows errors and does nothing else:
async function closeWatcher(handle: WatcherHandle): Promise<void> {
const { watcher } = handle
// Keep a swallow-only listener so sibling scans that reject after
// removeAllListeners() cannot escape as an unhandled rejection.
watcher.on('error', () => {})
await watcher.close()
}
Measured: 1 emitted / 16 _handleError / 0 unhandledRejection. The first error still reaches handleWatcherError, the complete: false semantics are preserved, and observability is not lost. This is the fix that best matches the fault's scope.
Option C: two upstream spots, either one alone is sufficient
Two changes in chokidar 5, each of which plugs the escape:
// 1) index.js: do not emit when there is no listener, avoiding Node's rethrow out of emit
if (this.listenerCount(EV.ERROR) === 0) return
this.emit(EV.ERROR, error)
// 2) handler.js:488: attach a catch to the fire-and-forget recursive scan
this._addToNodeFs(path, initialAdd, wh, depth + 1).catch((err) => {
this._handleError(err)
})
Note that the sixth arm above already showed: a .catch on the outer add() chain does not help — it must be attached at the _addToNodeFs level (or the equivalent handler.js:488).
On-site stopgap (no code change): watch: false
If you just need it running right now, you can turn off hot reload in the same config block:
- id: skill-filesystem
name: '@deepseek-ai/dsh-skill-filesystem'
config:
includeDefaultRoots: false
customSkillDirs:
- D:/codes/skills
watch: false # ← no watcher is started (retainRoot / ensureWatcher)
The cost is losing skill hot reload: after editing a SKILL.md you must restart for it to take effect. Also do not forget the precondition — at least two unreadable entries must exist to crash, so "it stops crashing after I delete the only .pytest_cache" follows the rule rather than being luck.
Fix two (mechanism two): confine the failure to that tool's own scope
Core idea: session/create should not return 500 just because an optional tool cannot mount; the correct behavior is "this tool is unavailable in this session" plus a clear notice.
Option A: a one-line stopgap on the user side (immediate recovery)
- id: 1afc9ad1
name: '@deepseek-ai/dsh-experimental-browser-use-playwright-mcp'
config:
mode: launch
headless: true
executablePath: <EDGE>\msedge.exe
failOnStartupError: false # ← change from true to false
New sessions recover immediately after the change. But be clear: this is only a stopgap; the browser tool itself is still unavailable (see FAQ item 4 above). If you also want it to self-heal, turn reconnect on too:
reconnect: { enabled: true }
Option B: upstream — close it on the existing degradation branch
The same file's :183-187 is already this shape: when the resource is occupied it only marks blocked and then returns. Just route an MCP mount failure onto the same path:
try {
await scope.ctx.plugin(McpClient, McpClient.Config({
transport: 'stdio',
serverName: options.name,
command: options.command,
args: options.args,
failOnStartupError: false, // let the catch below decide on failure
reconnect: { enabled: true }, // allow later self-healing
}))
} catch (error) {
// A failed optional tool must not fail session creation.
markToolUnavailable(agent, options.name, error)
return
}
Three points: ① failures only degrade, never throw; ② degradation must be visible, not silent; ③ turn reconnect on, otherwise this Host process stays crippled.
Option C: include cause in the error text
mcp-client already builds { cause: outcome.error }, it just is not shown. Change the Host log and RPC error text to expand cause (including the MCP child process's stderr), so the user side can tell whether it is ENOENT, permissions, or the handshake:
throw new Error(
`mcp-client(${config.serverName}): initial connection or tool synchronization failed: ` +
`${formatCause(outcome.error)}`,
{ cause: outcome.error },
)
Option D: resource reclamation
When a session is released or startup fails, make sure the playwright-mcp process and the --isolated headless browser process tree are cleaned up. This complements the stopgap above: no cleanup → leftover processes raise the chance of the next failure → making the collateral more likely to recur.
Troubleshooting notes and summary
Troubleshooting notes
- First tell apart "crash at startup" and "crash at runtime". The former shows
dsh: fatal load failure, the latter shows an RPC status code plus a page error — two root causes, two sets of fixes; do not apply one to the other. - Seeing
realpathand assuming it is a permissions problem is a misdiagnosis. Here the provider already caught the error; what is fatal is the uncaught rejection. The tell is that the stack has only two frames and the error is printed by the launcher'sinstallFailLoud(anunhandledRejection). - A single entry not reproducing is normal. You need ≥2 unreadable entries, and the watcher must
close()mid-scan. So "it seems fine after I deleted one directory" does not mean it is fixed. watch: falseis a workaround, not a fix: it removes the watcher, and with it hot reload. Good for an on-site stopgap.- Do not treat
ignorePermissionErrors: trueas a free switch. It also silences an unreadable root,completeflips fromfalsetotrue, and the layer above no longer falls back to cache. - For mechanism two, do the offline equivalent-mount comparison first (same config, same
@playwright/mcp, sameexecutablePath, samecwd). Measured 483 ms success and 24 registered tools, which says the config is fine — do not waste time on the YAML. - A
session/create500 is a symptom, not the root cause. The sentence inside the 500 has nocause; you need to go to the Host log or instrument it yourself to get the underlying error. - Before upgrading, confirm the version line actually contains the fix. Mechanism two is explicitly unfixed in rc.2 (git log shows only two commits, and
npm packis hash-identical per file), so do not expect an upgrade to cure it. - Watch out for the npm
latesttrap: at the time it was still stuck on0.1.5-rc.3, older than0.1.7-rc.2; usenpm i -g @deepseek-ai/dsh@next. - Clear the orphan processes before troubleshooting: the leftover headless Edges (12 this time, ~1.5 GB) keep holding memory and browser profiles, and if you do not clean them out they will knock an "already fixed" illusion right back into place.
Summary: two chains, one shared design defect
Put the two chains side by side and the shape is identical:
Mechanism one: realpath EPERM
→ provider catches it (not fatal)
→ watcher.close() removes the listener
→ the sibling scan's rejection has no one to catch it
→ unhandledRejection
→ installFailLoud → exit 1 ← failure radius = the whole process
Mechanism two: MCP connection fails (optional tool)
→ connection.ready carries back an error
→ failOnStartupError: true decides to throw
→ bubbles up through the awaited agent/created hook
→ session/create 500 ← failure radius = all session creation
In both places the failure radius is far larger than the fault source. The fix is therefore consistent too: let "this link fails" affect only itself — for mechanism one, the swallow-only listener after close (or those two upstream spots); for mechanism two, the existing blocked degradation branch plus reconnect plus an expanded cause. This is not patching a particular bug; it is explicitly modeling the common path of "an optional component is unavailable".
Sources: Discussion #7537, Discussion #7865.
Once you have triaged this class of "one optional component drags down the whole tree" failure, you will usually want to check which plugins are actually installed, at which versions, and which are broken yet still listed. DSH Plugin Hub provides five screens — plugin market, installed list, custom install, settings, and system logs — so you can filter by source (catalog plugins / custom installs), see versions and update times, and trace install, uninstall, and settings-change execution trails by level and category on the system log page. That makes it a good way to confirm "which plugin is the one dragging its feet" before deciding whether to uninstall it or change its config.

FAQ
Because printing the URL and loading plugins advance in parallel: the web server comes up first, and then skill-filesystem scans customSkillDirs, where chokidar calls realpath() on a watched directory and gets an EPERM (on Windows this is common for a .pytest_cache with a broken ACL). The provider already catches that error, so the error itself is not fatal. What is fatal is a Promise nobody catches escaping as an unhandledRejection **after** the watcher has been closed, which the launcher's installFailLoud captures and treats as a fatal error. In other words the crash point is "an unhandled rejection", not "an unreadable directory".
Because the trigger is not "a particular directory", it is "**at least two** unreadable entries in the same scan". Measured: one unreadable entry inside the watched tree reproduced 0 times out of 10; two reproduced 8–10 out of 10. With a single entry the timing of that rejection happens to be absorbed by _handleError before removeAllListeners(), so it looks fine; as soon as there is a concurrent sibling scan, the second rejection lands after the listener has been removed. In other words, fixing the directory is just luck — you need to close this at the watch configuration or in the upstream code.
The cost is that it silences an unreadable **root** directory too. That option makes chokidar emit no error event at all for EPERM/EACCES, so a root directory that cannot be read becomes "zero candidates, no error", the provider returns complete: true, and the layer above does not fall back to cache or log skill provider "filesystem" skipped. If you can accept "a whole unreadable skills directory behaves like having no skills", it is enough; if you want to keep observability, choose the narrower fix (keep a swallow-only listener after close()), which still routes the first error to handleWatcherError and keeps complete: false.
Two things got mixed together. failOnStartupError: false only changes "should a mount failure fail session creation", not "can the MCP connect" — so after switching to false the session can be created, but that session still has no browser tool. What you actually need to investigate is why the MCP cannot connect (usually accumulated mount/resource state inside the host process: measured, an equivalent offline mount with the same config succeeds in 483 ms and registers 24 tools, which says the config itself is fine). Note also that this combination carries reconnect: { enabled: false }, so it will not self-heal after a failure — you can only restart the Host.
As of the reports cited in this article, both are still in the discussion stage and have not been merged as fixes. For triage in the meantime: for #7537 first confirm "are there at least two unreadable entries at once" (a single entry does not reproduce), clean up anything you can delete or re-ACL, or set watch: false for that provider to give up hot reload; for #7865 the most direct stopgap is to set failOnStartupError to false in that plugin's config to restore session creation, then use taskkill /T /F (Windows) or pkill -f playwright-mcp to clear the orphan browser processes, and finally move to the @next channel (do not use npm latest, which is still stuck on the older 0.1.5-rc.3).
Related Terms
- unhandledRejection escape
- When an async function returns a Promise that is neither `await`ed nor given a `.catch()`, nothing catches its rejection, and Node promotes it to a process-level `unhandledRejection`. Here chokidar's `this._addToNodeFs(path, initialAdd, wh, depth + 1)` is exactly this kind of fire-and-forget call; when it rejects later than the watcher's `close()` (and `close()` calls `removeAllListeners()`), there is nowhere for even `emit('error')` to deliver, so the error "escapes" from the event channel up to the process level.— https://github.com/deepseek-ai/deepseek-harness/discussions/7537
- ignorePermissionErrors (chokidar option)
- A chokidar watch option that, when `true`, neither logs nor emits an `error` event for `EPERM`/`EACCES` (equivalent to treating those two classes as `ENOENT`). It is a "one-line stopgap" but very coarse-grained: because it silences at the directory-tree level, even an unreadable **root** produces no error, so the layer above cannot distinguish "the directory really has no skills" from "the directory cannot be read at all".— https://github.com/deepseek-ai/deepseek-harness/discussions/7537
- complete flag (skill provider contract)
- A completeness declaration a provider carries when returning a candidate list. `complete: true` means "this is the whole result", and the layer above treats the directory as genuinely empty; `complete: false` means the result may be incomplete (for example, one source was skipped), and the layer above falls back to cache and logs `skill provider "filesystem" skipped`. The key to fault isolation is making "could not read everything" show up as `false` rather than faking a complete empty set.— https://github.com/deepseek-ai/deepseek-harness/discussions/7537
- failOnStartupError (MCP client config)
- A startup-policy switch in `@deepseek-ai/dsh-mcp-client`. When `true`, as soon as `connection.ready` carries an error the plugin registration stage `throw`s; because that registration happens inside the `agent/created` hook and is `await`ed, the throw bubbles all the way up and makes `session/create` return 500. Changing it to `false` only adjusts the **scope of the failure** (degrading to "this session does not have this tool"); it does not fix the connection itself.— https://github.com/deepseek-ai/deepseek-harness/discussions/7865
Sources
- #7537 — [Bug][Windows] skill-filesystem crashes dsh web when a custom skill root contains an inaccessible directory· deepseek-ai (GitHub Discussions)
- #7865 — [bug] A failed experimental Playwright MCP makes all session creation fail (failOnStartupError: true + reconnect: false) and leaves orphan browser processes behind· deepseek-ai (GitHub Discussions)