What Is mnemon? LLM-Supervised Memory for AI Agents

GuidePublished 2026-08-30Author: DeepSeek Plugin Market
mnemonpersistent memoryknowledge graphguide
mnemon is a DSH plugin giving AI agents LLM-supervised cross-session memory: one binary, zero API keys. Learn installation, four-graph storage, and recall.

mnemon is an LLM-supervised persistent memory plugin in the DeepSeek Harness (DSH Plugin) ecosystem that lets AI agents remember key decisions across sessions: a single local binary with zero API keys turns memory into a four-graph knowledge store and recalls it by intent when needed. This article is an independent guide based on the official README, covering what it is, core features, installation, typical usage, and common troubleshooting.

What Is mnemon?

mnemon solves the problem that AI agents forget everything between sessions: context compaction drops critical decisions, cross-session knowledge vanishes, and long conversations push early information out of the window. The following positioning and facts come from the official README (source).

mnemon is developed by mnemon-dev in Go and open-sourced under Apache-2.0, positioned as "LLM-supervised persistent memory for AI agents." Its approach differs from most memory tools: the host LLM is the supervisor (deciding what to remember, how to relate it, and when to forget), while the binary handles deterministic computation (storage, graph indexing, search, decay). There is no embedded LLM and no middleman inference cost. The memory command path mnemon is a local binary that connects with one setup command. It also fills a gap in the protocol stack with an intent-native protocol: the remember, link, and recall primitives map command names to the LLM's cognitive vocabulary and output structured JSON instead of raw database rows. DeepSeek Harness integrates it through the dsh-mnemon plugin, combining DSH runtime memory, hosted project documents, and the Mnemon long-term memory space into one supervised three-tier memory system.

What Are the Core Features of mnemon?

mnemon's core capabilities are four-graph knowledge storage, intent-aware recall, importance decay with automatic deduplication, zero user-side operations, and multi-runtime support: the agent's memory grows like compound interest through use. These capabilities all come from the official README (source).

  • Four-graph knowledge storage: temporal, entity, causal, and semantic edges form a knowledge graph, not just vector similarity; a real example graph contains 87 insights and 2150 edges.
  • Intent-aware recall: graph traversal plus optional vector search (RRF fusion), enabled by default for all queries; recall reads the active memory space.
  • Built-in dedup and retention lifecycle: remember automatically detects duplicates and conflicts (skipping or auto-replacing them); importance decay, access-count weighting, and garbage collection manage retention.
  • LLM-supervised pattern: the host LLM makes judgments (what to remember, how to relate, when to forget), and the binary performs deterministic computation; no extra API key is needed, and a Claude subscription is the intelligence layer.
  • Intent-native protocol: the remember / link / recall primitives output structured JSON with signal transparency and no database commands.
  • Zero user-side operations: install once and it works; supported runtimes use hooks or plugins, while minimal runtimes use persistent rules; privacy-safe receipts export hashed operation credentials for memory-boundary audits.

How to Install and Enable mnemon?

Enabling mnemon in DeepSeek Harness takes two steps: install the mnemon binary on your host, then install the dsh-mnemon plugin and restart the Web profile. Installation commands come from the official README (source).

1. Install the mnemon binary: on macOS use the Homebrew Cask; on other platforms use Go install.

bash
brew install --cask mnemon-dev/tap/mnemon

2. Add the dsh-mnemon plugin: run the install command in the DSH terminal (fresh installs resolve the latest dsh-mnemon from npm).

bash
dsh plugin --profile web add dsh-mnemon

3. Start and enable: restart the Web profile, choose the storage scope under Settings, then create or activate a memory space in the Memory System tab of a session.

bash
dsh --profile web

4. Update / remove: re-run the install command to update; to remove the plugin, remove it and run mnemon setup --eject on the host side to strip all integrations.

bash
dsh plugin --profile web update dsh-mnemon
dsh plugin remove dsh-mnemon

Verify after installation: mnemon --version confirms the binary is available, and the DSH settings page confirms the plugin is enabled. Memory data lives in ~/.mnemon by default (configurable with MNEMON_DATA_DIR).

Typical mnemon Usage

Day-to-day use barely requires you to type commands — hooks nudge the agent automatically at four boundaries (session start, question, post-answer, and before compaction); use the store commands only when you need isolation or inspection. The following usage comes from the official README (source).

1. Isolate memory per project: all sessions share the default store; use named stores to isolate.

bash
mnemon store create work
mnemon store set work

2. Activate a memory space: create or activate a space in the Memory System tab of a DSH session; recall reads only the active space, and persistent writes go through a supervised subagent.

3. Check embedding status: embeddings are optional and plain graph traversal works without them; check status once configured.

bash
mnemon embed --status

4. Configure local embeddings: Ollama (localhost:11434) is the default; OpenAI-compatible endpoints also work.

bash
export MNEMON_EMBED_ENDPOINT=http://127.0.0.1:18000/v1
export MNEMON_EMBED_MODEL=bge-m3-mlx-8bit
export MNEMON_EMBED_API_KEY=sk-...

The four hook stages are Prime (making skills, guides, and the active store visible), Remind (deciding whether recall could change the task), Nudge (deciding whether a write-back is worth it), and Compact (keeping only critical continuity before compaction). They are reminders, not a hard workflow.

mnemon Troubleshooting

mnemon issues concentrate on four areas — "no memory after install", "direct GitHub-source install errors", "missing features on Windows", and "poor vector recall" — with symptoms, causes, and fixes below. The troubleshooting basis comes from the official README (source).

1. Symptom: new sessions cannot recall old memories. Cause: the mnemon binary is not installed, the storage scope is unset, or the memory space is inactive. Fix: confirm the binary is installed on the host (mnemon --version), choose the storage scope under DSH Settings, and activate a space in the session Memory System tab; recall only reads the active space.

2. Symptom: direct add github:mnemon-dev/mnemon fails to install or behaves oddly. Cause: fresh installs resolve the latest dsh-mnemon from npm, and source-path installs are unstable. Fix: switch to the npm channel.

bash
dsh plugin --profile web add dsh-mnemon

3. Symptom: Agency features are missing on Windows. Cause: Agency preview requires native Windows security support at its local authority boundary. Fix: Windows supports core Memory commands and memory features are unaffected; enable Agency once native support lands.

4. Symptom: semantic recall results are underwhelming. Cause: no embedding is configured (optional), so recall uses graph traversal only. Fix: configure an Ollama or OpenAI-compatible embedding endpoint and restart (see Typical Usage); confirm with mnemon embed --status.

Use Cases and Notes

mnemon fits AI-agent workflows that need long-term memory across frameworks and sessions, especially DeepSeek Harness users who already subscribe to Claude; zero API keys and a single binary keep deployment cost extremely low. Facts come from the official README (source).

Good use cases: multi-session projects that must remember key decisions; several runtimes such as Claude Code, Codex, Cursor, and TRAE sharing one memory (the vision is that all local agents share one active memory in ~/.mnemon); and workflows where context compaction is frequent and critical information is easily lost.

Notes and limitations:

  • Windows does not support Agency preview (core Memory commands work);
  • embeddings are optional and plain graph traversal is sufficient; configure an embedding endpoint only when hybrid retrieval is needed;
  • memory data lives in ~/.mnemon on your machine by default, a private asset that grows with you;
  • compared with memory tools that embed an LLM, mnemon uses the LLM-supervised pattern with no embedded LLM and no extra inference cost.

mnemon is an Apache-2.0 open-source project maintained by mnemon-dev, currently at version v0.2.5.

Visit the plugin detail page for full information: mnemon.

This page is an independent guide rewritten from the plugin's official README — for authoritative documentation and the latest changes, please refer to the source: mnemon-dev/mnemon. A plugin is third-party code that runs on your machine once installed; listing it here is not an endorsement — please review the source before installing.

FAQ

How do I enable the mnemon DSH plugin's cross-session memory in DeepSeek Harness?

After enabling Mnemon in DSH, run `mnemon setup` to auto-detect and configure Claude Code, or use `mnemon agency setup` to enable Agency preview for Pi agents. Once set up, memory works automatically in new sessions.

What do Mnemon's memory primitives remember, link, and recall do?

Mnemon's three primitives form an intent-native protocol: `remember` stores new knowledge, `link` connects knowledge, and `recall` retrieves relevant memories. They output structured JSON with signal transparency, letting the LLM use cognitive vocabulary instead of database commands.

Which Mnemon features are supported on Windows? Is Agency preview available?

On Windows, mnemon supports core Memory commands, but Agency preview is unavailable until its local authority boundary has native Windows security. Thus, Windows users can use memory features but not Agency's project-local responsibility and effect admission.

How does Mnemon implement importance decay and automatic deduplication?

Mnemon's binary handles deterministic computation including storage, graph indexing, search, and decay. Importance decay reduces the weight of old memories over time, and automatic deduplication removes redundant knowledge to keep the store precise. The LLM supervises when to forget.

Does Mnemon require a separate API key? How does it work with a Claude subscription?

No. Mnemon works entirely through your existing Claude subscription, no separate API key required. Your LLM subscription is the intelligence layer, and the binary only does deterministic computation, so there's no extra inference cost. Two commands and you're done.

How is Mnemon different from MCP servers in terms of memory protocol?

Mnemon differs from MCP servers in memory protocol: MCP standardizes how LLMs discover and invoke tools, but lacks a memory-semantics layer. Mnemon's `remember`, `link`, and `recall` primitives form an intent-native protocol, mapping command names to the LLM's cognitive vocabulary and outputting structured JSON instead of raw database rows.

Related Terms

mnemon
mnemon is an LLM-supervised persistent memory program for AI agents, providing cross-session memory in a single local binary. The host LLM decides what to store and when to forget, while the binary handles storage, graph indexing, search, and decay.— mnemon README
LLM-supervised pattern
The LLM-supervised pattern is mnemon's core architecture: the host LLM acts as an external supervisor for judgment, and a standalone binary performs deterministic computation. It embeds no LLM, requires no extra API key, and adds no middleman inference cost.— mnemon README
Intent-native protocol
The intent-native protocol is mnemon's memory-semantics layer built from the remember, link, and recall primitives. Command names map to the LLM's cognitive vocabulary and output structured JSON, filling the protocol gap between MCP and raw database rows.— mnemon README
Four-graph knowledge store
The four-graph knowledge store is mnemon's memory structure, organizing a knowledge graph with temporal, entity, causal, and semantic edges rather than vector similarity alone. Combined with intent-aware recall, it delivers precise retrieval.— mnemon README

Sources

View all articles