What Is ModLens? Text-Only Model Vision for DeepSeek Harness
ModLens is a vision plugin for DeepSeek Harness (DSH) that solves the problem of text-only models like DeepSeek and GLM being unable to read images: paste an image straight into the chat and get structured JSON evidence with OCR, layout, and semantics, so the model quotes specifics instead of imagining. This article covers what ModLens is, its core features, the full install/update/uninstall commands, typical usage, and common troubleshooting, so you can get vision in DSH Plugin with zero configuration.
What Is modlens?
ModLens is a plug-in vision engine that gives a text-only model sight, letting it read images pasted straight into the chat. ModLens is maintained by liustack under the MIT license and positioned as the most capable vision plugin in the DSH ecosystem (source). DeepSeek's flagship chat models and GLM-5.3 itself are text-only and cannot read images; GLM-5.3-Flash is the native multimodal one. The old workaround was saving an image to a file and passing a path, which a text-only model still cannot see. ModLens reduces this to paste-and-read: the image enters the conversation and the native modlens_read_image tool takes over, returning full transcription, reading-order layout regions, and entity and relation lists. Install once and reuse across Claude Code, Codex, Pi, and OpenCode, as part of the DSH Plugin ecosystem.
What Are the Core Features of modlens?
ModLens's core features are paste-to-read vision, auto-discovery and wrapping of text-only model routes, zero-config startup, multi-key rotation with failover, and evidence-based structured output. All capabilities come from the official README (source):
- Paste-to-read: paste an image directly into the chat with no saving to a file first; after one paste, follow-up questions about the same image need no second paste.
- Auto-discovery of model routes: every provider route carrying eligible text-only DeepSeek, GLM, or MiMo Pro models is auto-discovered and wrapped; a stock install gets
DeepSeek-V4-Flash (modlens vision)andDeepSeek-V4-Pro (modlens vision), while native vision models like GLM-5.3-Flash are excluded. - Zero-config startup: reuses existing logins in Claude Code, Codex, OpenCode, and Pi; with nothing installed, the free Antigravity CLI needs no key, and a free Gemini key brings each read down to 5-10 seconds.
- Multi-key rotation and failover: comma-separated API keys rotate on auth, rate-limit, or quota failures; network, 5xx, and parse failures skip remaining keys and keep provider failover.
- Evidence, not imagination: full transcription, reading-order layout regions, and entity and relation lists ground the model in specifics;
meta.attemptsrecords every attempt so a fallback is never silent.
How to Install and Enable modlens?
Installing, updating, and uninstalling ModLens in DeepSeek Harness are each a single dsh plugin command; once enabled, just paste an image to start. The official install command pins the explicit version @liustack/modlens@3.25.2 to avoid pnpm 11 holding back a fresh release and resolving to an older version (source):
1. Install ModLens: run the install command in the DeepSeek Harness terminal to add the plugin to the web profile:
npx -y @deepseek-ai/dsh plugin --profile web add @liustack/modlens@3.25.2
Wait for the command to report success; if you use the DSH plugin market, search "modlens" in Settings → Plugins and install with one click.
2. Enable ModLens: DSH auto-loads the plugin on startup, so no extra command is needed. Confirm the ModLens card appears in Settings → Plugins — that means it is enabled, with zero config required.
3. Update ModLens: run the update command to bring the installed version to the latest:
dsh plugin --profile web update @liustack/modlens
Re-running the install command is equivalent and also fetches the latest version.
4. Uninstall ModLens: run the uninstall command to remove the plugin from this environment:
dsh plugin remove @liustack/modlens
On skill-style harnesses like Claude Code or Codex, uninstalling is deleting the skill folder and your agents are back to stock.
Typical modlens Usage
Typical ModLens usage is pasting an image straight into the chat, or picking a (modlens vision) entry in the model selector first; power users plug in any OpenAI-compatible vision model with modlens config set. (source)
1. Paste directly: the image lands as a private temp file with its path in the composer and modlens_read_image takes it from there, converting it to structured evidence — no command needed.
2. Pick a wrapped entry: select DeepSeek-V4-Flash (modlens vision) in the model selector (your choice is remembered, once is enough), then paste; the thumbnail stays in the message and is converted at request time, answered by the same underlying route — no command needed.
3. Run a health check: run the health-check command to see which local agent CLIs can be reused and whether the engine pool is ready:
modlens doctor
4. Set a vision engine: state a preferred engine (such as the fast, free gemini-api) with the command:
modlens config set provider gemini-api
To drive an OpenAI-compatible endpoint (such as qwen-vl), run these three config commands in order:
modlens config set openai.baseUrl https://dashscope.aliyuncs.com/compatible-mode/v1
modlens config set openai.apiKey <key>
modlens config set openai.model qwen3-vl-plus
5. Multi-key and proxy: openai.apiKey accepts a comma-separated list for auto rotation; behind a proxy, run the command so API providers route through it:
modlens config set proxy <url>
modlens Troubleshooting
The three most common ModLens problems are the model not reading the image, reads failing or timing out, and requests failing behind a proxy; fix them by picking a vision entry, running modlens doctor, and configuring the proxy respectively. The project also ships a dedicated troubleshooting doc (source):
1. The model does not read a pasted image: the symptom is the model answering off-topic or ignoring the image; the cause is that the current model is not taken over by ModLens — it only wraps routes whose metadata positively confirms text-only. Fix: pick a (modlens vision) entry in the model selector, or switch to a text-only DeepSeek or GLM model.
2. A read fails or hangs: the symptom is modlens_read_image erroring or spinning; the cause is an unconfigured engine or exhausted key quota. Fix: first run the health-check command to inspect engine status:
modlens doctor
Then configure a free Gemini key to bring reads to 5-10 seconds, and separate multiple keys with commas so ModLens rotates on auth, rate-limit, or quota failures.
3. Install or reads fail behind a proxy: the symptom is network requests timing out or being rejected; the cause is that requests are not going through the system proxy. Fix: set the HTTPS_PROXY environment variable, or run the command so API providers route through it:
modlens config set proxy <url>
Use Cases and Notes
ModLens fits any DeepSeek Harness session that runs text-only models but still needs to read images; note that it only takes over text-only models, requires Node.js, and does not accept pull requests. (source)
Good use cases: text-only models like DeepSeek-V4 reading screenshots, reading data charts (the README shows a 128-model scatter plot read axis by axis), analyzing several images pasted at once one by one, and reading an author, caption, and engagement numbers from a tweet screenshot. Notes:
- ModLens only wraps models whose metadata positively confirms text-only; native vision models are auto-excluded and unconfirmed models keep their native paste.
- It is built with TypeScript and requires a Node.js environment.
- The project does not accept pull requests — the author reviews every line; contribute by opening an issue or forking for your own use under MIT.
- Upstream engines (Antigravity CLI, Gemini/OpenAI/Anthropic APIs, and any compatible endpoint) are governed by their own terms and quotas, which you are responsible for.
Project Links
ModLens is an MIT open-source project maintained by liustack. Plugin detail page: ModLens plugin detail.
This page is an independent guide rewritten from the plugin's official README — for the authoritative documentation and the latest changes, defer to the source: liustack/modlens. A plugin is third-party code that runs on your machine once installed; inclusion is not an endorsement — review the source before installing.
FAQ
ModLens, a DeepSeek Harness (DSH) vision plugin, is not limited to DeepSeek: it auto-wraps eligible text-only DeepSeek, GLM, or MiMo Pro model routes, while native vision models like GLM-5.3-Flash are excluded automatically. It only takes over models whose metadata positively confirms text-only; the rest keep their native paste.
Just paste an image in DeepSeek Harness: it lands as a private temp file in the composer and the modlens_read_image tool takes it from there. Or pick a (modlens vision) entry such as DeepSeek-V4-Flash (modlens vision) in the model selector, and the thumbnail stays visible, converted to structured evidence at request time.
ModLens, a DeepSeek Harness (DSH) vision plugin, starts with zero config: existing Claude Code, Codex, OpenCode, or Pi logins are reused, and free channels include the Antigravity CLI (no key) and a Gemini key. It works with nothing configured, and adding a key brings each read down to 5-10 seconds.
ModLens in DeepSeek Harness (DSH) accepts comma-separated API keys and rotates to the next one on authentication, rate-limit, or quota failures; network, 5xx, and parse failures skip remaining keys and keep the existing provider failover. meta.attempts records every attempt, so a fallback is never silent.
No. ModLens is a DeepSeek Harness (DSH) plugin — exactly one skill folder on skill harnesses and one plugin on DSH, with no hooks, no wrappers, no local proxy daemon, and no changes to any harness config. Uninstalling is deleting a folder, and your agents are back to stock.
ModLens's openai provider is a universal socket: run modlens config set openai.baseUrl <endpoint>, modlens config set openai.apiKey <key>, and modlens config set openai.model <model>. Any compatible endpoint like qwen-vl, GLM, or a self-hosted vLLM/Ollama plugs straight in.
Related Terms
- ModLens
- ModLens is a vision plugin for DeepSeek Harness (DSH) that gives text-only models sight by converting pasted images into structured JSON evidence.— ModLens README
- (modlens vision) entry
- A (modlens vision) entry is a wrapped model-selector entry ModLens appends to each eligible text-only model route; picking one keeps the pasted thumbnail visible and converts it to structured evidence at request time.— ModLens README
- Structured JSON evidence
- Structured JSON evidence is ModLens's output for a pasted image: full transcription, reading-order layout regions, and entity and relation lists that let the model quote specifics.— ModLens README
- modlens_read_image
- modlens_read_image is the native tool provided by ModLens that reads an image pasted into the chat and converts it into structured JSON evidence.— ModLens README
- Vision engine
- A vision engine is ModLens's backend source for reading images, chosen from six built-in providers (gemini-api, openai, anthropic, antigravity-cli, claude-cli, kimi-cli) plus reusable local CLIs, forming one failover chain.— ModLens README