modlens: A Vision Plugin for DeepSeek Harness, Giving Text-Only Models Sight
liustack/modlens
ModLens is a vision plugin for DeepSeek Harness, giving text-only models sight: paste an image and get structured JSON evidence (OCR, layout, semantics).
modlens is a vision plugin for DeepSeek Harness that solves the problem of text-only models like DeepSeek and GLM being unable to read images. Once installed, paste an image directly into the chat and get structured JSON evidence including OCR, layout, and semantics, so the model quotes specifics instead of imagining. It starts with zero config, auto-discovers and wraps eligible text-only model routes, and excludes native vision models automatically. Install once and reuse across Claude Code, Codex, Pi, OpenCode, and more.
How to Install
dsh plugin --profile web add @liustack/modlens- Category
- Tools & Capabilities
- Platform
- DSH Plugin
- Author
- liustack
- Distribution
- Plugin
modlens Key Features
modlens Repo Summary
What Does It Do?
ModLens is a DSH plugin for DeepSeek Harness that gives text-only models vision capabilities. It solves the problem that DeepSeek, GLM, and other text-only models cannot read images, letting you paste an image directly into the chat and get structured JSON evidence (OCR, layout, semantics). Core capabilities include auto-discovering and wrapping eligible text-only model routes, zero-config startup, multi-key rotation, and cross-platform reuse.
Core Features
- Paste-to-read: paste an image directly into the chat without saving to a file first, automatically converting it to structured JSON evidence.
- Auto-discovery of model routes: automatically detects and wraps eligible text-only DeepSeek, GLM, or MiMo Pro models, while native vision models (like GLM-5.3-Flash) are excluded automatically.
- Zero-config startup: reuses existing setups in Claude Code, Codex, OpenCode, Pi, and also supports the Antigravity CLI free channel or a Gemini key.
- Multi-key rotation: comma-separated API keys rotate on auth, rate-limit, or quota failures; other failures skip remaining keys and keep existing provider failover.
- Evidence, not imagination: provides full transcription, reading-order layout regions, entity and relation lists, so the model quotes specifics.
- Install once, use everywhere: verified on real machines in Claude Code, Codex, Pi, and OpenCode.
How to Use This Plugin?
After enabling it in DSH, there are two paste methods: first, just paste the image, which lands as a private temp file and enters the composer, then the modlens_read_image tool takes it from there; second, pick a (modlens vision) entry in the model selector (it remembers your choice, so once is enough), then paste, and the thumbnail stays visible in your message, converted to structured evidence at request time. A stock install gets DeepSeek-V4-Flash (modlens vision) and DeepSeek-V4-Pro (modlens vision) entries, while extra routes like opencode-go or zai get their own.
This page is an independent rewrite of the plugin's official README — for authoritative documentation and the latest changes, refer to the source: liustack/modlens. The plugin is third-party code that runs on your machine once installed; inclusion does not imply endorsement — please review the source before installing.
More DSH Plugin articles
View all 2 articlesWhat Is ModLens? Text-Only Model Vision for DeepSeek Harness
2026-08-30ModLens is a DeepSeek Harness vision plugin that gives text-only models sight: paste an image and get structured JSON evidence (OCR, layout, semantics).
How to Use ModLens? Give DeepSeek Harness Text-Only Models Sight
2026-08-18ModLens is a vision plugin for DeepSeek Harness: paste an image and text-only models output JSON evidence. This guide covers install and paste modes.
