dsh-vision-router: A Vision Routing Plugin for DeepSeek Harness, Giving Text-Only Agents Eyes
ysr666/dsh-vision-router
Eyes for text-only DeepSeek Harness agents: built-in free vision chain and pixel-level vision tools (Q&A, grounding, crop, pixel diff, OCR, SVG trace, cutout, screenshots). One-command install, no Python, image turns work like ordinary tool-calling turns.
dsh-vision-router is a vision routing plugin for DeepSeek Harness that gives text-only agents eyes. It solves the problem of DeepSeek not being able to directly see images by turning image inputs into ordinary tool calls, with a built-in free vision chain (no key) and 14 pixel-level vision tools for Q&A, grounding, cropping, pixel diff, colors, OCR, SVG trace, cutout, and screenshots. It requires no Python environment, wires the composition patch automatically, and lets you use image turns like ordinary tool-calling turns by enabling the "👁 Vision" control in the composer.
How to Install
dsh plugin --profile web add dsh-vision-router- Category
- Tools & Capabilities
- Platform
- DSH Plugin
- Author
- ysr666
- Distribution
- Plugin
dsh-vision-router Key Features
dsh-vision-router Repo Summary
What Does It Do?
dsh-vision-router is a DSH plugin for DeepSeek Harness that gives text-only agents eyes. It solves the problem of DeepSeek not being able to directly see images by turning image inputs into ordinary tool calls, with a built-in free vision chain (no key) and pixel-level vision tools for Q&A, grounding, cropping, pixel diff, colors, OCR, SVG trace, cutout, and screenshots. Maintained by ysr666 under the MIT license, it was last updated in 2026-08 and requires no Python environment.
Core Features
- Pixel-level vision tools: provides 14 deep vision tools including visual Q&A, grounding, cropping, pixel diff, color extraction, OCR, SVG tracing, cutout, and HTML screenshots, supporting continuous multi-step image work.
- Free built-in vision chain: defaults to a five-model OVHcloud anonymous fallback with no account or key; user-provided vision models run first.
- Routing bridge instead of description bridge: hands images directly to the vision model for pixel fidelity, keeping DeepSeek as the reasoning brain and leaving text turns untouched.
- Transparent and cached: uploaded images render normally in the conversation UI, the rewrite happens only inside the model call, and answers are cached by image content with JSON mode support.
- Capability-aware auto routing: since v2.0.0, supports capability-aware auto routing and benchmarks, with an explicit "👁 Vision" control.
How to Use This Plugin?
After enabling it in DSH, when you need to process an image, turn on the "👁 Vision" control in the composer and use image turns like ordinary tool-calling turns. The plugin automatically routes the image to the vision model and executes the appropriate vision tools (such as vision_ground, vision_crop, vision_describe, vision_pixel_diff, etc.), allowing the agent to iterate until the task is done. No manual file edits are needed; the composition patch is wired automatically during installation.
This page is an independent rewrite of the plugin's official README — for authoritative documentation and the latest changes, refer to the source: ysr666/dsh-vision-router. The plugin is third-party code that runs on your machine once installed; inclusion does not imply endorsement — please review the source before installing.
