What Is dsh-vision-toolkit? Vision for Text-Only DSH Models
dsh-vision-toolkit is the first comprehensive vision-tool plugin in the DeepSeek Harness (DSH) ecosystem: it gives text-only models eyes — 10 vision tools plus a bundled vision-skills Skill for image Q&A, long-screenshot OCR, UI restoration, pixel diff, grounding, and cropping. This article explains what dsh-vision-toolkit is, its core features, the full install/update/uninstall commands, typical usage, and common troubleshooting, so you can give DeepSeek Harness vision capabilities.
What Is dsh-vision-toolkit?
dsh-vision-toolkit solves the problem that text-only models cannot understand images: it is the first comprehensive vision-tool plugin in the DSH ecosystem, letting the model paste an image and ask directly, read screenshots, restore front-end UI, and run pixel diff, grounding, and cropping tasks. The positioning and facts below all come from the official README (source).
dsh-vision-toolkit (GitHub repo Anionex/dsh-vision-toolkit) is maintained by Anionex and open sourced under the MIT license, built on the upstream agent-vision-toolkit. It is the native DeepSeek Harness integration of that visual workflow — read, ground, crop, trace, rebuild, and verify — brought into Web and Headless Profiles. In DSH Web, pasting an image automatically switches the text-only model to its (Vision Toolkit) variant with native thumbnails, session history, and workspace paths intact, and Web can preview artifacts; Headless mode invokes vision through tools such as read_image and OCR. It is an out-of-the-box vision plugin in the DSH Plugin ecosystem, with a bilingual UI.
What Are the Core Features of dsh-vision-toolkit?
The core capabilities revolve around five things: ten vision tools, the focus hint mechanism, the vision-skills Skill, transparent routing, and an isolated Python runtime — configured once, then paste an image to start. The capabilities below all come from the official README (source).
- Image Q&A and multi-image comparison: vision_glance answers questions such as "What is this?" or "Where is the error?" and can compare several images at once.
- Long-screenshot OCR: vision_long_screenshot_ocr reads very long screenshots and outputs Markdown, chunked results, an element list, and an audit.
- UI restoration: rebuild front-end UI from a screenshot into UI code or a structural description for development and testing.
- Pixel diff: vision_pixel_diff outputs the difference percentage, ranked diff regions, a heatmap, and JSON.
- Grounding and cropping: vision_ground returns raw pixel coordinates (x1,y1,x2,y2) with an optional boxed preview; vision_crop crops PNG/JPEG; vision_trace vectorizes into SVG.
- Transparent routing: pasting an image auto-switches to the (Vision Toolkit) variant with no manual model switching, keeping history and workspace paths intact.
How to Install and Enable dsh-vision-toolkit?
Install with one npm command into the DeepSeek Harness Web profile, restart it, configure a vision provider in Settings → Vision Toolkit, test the model, then paste an image or call /vision-skills. The install, update, and uninstall commands below all come from the official README (source).
1. Install — the stable release is recommended from npm into the Web profile:
dsh plugin --profile web add @anionex/dsh-vision-toolkit
Headless and Desktop use their own profiles:
dsh plugin --profile headless add @anionex/dsh-vision-toolkit
2. Enable — restart the Web profile, open Settings → Vision Toolkit, configure a vision provider (for example Gemini 3.7 Flash), and click Test vision model. The first launch prepares an isolated Python runtime: system Python 3.11+ is preferred; otherwise a standalone ~35 MB Python is downloaded and Pillow, NumPy, and vtracer are installed from a mirror.
3. Update — re-run the same add command pinned to the latest release (update to latest):
# update to latest
dsh plugin --profile web add @anionex/dsh-vision-toolkit@latest
4. Uninstall — remove it from the Web profile:
dsh plugin --profile web remove @anionex/dsh-vision-toolkit
Typical dsh-vision-toolkit Usage
Once installed and configured, paste an image or invoke the visual Skill and the model does the rest: Q&A, OCR, UI restoration, pixel diff, and grounding all happen through tools. The usage examples below all come from the official README (source).
1. Paste an image and ask directly — paste a screenshot and ask "Where is the error in this screenshot?"; the model auto-switches to the vision variant and calls vision_glance to ground the answer, with multi-image comparison supported.
2. Invoke the visual Skill — use /vision-skills to let the model pick from five playbooks (long-screenshot OCR, UI restoration, graphic restoration, structure restoration, GUI operation):
/vision-skills
3. Long-screenshot OCR — ask the model to OCR a very long screenshot with vision_long_screenshot_ocr, returning Markdown, chunked results, and an element list:
OCR this long screenshot and output Markdown with all buttons and form elements
4. UI restoration — ask the model to rebuild a design or interface screenshot into front-end UI:
Restore this screenshot into front-end UI and output the HTML structure
5. Pixel diff — compare two screenshots in regression testing with vision_pixel_diff for the difference percentage and ranked regions:
Compare these two screenshots and output the difference percentage and regions
6. Grounding and cropping — let vision_ground return raw pixel coordinates (x1,y1,x2,y2), then crop with vision_crop:
Locate the error dialog in this screenshot and crop it out
dsh-vision-toolkit Troubleshooting
The most common dsh-vision-toolkit issues cluster in four places: an incompatible vision API response, paste still reported as unsupported, 429 rate limiting, and first-run environment setup failures. The troubleshooting notes below all come from the official README (source).
1. "Vision API returned an incompatible response structure" — the symptom is the API complaining about the response structure after configuring the provider; the cause is a baseUrl missing the /v1 path prefix, for example LM Studio or Ollama need http://127.0.0.1:1234/v1. Fix: add the /v1 prefix in the configuration and retry.
2. Pasting an image still reports unsupported — the symptom is the model still saying it cannot handle the image; the cause is a stale route or a missing restart. Fix: restart the Web profile, confirm the route carries the (Vision Toolkit) suffix, or put the image in the workspace and call /vision-skills.
3. The vision API returns 429 — the symptom is the request being rate limited; the cause is insufficient quota on the endpoint. Fix: wait for the Retry-After window, or switch to your own endpoint.
4. First-run setup fails — the symptom is an error while preparing the environment on first launch; the cause is a network or disk problem preventing the standalone Python download. Fix: check the network and disk, install system Python 3.11+, or set runtime.python in the configuration.
Use Cases and Notes
dsh-vision-toolkit suits anyone who needs DeepSeek Harness to handle images: reading screenshots, answering image questions, restoring UI, comparing pixels, and locating elements in both Web and Headless. The scenarios and limits below all come from the official README (source).
- Use cases: paste an error screenshot during development and testing to locate the problem directly; OCR very long screenshots in bulk; restore front-end UI from a design; compare pixel differences between two screenshots in regression tests; guide the model through GUI operations with the playbook.
- Notes: transparent routing is enabled by default and can be disabled under the advanced setting "Transparent variant routing" to restore explicit variant entries; each call sends only the necessary intent plus images and context does not accumulate across calls, keeping cost small; you can deploy small local multimodal models such as Gemma 4 or Qwen 3.5/3.6 to reduce cost further; set
VISION_SSL_VERIFY=0for self-signed certificates or MITM setups; the HTML screenshot tool needs Chrome/Chromium/Edge installed, and the other tools work regardless.
Project Links
dsh-vision-toolkit is an MIT open-source project maintained by Anionex.
Plugin details: dsh-vision-toolkit plugin page.
This page is an independent guide rewritten from the plugin's official README — for the authoritative documentation and the latest changes, please refer to the source: Anionex/dsh-vision-toolkit. Plugins are third-party code that runs on your machine after installation; listing here is not an endorsement — please review the source code yourself before installing.
FAQ
dsh-vision-toolkit is the first comprehensive vision-tool plugin in the DeepSeek Harness ecosystem. It gives text-only models image Q&A, long-screenshot OCR, UI restoration, and pixel diff, so you can paste an image and ask directly.
When you paste an image in DSH Web, dsh-vision-toolkit's transparent routing automatically switches the text-only model to its (Vision Toolkit) variant — no manual model changes, and thumbnails, history, and workspace paths stay intact.
dsh-vision-toolkit ships 10 tools: vision_glance for image Q&A, vision_long_screenshot_ocr for long-screenshot OCR, vision_pixel_diff for pixel diff, vision_ground for grounding, vision_crop for cropping, and more.
dsh-vision-toolkit's vision-skills Skill bundles five playbooks: long-screenshot OCR, UI restoration, graphic restoration, structure restoration, and GUI operation, guiding tool choice and result verification.
dsh-vision-toolkit needs a vision provider before use: restart the Web profile, configure it in Settings → Vision Toolkit, and test the model. First launch prepares an isolated Python runtime, with system Python 3.11+ preferred.
dsh-vision-toolkit is cheap to run: each call sends only the necessary intent plus images and context does not accumulate across calls, so cost stays low; local models such as Gemma 4 or Qwen 3.5 reduce it further.
Related Terms
- dsh-vision-toolkit
- dsh-vision-toolkit is the vision toolbox plugin for DeepSeek Harness (DSH), shipping 10 vision tools and a vision-skills Skill that give text-only models image Q&A, long-screenshot OCR, UI restoration, and pixel diff.— dsh-vision-toolkit README
- Transparent variant routing
- Transparent variant routing is dsh-vision-toolkit's default-on mechanism that auto-switches the text-only model to its (Vision Toolkit) variant when an image is pasted; it can be disabled in advanced settings.— dsh-vision-toolkit README
- vision-skills
- vision-skills is the bundled Skill with five playbooks — long-screenshot OCR, UI restoration, graphic restoration, structure restoration, and GUI operation — guiding tool selection, execution order, and result verification.— dsh-vision-toolkit README
- focus hint
- focus hint is dsh-vision-toolkit's mechanism for passing the user message or the model's reasoning as a focus prompt to the vision model, producing task-aware descriptions with fewer tokens and higher accuracy.— dsh-vision-toolkit README
- (Vision Toolkit) variant
- The (Vision Toolkit) variant is the vision form of a text-only model in DSH, switched to automatically by transparent routing when you paste an image, or selectable manually in the model picker.— dsh-vision-toolkit README