What Is dsh-vision-toolkit? Vision for Text-Only DSH Models

GuidePublished 2026-08-30Author: DeepSeek Plugin Market
dsh-vision-toolkitvision toolsDeepSeek HarnessDSHOCR
dsh-vision-toolkit is the first comprehensive vision-tool plugin for DeepSeek Harness: image Q&A, long-screenshot OCR, and UI restoration for text-only models.

dsh-vision-toolkit is the first comprehensive vision-tool plugin in the DeepSeek Harness (DSH) ecosystem: it gives text-only models eyes — 10 vision tools plus a bundled vision-skills Skill for image Q&A, long-screenshot OCR, UI restoration, pixel diff, grounding, and cropping. This article explains what dsh-vision-toolkit is, its core features, the full install/update/uninstall commands, typical usage, and common troubleshooting, so you can give DeepSeek Harness vision capabilities.

What Is dsh-vision-toolkit?

dsh-vision-toolkit solves the problem that text-only models cannot understand images: it is the first comprehensive vision-tool plugin in the DSH ecosystem, letting the model paste an image and ask directly, read screenshots, restore front-end UI, and run pixel diff, grounding, and cropping tasks. The positioning and facts below all come from the official README (source).

dsh-vision-toolkit (GitHub repo Anionex/dsh-vision-toolkit) is maintained by Anionex and open sourced under the MIT license, built on the upstream agent-vision-toolkit. It is the native DeepSeek Harness integration of that visual workflow — read, ground, crop, trace, rebuild, and verify — brought into Web and Headless Profiles. In DSH Web, pasting an image automatically switches the text-only model to its (Vision Toolkit) variant with native thumbnails, session history, and workspace paths intact, and Web can preview artifacts; Headless mode invokes vision through tools such as read_image and OCR. It is an out-of-the-box vision plugin in the DSH Plugin ecosystem, with a bilingual UI.

What Are the Core Features of dsh-vision-toolkit?

The core capabilities revolve around five things: ten vision tools, the focus hint mechanism, the vision-skills Skill, transparent routing, and an isolated Python runtime — configured once, then paste an image to start. The capabilities below all come from the official README (source).

  • Image Q&A and multi-image comparison: vision_glance answers questions such as "What is this?" or "Where is the error?" and can compare several images at once.
  • Long-screenshot OCR: vision_long_screenshot_ocr reads very long screenshots and outputs Markdown, chunked results, an element list, and an audit.
  • UI restoration: rebuild front-end UI from a screenshot into UI code or a structural description for development and testing.
  • Pixel diff: vision_pixel_diff outputs the difference percentage, ranked diff regions, a heatmap, and JSON.
  • Grounding and cropping: vision_ground returns raw pixel coordinates (x1,y1,x2,y2) with an optional boxed preview; vision_crop crops PNG/JPEG; vision_trace vectorizes into SVG.
  • Transparent routing: pasting an image auto-switches to the (Vision Toolkit) variant with no manual model switching, keeping history and workspace paths intact.

How to Install and Enable dsh-vision-toolkit?

Install with one npm command into the DeepSeek Harness Web profile, restart it, configure a vision provider in Settings → Vision Toolkit, test the model, then paste an image or call /vision-skills. The install, update, and uninstall commands below all come from the official README (source).

1. Install — the stable release is recommended from npm into the Web profile:

bash
dsh plugin --profile web add @anionex/dsh-vision-toolkit

Headless and Desktop use their own profiles:

bash
dsh plugin --profile headless add @anionex/dsh-vision-toolkit

2. Enable — restart the Web profile, open Settings → Vision Toolkit, configure a vision provider (for example Gemini 3.7 Flash), and click Test vision model. The first launch prepares an isolated Python runtime: system Python 3.11+ is preferred; otherwise a standalone ~35 MB Python is downloaded and Pillow, NumPy, and vtracer are installed from a mirror.

3. Update — re-run the same add command pinned to the latest release (update to latest):

bash
# update to latest
dsh plugin --profile web add @anionex/dsh-vision-toolkit@latest

4. Uninstall — remove it from the Web profile:

bash
dsh plugin --profile web remove @anionex/dsh-vision-toolkit

Typical dsh-vision-toolkit Usage

Once installed and configured, paste an image or invoke the visual Skill and the model does the rest: Q&A, OCR, UI restoration, pixel diff, and grounding all happen through tools. The usage examples below all come from the official README (source).

1. Paste an image and ask directly — paste a screenshot and ask "Where is the error in this screenshot?"; the model auto-switches to the vision variant and calls vision_glance to ground the answer, with multi-image comparison supported.

2. Invoke the visual Skill — use /vision-skills to let the model pick from five playbooks (long-screenshot OCR, UI restoration, graphic restoration, structure restoration, GUI operation):

bash
/vision-skills

3. Long-screenshot OCR — ask the model to OCR a very long screenshot with vision_long_screenshot_ocr, returning Markdown, chunked results, and an element list:

bash
OCR this long screenshot and output Markdown with all buttons and form elements

4. UI restoration — ask the model to rebuild a design or interface screenshot into front-end UI:

bash
Restore this screenshot into front-end UI and output the HTML structure

5. Pixel diff — compare two screenshots in regression testing with vision_pixel_diff for the difference percentage and ranked regions:

bash
Compare these two screenshots and output the difference percentage and regions

6. Grounding and cropping — let vision_ground return raw pixel coordinates (x1,y1,x2,y2), then crop with vision_crop:

bash
Locate the error dialog in this screenshot and crop it out

dsh-vision-toolkit Troubleshooting

The most common dsh-vision-toolkit issues cluster in four places: an incompatible vision API response, paste still reported as unsupported, 429 rate limiting, and first-run environment setup failures. The troubleshooting notes below all come from the official README (source).

1. "Vision API returned an incompatible response structure" — the symptom is the API complaining about the response structure after configuring the provider; the cause is a baseUrl missing the /v1 path prefix, for example LM Studio or Ollama need http://127.0.0.1:1234/v1. Fix: add the /v1 prefix in the configuration and retry.

2. Pasting an image still reports unsupported — the symptom is the model still saying it cannot handle the image; the cause is a stale route or a missing restart. Fix: restart the Web profile, confirm the route carries the (Vision Toolkit) suffix, or put the image in the workspace and call /vision-skills.

3. The vision API returns 429 — the symptom is the request being rate limited; the cause is insufficient quota on the endpoint. Fix: wait for the Retry-After window, or switch to your own endpoint.

4. First-run setup fails — the symptom is an error while preparing the environment on first launch; the cause is a network or disk problem preventing the standalone Python download. Fix: check the network and disk, install system Python 3.11+, or set runtime.python in the configuration.

Use Cases and Notes

dsh-vision-toolkit suits anyone who needs DeepSeek Harness to handle images: reading screenshots, answering image questions, restoring UI, comparing pixels, and locating elements in both Web and Headless. The scenarios and limits below all come from the official README (source).

  • Use cases: paste an error screenshot during development and testing to locate the problem directly; OCR very long screenshots in bulk; restore front-end UI from a design; compare pixel differences between two screenshots in regression tests; guide the model through GUI operations with the playbook.
  • Notes: transparent routing is enabled by default and can be disabled under the advanced setting "Transparent variant routing" to restore explicit variant entries; each call sends only the necessary intent plus images and context does not accumulate across calls, keeping cost small; you can deploy small local multimodal models such as Gemma 4 or Qwen 3.5/3.6 to reduce cost further; set VISION_SSL_VERIFY=0 for self-signed certificates or MITM setups; the HTML screenshot tool needs Chrome/Chromium/Edge installed, and the other tools work regardless.

dsh-vision-toolkit is an MIT open-source project maintained by Anionex.

Plugin details: dsh-vision-toolkit plugin page.

This page is an independent guide rewritten from the plugin's official README — for the authoritative documentation and the latest changes, please refer to the source: Anionex/dsh-vision-toolkit. Plugins are third-party code that runs on your machine after installation; listing here is not an endorsement — please review the source code yourself before installing.

FAQ

What is dsh-vision-toolkit and what problems does it solve?

dsh-vision-toolkit is the first comprehensive vision-tool plugin in the DeepSeek Harness ecosystem. It gives text-only models image Q&A, long-screenshot OCR, UI restoration, and pixel diff, so you can paste an image and ask directly.

How does dsh-vision-toolkit switch a text-only DeepSeek Harness model to its vision variant automatically?

When you paste an image in DSH Web, dsh-vision-toolkit's transparent routing automatically switches the text-only model to its (Vision Toolkit) variant — no manual model changes, and thumbnails, history, and workspace paths stay intact.

Which vision tools does dsh-vision-toolkit include?

dsh-vision-toolkit ships 10 tools: vision_glance for image Q&A, vision_long_screenshot_ocr for long-screenshot OCR, vision_pixel_diff for pixel diff, vision_ground for grounding, vision_crop for cropping, and more.

What is the vision-skills Skill in dsh-vision-toolkit?

dsh-vision-toolkit's vision-skills Skill bundles five playbooks: long-screenshot OCR, UI restoration, graphic restoration, structure restoration, and GUI operation, guiding tool choice and result verification.

What does dsh-vision-toolkit need before first use?

dsh-vision-toolkit needs a vision provider before use: restart the Web profile, configure it in Settings → Vision Toolkit, and test the model. First launch prepares an isolated Python runtime, with system Python 3.11+ preferred.

Does using dsh-vision-toolkit cost a lot?

dsh-vision-toolkit is cheap to run: each call sends only the necessary intent plus images and context does not accumulate across calls, so cost stays low; local models such as Gemma 4 or Qwen 3.5 reduce it further.

Related Terms

dsh-vision-toolkit
dsh-vision-toolkit is the vision toolbox plugin for DeepSeek Harness (DSH), shipping 10 vision tools and a vision-skills Skill that give text-only models image Q&A, long-screenshot OCR, UI restoration, and pixel diff.— dsh-vision-toolkit README
Transparent variant routing
Transparent variant routing is dsh-vision-toolkit's default-on mechanism that auto-switches the text-only model to its (Vision Toolkit) variant when an image is pasted; it can be disabled in advanced settings.— dsh-vision-toolkit README
vision-skills
vision-skills is the bundled Skill with five playbooks — long-screenshot OCR, UI restoration, graphic restoration, structure restoration, and GUI operation — guiding tool selection, execution order, and result verification.— dsh-vision-toolkit README
focus hint
focus hint is dsh-vision-toolkit's mechanism for passing the user message or the model's reasoning as a focus prompt to the vision model, producing task-aware descriptions with fewer tokens and higher accuracy.— dsh-vision-toolkit README
(Vision Toolkit) variant
The (Vision Toolkit) variant is the vision form of a text-only model in DSH, switched to automatically by transparent routing when you paste an image, or selectable manually in the model picker.— dsh-vision-toolkit README

Sources

View all articles