dsh-vision-toolkit: A Vision Toolkit Plugin for DeepSeek Harness

anionex/dsh-vision-toolkit

Tools & CapabilitiesVerified
Listed on 2026-08-19
Page last updated 2026-10-04

DSH Vision Toolkit gives text-only models in DeepSeek Harness eyes: paste-and-ask image Q&A, long-screenshot OCR, UI restoration, grounding, and pixel diff—install with one command.

dsh-vision-toolkit is the first comprehensive vision-tool plugin in the DeepSeek Harness ecosystem, giving text-only models capabilities like image Q&A, long-screenshot OCR, UI restoration, and pixel diff. It solves the problem of text-only models being unable to process images, letting you paste an image and ask questions directly, recognize screenshot content, restore front-end UI, and handle multi-image Q&A. Maintained by Anionex under the MIT license, it deeply integrates with DSH's Web and Headless Profiles, making it the first comprehensive vision-tool plugin in the DSH ecosystem.

How to Install

install
npmdsh plugin --profile web add @anionex/dsh-vision-toolkit
Category
Tools & Capabilities
Platform
DSH Plugin
Author
anionex
Distribution
Plugin

dsh-vision-toolkit Key Features

Paste-and-ask image Q&ALong-screenshot OCRUI restoration & groundingPixel diff verification
Listed on dsh-plugin.org

dsh-vision-toolkit Repo Summary

What Does It Do?

DSH Vision Toolkit is a DSH plugin for DeepSeek Harness that gives text-only models powerful vision capabilities, including image Q&A, long-screenshot OCR, UI restoration, and pixel diff. It solves the problem of text-only models being unable to process images, letting you paste an image and ask questions directly, recognize screenshot content, restore front-end UI, and handle multi-image Q&A. Maintained by Anionex under the MIT license, it deeply integrates with DSH's Web and Headless Profiles, making it the first comprehensive vision-tool plugin in the DSH ecosystem.

Core Features

  • Image Q&A: paste an image and ask directly; the model automatically switches to its vision variant without manual switching.
  • Long-screenshot OCR: extract text from long screenshots to capture key information.
  • UI restoration: rebuild front-end UI from screenshots, aiding interface development and testing.
  • Pixel diff: compare pixel differences between two images for screenshot testing.
  • Grounding and cropping: locate targets (e.g., buttons, error messages) in images, with cropping and tracing support.
  • Built-in vision Skill: a bundled Skill guides the model on which tool to choose for different visual tasks, how to proceed, and how to verify results.

How to Use This Plugin?

After enabling the plugin in DSH, go to Settings → Vision Toolkit to configure a vision provider, then start using the tools. In the Web UI, simply paste an image and ask a question; the model will automatically switch to its vision variant and respond. In Headless mode, you can invoke vision capabilities through the corresponding tools. To adjust behavior, modify options like transparent routing in advanced settings or refer to the official documentation for model configuration.

This page is an independent rewrite of the plugin's official README — for authoritative documentation and the latest changes, refer to the source: anionex/dsh-vision-toolkit. The plugin is third-party code that runs on your machine once installed; inclusion does not imply endorsement — please review the source before installing.

Install, update, and uninstall this plugin in the Plugin Market of DeepSeek Harness's DSH Plugin Hub

More DSH Plugin articles

View all 1 articles

DSH Plugin FAQ