dsh-vision-toolkit: A Vision Toolkit Plugin for DeepSeek Harness
anionex/dsh-vision-toolkit
DSH Vision Toolkit gives text-only models in DeepSeek Harness eyes: paste-and-ask image Q&A, long-screenshot OCR, UI restoration, grounding, and pixel diff—install with one command.
dsh-vision-toolkit is the first comprehensive vision-tool plugin in the DeepSeek Harness ecosystem, giving text-only models capabilities like image Q&A, long-screenshot OCR, UI restoration, and pixel diff. It solves the problem of text-only models being unable to process images, letting you paste an image and ask questions directly, recognize screenshot content, restore front-end UI, and handle multi-image Q&A. Maintained by Anionex under the MIT license, it deeply integrates with DSH's Web and Headless Profiles, making it the first comprehensive vision-tool plugin in the DSH ecosystem.
How to Install
dsh plugin --profile web add @anionex/dsh-vision-toolkit- Category
- Tools & Capabilities
- Platform
- DSH Plugin
- Author
- anionex
- Distribution
- Plugin
dsh-vision-toolkit Key Features
dsh-vision-toolkit Repo Summary
What Does It Do?
DSH Vision Toolkit is a DSH plugin for DeepSeek Harness that gives text-only models powerful vision capabilities, including image Q&A, long-screenshot OCR, UI restoration, and pixel diff. It solves the problem of text-only models being unable to process images, letting you paste an image and ask questions directly, recognize screenshot content, restore front-end UI, and handle multi-image Q&A. Maintained by Anionex under the MIT license, it deeply integrates with DSH's Web and Headless Profiles, making it the first comprehensive vision-tool plugin in the DSH ecosystem.
Core Features
- Image Q&A: paste an image and ask directly; the model automatically switches to its vision variant without manual switching.
- Long-screenshot OCR: extract text from long screenshots to capture key information.
- UI restoration: rebuild front-end UI from screenshots, aiding interface development and testing.
- Pixel diff: compare pixel differences between two images for screenshot testing.
- Grounding and cropping: locate targets (e.g., buttons, error messages) in images, with cropping and tracing support.
- Built-in vision Skill: a bundled Skill guides the model on which tool to choose for different visual tasks, how to proceed, and how to verify results.
How to Use This Plugin?
After enabling the plugin in DSH, go to Settings → Vision Toolkit to configure a vision provider, then start using the tools. In the Web UI, simply paste an image and ask a question; the model will automatically switch to its vision variant and respond. In Headless mode, you can invoke vision capabilities through the corresponding tools. To adjust behavior, modify options like transparent routing in advanced settings or refer to the official documentation for model configuration.
This page is an independent rewrite of the plugin's official README — for authoritative documentation and the latest changes, refer to the source: anionex/dsh-vision-toolkit. The plugin is third-party code that runs on your machine once installed; inclusion does not imply endorsement — please review the source before installing.
