dsh-local-models: Run Local GGUF Models inside DeepSeek Harness
vmarcelo49/dsh-local-models
Adds a Local Models tab to the DeepSeek Harness Web GUI: pick a .gguf file, tune context and speculative decoding, watch a live VRAM estimate, load it via llama-server, and register the running server as a dsh LLM provider in one click.
dsh-local-models is a DSH plugin that brings local model support to DeepSeek Harness by adding a Local Models tab to the Web GUI settings. It removes the friction of running a local model and then wiring it into an agent framework by hand: pick a .gguf file, tune context and speculative decoding, watch a live VRAM estimate, load it through llama-server, and register the running server as a DSH provider in one click. Built against stock upstream llama.cpp with no patches or build step, under the MIT license.
How to Install
dsh plugin --profile web add github:vmarcelo49/dsh-local-models- Category
- Models & Reasoning
- Platform
- DSH Plugin
- Author
- vmarcelo49
- Distribution
- Plugin
dsh-local-models Key Features
dsh-local-models Repo Summary
What Does It Do?
dsh-local-models is a DSH plugin for DeepSeek Harness that adds a Local Models tab to the DSH Web GUI, letting you pick a .gguf file, tune context and speculative decoding, watch a live VRAM estimate, and load the model through llama-server. It solves the friction of getting a local model running and then wiring it into an agent framework: once the server is up, one click registers it as an LLM provider inside DSH. Maintained by Vmarcelo49 under the MIT license, it is built against stock upstream llama.cpp with no fork, no patches, and no build step.
Core Features
- Model picker: an in-app file browser that shows directories and .gguf files only, with a header-only GGUF parse reporting architecture, quantization, layer count, context length, and MoE detection.
- Launch options: a context slider in 8K steps capped at the model's trained context, plus MTP draft depth, thinking level, optional vision mmproj with GPU or CPU offload, and MoE expert placement with a fit-to-VRAM helper.
- Live VRAM estimate: weights, KV cache, recurrent state, and compute/graph overhead are summed and compared against 16 GB, with fits, safe-margin, and max-context-that-fits rows.
- Profiles and router mode: save named launch configurations and reload them in one click; router mode serves all saved profiles from one OpenAI-compatible endpoint, loading models on demand with one resident at a time by default.
- Register in DSH: writes the ready server as a provider route, including vision modality and thinking levels.
- Terminal overlay: a live tail of the llama-server log inside the tab.
How to Use This Plugin?
After enabling it in the DSH web profile, open the Local Models tab in Settings and click "Choose GGUF…" to pick a model file, using the Home / Models shortcuts and Up navigation to browse. Tune context, MTP draft depth, and thinking level, optionally configure mmproj and MoE settings, then click "Load model" and watch the status card, inspecting output via "Open terminal" if needed. When the server is ready, click "Register in dsh" and the route appears in the Models picker; alternatively save profiles and use "Start router (from profiles)" for a multi-model endpoint. Port, server binary path, file-browser shortcut directories, and router concurrency are configurable through environment variables, while the VRAM estimate constants live at the top of the client script for machines with different GPUs.
How to Troubleshoot This Plugin?
The VRAM estimate has a known accuracy gap for the Gemma family, so when the estimate disagrees with real usage, trust the actual runtime log. The plugin depends on a working llama-server binary; if loading fails, first confirm that binary exists and matches your Vulkan, CUDA, or CPU setup. The VRAM estimate constants default to a 16 GB GPU and need adjusting per the official instructions on other hardware. Changes to server routes or the inject list require a DSH restart, while client-side changes only need a page refresh.
This page is an independent rewrite of the plugin's official README — for authoritative documentation and the latest changes, refer to the source: vmarcelo49/dsh-local-models. The plugin is third-party code that runs on your machine once installed; inclusion does not imply endorsement — please review the source before installing.
