billion-context: A Context Compression Proxy for Long Sessions in DeepSeek Harness
ranxianglei/billion-context
billion-context is a universal context-compression proxy for AI coding agents; mounted as a DSH plugin it folds consumed conversation into layered summaries so one session can push billions of tokens without blowing the context window.
billion-context is a context-compression proxy for DeepSeek Harness that sits between any AI coding agent and its model API, rewriting consumed conversation into layered summaries. It targets the pain of long sessions blowing past the context window, steadily rising per-token billing, and sessions degrading or dying once the window is exceeded, letting a single session run for days. The model decides when and what to compress, summaries are written incrementally and can be decompressed on demand, the prefix cache stays intact, and four context-management tools are injected into the conversation. Maintained by ranxianglei under the MIT License in TypeScript.
How to Install
dsh plugin --profile web add billion-context- Category
- Memory & Context
- Platform
- DSH Plugin
- Author
- ranxianglei
- Distribution
- Plugin
billion-context Key Features
billion-context Repo Summary
What Does It Do?
billion-context is a DSH plugin for DeepSeek Harness that acts as a universal context-compression proxy between any AI coding agent and its model API. It solves the problem of long coding sessions blowing past the context window, where sessions degrade or die and every token is billed. Maintained by ranxianglei under the MIT License and written in TypeScript, it has 216 stars and was last updated in 2026-09.
Core Features
- Universal proxy: any agent that can set a base URL works, with zero per-agent adapter code.
- Model-driven compression: the model decides when and what to compress into high-fidelity summaries, instead of hitting a hard truncation limit.
- Incremental and reversible: summaries are written in small ranges, can be decompressed on demand, and keep the prefix cache intact.
- Built-in context tools: four context-management tools (compress, decompress, search_context, acp_status) are injected into the conversation and executed server-side.
- Opt-in extras: absorb compresses individual tool results the moment they arrive, acp_rule records persistent principle-level reminders that are hard-protected from compression, and protectedLatestTools keeps the latest snapshot of a cumulative tool un-compressible.
How to Use This Plugin?
After enabling it in DSH, point your coding agent's base URL at the proxy. It parses the request in Anthropic or OpenAI shape, runs compression on the conversation, forwards to the real model API, and rewrites the streaming response. When the conversation grows, the model calls compress and the compressed ranges are folded into history before the next turn; use decompress or search_context to bring details back. Behavior can be tuned through configuration options such as absorb, acp_rule, and protectedLatestTools, whose exact keys are documented in the repository's configuration guide.
How to Troubleshoot This Plugin?
When compression behaves unexpectedly, call acp_status first to inspect the current compression state and recorded rules, so you can tell whether configuration took effect or compression never triggered. If a tool result was folded unexpectedly, check whether protectedLatestTools covers that tool name; if principle-level reminders disappear, confirm acp_rule is enabled. If the issue persists, follow the official instructions to repair it and cross-check options against the repository's configuration guide.
This page is an independent rewrite of the plugin's official README — for authoritative documentation and the latest changes, refer to the source: ranxianglei/billion-context. The plugin is third-party code that runs on your machine once installed; inclusion does not imply endorsement — please review the source before installing.
