chocobo77/dsh-infinite-context

DeepSeek Harness plugin: multi-tier memory management, semantic retrieval, structured memory, and model-context awareness for infinite context.

Project Overview项目介绍

This is a native plugin built exclusively for DeepSeek Harness (DSH) that enables infinite context for long-running conversations through a multi-layer memory management system. It implements progressive token compression that prioritizes summarizing older messages while keeping recent full conversations intact, and dynamically adjusts compression thresholds based on the real running context window of the currently used model. It can even intervene mid-generation when the combined input and output context approaches the window limit, and organizes memories into a three-layer pyramid structure that gets persisted to a SQLite database for retention across application restarts. It also adds built-in semantic retrieval and three-layer deduplication to prevent redundant memory entries.

This plugin is designed for DSH users who regularly work on long multi-turn conversations or complex development tasks, and solves common problems like context window overflow and lost memory during extended sessions. During normal use, it automatically tracks token usage for the active session, and triggers compression automatically when usage hits the configured threshold, so users do not have to manually manage context most of the time. It also ships with 10 manual management tools that let users directly perform actions like semantic memory search, status reporting, structured index generation, forced compression, and full memory reset as needed.

The plugin is released under the permissive MIT open source license, and requires a Node.js environment to run. It supports multiple installation methods, including temporary patching for local development, tarball installation to a DSH profile, and manual copying to the DSH plugins directory. It defaults to using a dependency-free lightweight embedder for vector indexing, but can also be configured to use Transformers for embedding if preferred. For local models, it can actively probe the real context window of llama, ollama, and openai-compatible services to avoid overflow caused by inflated declared window sizes.

这是一个专为 DeepSeek Harness (DSH) 开发的原生插件,通过多层记忆管理为长对话提供「无限上下文」体验。它支持 token 压力驱动的渐进式压缩,能根据模型真实上下文窗口推导动态压缩阈值,还能在模型深度思考过程中介入处理溢出的上下文。它采用三层记忆金字塔存储,用 SQLite 做持久化存储,自带语义检索和三层去重避免重复入库。

该插件适合需要进行多轮长对话开发或复杂任务处理的 DSH 用户,能够解决长对话过程中上下文溢出、记忆丢失的问题。日常使用中,它会自动跟踪当前会话的 token 占用,在达到阈值时自动对最早的对话做摘要压缩,保留近期对话原文。用户也可以调用 10 种手动工具,执行记忆检索、状态查看、索引生成、强制压缩、清空记忆等操作。

该插件采用 MIT 开源许可,依赖 Node.js 环境运行,支持本地开发补丁安装或通过 tarball 包安装到 DSH 配置文件。它默认使用无依赖的轻量级嵌入器,也可配置使用 Transformers 做文本嵌入。对本地模型,它支持主动探测 llama、ollama、openai 兼容服务的真实上下文窗口,避免因声明窗口虚高导致上下文溢出。

Pre-install check安装前体检Compatibility · Security兼容性 · 安全性 1 warning1 项注意
  • Only 2 stars - very few users, little community feedback星标只有 2,几乎没人在用,遇到问题缺少社区反馈
DSH walks through these 9 checksDSH 会逐条核对这 9 项

Compatibility兼容性

  • DSH, Node, OS and profile requirementsDSH 版本 / Node 版本 / 操作系统 / profile 是否满足要求
  • External dependencies and runtimes (Electron / Python / Docker, ...)外部依赖与运行时(Electron / Python / Docker 等)是否齐备
  • Conflicts with installed plugins: command names, skill / tool names, ports, duplicate MCP registration与已装插件是否冲突:命令名、skill / tool 重名、端口占用、重复 MCP 注册

Security安全性

  • Repo matches the facts registered here; archived or abandoned?仓库是否与页面登记一致,是否归档或长期停更
  • Safety of preinstall / install / postinstall and install.sh / setup.ps1preinstall / install / postinstall 与 install.sh、setup.ps1 是否安全
  • curl|bash, download-then-execute, obfuscation, unrelated domains → stop immediatelycurl|bash、下载即执行、混淆代码、无关域名 → 立刻停止
  • Typosquatting or unmaintained packages among the new dependencies新增依赖里有没有 typosquatting 或无人维护的包
  • Requested permissions vs. what the feature actually needs申请了哪些权限、是否超出功能所需(filesystem / network / shell / clipboard)
  • Any sudo / admin requirement, plus uninstall and rollback是否要求 sudo / 管理员权限,以及卸载与回滚方式

Anything uncertain must be marked unknown with a note on how to confirm it. This site's signal screen is a static snapshot, not a security audit.拿不准的必须标「未知」并说明要我怎么确认。本站的信号筛查是静态快照,不能替代安全审计。

Or use CLI install (for developers)或使用命令行安装(适合开发者)

CLI Install命令行安装

dsh plugin --profile web add github:chocobo77/dsh-infinite-context

把 chocobo77/dsh-infinite-context 加入你的 DSH 配置(web profile)即可启用。

READMEREADME

dsh-infinite-context

🇨🇳 中文 | 🇬🇧 English


简介

一个 DeepSeek Harness (DSH) 插件,通过多层记忆管理让长对话拥有「无限上下文」体验:

  • 渐进式压缩 — token 压力驱动,最老消息优先摘要,近期对话原样保留
  • 动态压缩阈值 — 按路由模型的真实 CTX 推导触发水位;单轮超长 think 造成的一次性增长可越过轮次间隔立即介入
  • 深度思考介入 — 包装 agent 的 LLM 流:input + output 逼近当前模型真实窗口时注入溢出信号 → 持久压缩 → 带余量重试,模型思考中也能被介入
  • 三层记忆金字塔 — short(近期原文)→ mid(LLM 摘要)→ long(合并摘要)
  • 持久化存储 — SQLite(node:sqlite),重启不丢记忆
  • 语义检索 — 记忆嵌入、索引,每轮注入最相关的 top-K 记忆
  • 三层去重 — 精确 + 归一化模糊 + 语义余弦,防止重复入库
  • 结构化记忆 — 四分类(user/feedback/project/reference)+ 索引 + 审计 + 忘得可见
  • 模型上下文感知 — 自动采纳 DSH 解析的真实模型 CTX,本地小模型提前压缩
  • 高价值过滤 — 只入库高价值工具结果,低价值工具自动过滤
  • 手动工具 — 10 个:search / status / index / maintain / model_probe / forget / consolidate / reset / force_compress / ingest

核心特性

特性 说明
渐进式压缩 compress_trigger_ratio: 0.75 — 上下文 >75% 就压缩(本地模型长上下文 TPS 骤降,提前介入);compress_target_ratio: 0.6 — 只摘要溢出部分
动态压缩阈值 compaction_dynamic_threshold: true — 真实窗口 < 声明窗口时按真实窗口强制压缩(探测/modelWindows 驱动);thresholdRatio: 0.7 + compaction_dynamic_floor: 0.5 — 触发比例随窗口填充从 0.7 滑向 0.5(~60% 触发,本地模型留在高速区);单轮激增(≥20% 窗口)越过轮次间隔;同一会话两次强制压缩间隔 ≥10s
深度思考介入 thinking_guard_enabled: true — 包装 llm/stream:input + output 逼近 窗口 − (system/tools + 摘要估算 + 余量) 动态线时注入 CONTEXT_WINDOW_EXCEEDED → 持久压缩 → 重试;输入单独超线则生成前先压缩;thinking_guard_ratio: 0.9 为触发上限
三层去重 精确(hasText)+ 归一化(normalizeForDedup)+ 语义(cosine ≥ 0.92)
结构化记忆 memory_index(MEMORY.md 索引)+ memory_maintain(审计)+ 忘得可见
模型 CTX 感知 自动读取 DSH 模型目录的 contextWindow;本地模型主动探测真实运行窗口(llama/ollama/openai,含 llama-server meta.n_ctx);per-model 注册表按模型隔离
高价值过滤 denylist 过滤 23 个低价值工具;importance 分级(short=0.3/mid=0.6/long=0.6,long 继承批次 max)

架构

src/
├── types.ts              核心类型(无依赖)
├── embedder.ts           轻量级特征哈希嵌入器(无依赖)
├── vector-index.ts       内存向量索引(无依赖)
├── memory-store.ts       SQLite 持久化存储(无依赖)
├── token-budget.ts       CJK-aware token 估算 + 内容块计量(无依赖)
├── forgetting.ts         遗忘策略(无依赖)
├── memory-engine.ts      记忆引擎核心(无依赖)
├── model-context.ts      模型上下文跟踪器:探测时机 + per-model 注册表(无依赖)
├── model-probe.ts        主动探测:llama/ollama/openai + 本地/在线判定(无依赖)
├── compaction-policy.ts  压缩触发决策:动态比例、skip/force/delegate(无依赖)
├── summarization-target.ts 摘要目标路由(跟随会话模型,无依赖)
├── config.ts             schemastery 配置解析
├── memory-context.ts     Cordis 服务(动态 CTX 感知 + 探测接线)
├── memory-compaction.ts  压缩引擎(渐进式 + RAG + 清理 + thinking guard 接线)
├── thinking-guard.ts     深度思考介入:llm/stream 包装 + 动态触发线
├── OutputSanitizer.ts    工具结果清理
├── VectorRetriever.ts    RAG 检索/入库
├── strings.ts            共享字符串工具
├── core.ts               公共导出桶
├── index.ts              完整导出桶
└── tools.ts              10 个手动工具
tests/                    133 个单元测试

Showing the opening section of the README — the full document lives in the repository以上为 README 开头摘要,完整文档在仓库内 · View the full README on GitHub →在 GitHub 查看完整 README →

← 上一个 Prev dsh-comfyui-image 下一个 Next dsh-task-board →