PolinniZhong/dsh-knit 预览 preview

PolinniZhong/dsh-knit

面向 DeepSeek Harness 的任务上下文检索:根据当前对话,从项目工作区找到最相关的上下文,组织为主要 / 辅助 / 相关内容,并通过 Knit 面板与 knit_docs 提供给人和 Agent。纯本地、确定性、零模型调用、零网络。 Task-aware workspace context retrieval for DeepSeek Harness. Knit finds the project context most relevant to the current task, organizes it into primary / supporting / related context, and exposes the same context to humans a

Project Overview项目介绍

dsh-knit is a native DeepSeek Harness plugin developed to organize project documents based on their relevance to your current ongoing conversation. To install it, run the single command dsh plugin --profile web add dsh-knit in your terminal, then restart DSH and perform a hard refresh of your browser to fully activate it. It scans all Markdown files, images, and videos in your current session’s working directory recursively (up to depth 6, skipping node_modules, .git, and dist folders), and sorts them using a pure local BM25 algorithm. No external model calls, no third-party API requests, and no internet access are required at any step of the process.

In addition to providing a human-facing UI that displays sorted documents in the DSH right sidebar, with in-place preview of Markdown, images, and videos, it also exposes a knit_docs tool for the DSH AI agent. When the agent needs to find existing project documents related to the current topic, it can call this tool to get a pre-sorted list of relevant documents, cutting down on wasted tokens from guessing paths and reading irrelevant files. Unlike memory-based agent tools that start with an empty index and require the agent to store documents before they can be recalled, Knit has immediate access to all of your project’s existing documents right after installation, making it ideal for long-lived projects with many historical documents.

dsh-knit is released under the permissive MIT open source license, has zero third-party dependencies after installation, and requires no additional build steps. It never initiates outbound network requests, so all of your project data stays on your local machine at all times, which improves both privacy and security. It does have some known limitations: sorting will fall back to modification time when the conversation is too short or there are very few documents, and it cannot parse the content of images or videos, only matching their file names to the current conversation.

dsh-knit 是专为 DeepSeek Harness 开发的原生插件,用于根据当前对话的内容,将项目内的文档按相关性排序。它会扫描当前会话工作目录下的所有 Markdown、图片和视频,使用纯本地运行的 BM25 算法排序,不调用大模型、不需要联网,就能在 DSH 右侧边栏展示相关文档,支持在面板内就地预览文档、图片和视频。可通过 dsh plugin --profile web add dsh-knit 命令安装,安装后需重启 DSH 并硬刷新浏览器。

它还为 DSH AI 代理提供了 knit_docs 工具。当代理需要查找和当前话题相关的项目文档时,可直接通过该工具拿到按相关性排好序的结果,省去代理反复猜测路径、逐个读取文件的额外消耗。和需要先写入内容才能召回的记忆类插件不同,Knit 安装后就能直接使用项目内所有已存在的历史文档,非常适合已有大量文档的项目。

本项目采用 MIT 许可证开源,安装后零第三方依赖、无额外构建步骤,也不会发起任何外部网络请求,所有运算都在本地完成,保障了项目文档的数据安全。它存在一些已知限制,例如对话过短或文档数量过少时排序效果会退化,也不解析图片和视频的画面内容,仅能根据文件名匹配相关性。

Pre-install check安装前体检Compatibility · Security兼容性 · 安全性 1 note1 项提示
  • 40 stars - an early-stage project星标 40,属于早期项目
DSH walks through these 9 checksDSH 会逐条核对这 9 项

Compatibility兼容性

  • DSH, Node, OS and profile requirementsDSH 版本 / Node 版本 / 操作系统 / profile 是否满足要求
  • External dependencies and runtimes (Electron / Python / Docker, ...)外部依赖与运行时(Electron / Python / Docker 等)是否齐备
  • Conflicts with installed plugins: command names, skill / tool names, ports, duplicate MCP registration与已装插件是否冲突:命令名、skill / tool 重名、端口占用、重复 MCP 注册

Security安全性

  • Repo matches the facts registered here; archived or abandoned?仓库是否与页面登记一致,是否归档或长期停更
  • Safety of preinstall / install / postinstall and install.sh / setup.ps1preinstall / install / postinstall 与 install.sh、setup.ps1 是否安全
  • curl|bash, download-then-execute, obfuscation, unrelated domains → stop immediatelycurl|bash、下载即执行、混淆代码、无关域名 → 立刻停止
  • Typosquatting or unmaintained packages among the new dependencies新增依赖里有没有 typosquatting 或无人维护的包
  • Requested permissions vs. what the feature actually needs申请了哪些权限、是否超出功能所需(filesystem / network / shell / clipboard)
  • Any sudo / admin requirement, plus uninstall and rollback是否要求 sudo / 管理员权限,以及卸载与回滚方式

Anything uncertain must be marked unknown with a note on how to confirm it. This site's signal screen is a static snapshot, not a security audit.拿不准的必须标「未知」并说明要我怎么确认。本站的信号筛查是静态快照,不能替代安全审计。

Or use CLI install (for developers)或使用命令行安装(适合开发者)

CLI Install命令行安装

dsh plugin --profile web add dsh-knit

把 PolinniZhong/dsh-knit 加入你的 DSH 配置(web profile)即可启用。

READMEREADME

Knit

Agent 一天产出 20 篇文档,你找不到刚才那篇。 Knit 把它们放到对话旁边 —— 你正在聊什么,相关的那篇就在最上面。

你的 agent 也一样。 同一份排序也给它当工具用 —— 它问「哪几篇相关」,拿回排名加每篇里命中的那段原文。

Knit 面板真机截图:右侧栏按相关性列出工作区文档,就地展开 Markdown 预览,预览头下方是引用条

真机截图(不是原型):一个含 51 篇文档的工作区 —— 列表 + 就地预览 + 引用条。 引用条是 v0.12 加的:展开就能看到「这篇被谁引用」,点一项直接跳过去。 顶行会跟着你正在聊什么变;对话内容还不足时它退回按时间排,并如实说明依据。 v0.14 起文档列表永远单列:最左边独立一列是序号(与标题第一行垂直居中),右边第一行= Primary 点 + 标题 + 相对时间(时间靠右),摘要再往下;路径只出现在预览头里(目录收敛成一个 …/ 占位,完整相对路径在悬停提示里),列表行不再重复。 序号是跨三档连续的一条(先看 / 辅助 / 背景);「其他相关文档」不在包里、不编号,那一格用一个 · 占位 —— 空着会被读成「漏了一个号」(2026-10-01)。

排序跟着对话走:发一句话,右侧栏的文档列表按这句话重排

排序跟着对话走(同一个工作区、刚新开一个会话):还没聊什么时,顶行如实写着「对话内容还不足,暂按最新排序」, 列表按时间排;问一句「排序算法用的 BM25 是怎么加权的?」,顶上立刻换成排序那几篇 —— 序号与「主要上下文 / 辅助上下文 / 相关上下文」三层随之一起出现(跨三档连续编号)。 *⚠️ 面板是每 5 秒轮询一次的,所以重排落在发消息之后的 0–5 秒内、不是瞬时;动图比真实时间快。*

扫整个项目文件夹的 Markdown / 图片 / 视频 · 排序跟着对话走 · 不调模型、不联网

dsh plugin --profile web add dsh-knit

这不就是个「最近文件列表」吗?

是的,但有两个关键区别:

  1. 范围:最近打开列表只记你点开过的文件;Knit 扫整个项目文件夹。 重启 DSH、新开会话、跨天回来,它都还在。
  2. 排序:它按时间排;Knit 按你正在聊什么排。

三句话说完它是什么:

  1. 扫整个项目文件夹的 Markdown,不只是这一轮生成的那几篇 —— 重启、换会话、跨天都还在
  2. 排序跟着你正在聊什么走:聊架构,架构文档浮上来;聊竞品,竞品分析浮上来
  3. 不调模型、不联网:全是本地字符串运算,零延迟、零成本、文档不出本机

「相关」是怎么算出来的

没有玄学,就是字符串运算。三步:

1. 读当前会话。 取最近 6 条用户 / 助手消息,只认真人输入的用户消息 (agent.inject() 塞进来的合成上下文不算,那会把话题带偏)。越新的消息权重越高:3 / 2 / 1 / 1 …

2. 抽关键词。

  • 英文词:取值很高,出现 1 次就要(chokidar、mtime 这种精确词)
  • 中文 2/3-gram:出现 2 次,或出现在最新那条消息里
  • 丢掉跨词边界的碎片:中文没有词边界,n-gram 会把相邻两个词的字粘起来 (「图片和」「个插」「的排」)。这类碎片有个共同特征 —— 首字或尾字是纯虚词, 一律丢掉。不丢的话它们会占满候选位,把「图片」「排序」这些真词全挤出去
  • 虚词表过滤 + 贪心去重叠(选了「相关性排序」就不再算「相关性」和「排序」)

3. 给文档打分 —— BM25。

每个词先算 IDF:在语料里越罕见越值钱   ln(1 + (N - df + 0.5) / (df + 0.5))
再按字段加权求和:标题 ×4  +  摘要 ×2  +  正文前 2500 字 ×1
每个字段都按 BM25 饱和 + 长度归一化(k1 = 1.2,b = 0.3 / 0.5 / 0.75)
再叠 10% 的时间新鲜度微调(主排序仍是相关性)

为什么是 BM25 而不是「命中次数 × 权重」(那是最初的做法,已换掉):

  • 没有 IDF 时,语料里到处都是的词(比如项目名)和罕见词一样值钱, 于是高频词不产生任何区分度,还稀释掉罕见词的分辨力
  • 没有长度归一化时,长文档靠堆词就能赢
  • 命中次数封顶 6 次是个手写硬拐点;k1 / b 才是为这件事设计的

Showing the opening section of the README — the full document lives in the repository以上为 README 开头摘要,完整文档在仓库内 · View the full README on GitHub →在 GitHub 查看完整 README →

← 上一个 Prev dsh-thinking-effort 下一个 Next dsh-files →