Flyvhidbwo/dsh-vision-proxy
DeepSeek Harness插件:DeepSeek大脑+自动识图。GUI附加图片自动经OpenAI兼容VLM转译成文字后交给DeepSeek作答;支持百炼/智谱/OpenRouter等任意OpenAI兼容端点(默认qwen3.7-flash),无key自动探测本地Ollama(图片不出本机);安装时有一问式确认。
Project Overview项目介绍
dsh-vision-proxy is a native plugin built exclusively for DeepSeek Harness (DSH). It enables image support for text-only DeepSeek models like V4 Pro and the base Flash model, which DSH natively blocks image attachments for. The plugin adds a new deepseek-vision provider route that declares image input support to pass DSH's attachment pre-check, then translates any attached images to text via a configured visual language model (VLM) before passing the request to the original DeepSeek adapter. DeepSeek remains the core model for generating all final responses.
This plugin is targeted at DSH users who prefer to use DeepSeek's text-only flagship models while still needing to handle occasional image attachments. It supports multiple VLM providers including Alibaba Cloud Dashscope, Tongyi Qwen, Zhipu AI, OpenRouter, and any OpenAI-compatible VLM endpoint. It can also automatically detect a local running Ollama instance and add it to the fallback chain, enabling zero-configuration local image translation that never sends image data off your machine. Each endpoint in the fallback chain can point to a different provider for flexible configuration.
dsh-vision-proxy is released under the permissive MIT open-source license, and requires Node.js version 22.19 or newer to run. It includes multiple reliability-focused features: 20-second hard timeouts for anonymous endpoints, automatic 60-second cooldown for rate-limited or failed endpoints, and in-process caching of translation results by image content hash to avoid redundant re-translation. It defaults to inserting a clear placeholder text with error details when all VLM endpoints fail for an image, rather than aborting the entire conversation, which keeps your dialog flowing even when most VLM services are currently unavailable or unresponsive.
dsh-vision-proxy 是专为 DeepSeek Harness(DSH)开发的原生插件,核心作用是为纯文本版本的 DeepSeek 模型添加图片识别能力。DSH 原生会根据模型声明的输入模态拦截图片附件,本插件通过新增 deepseek-vision 路由,允许图片附件传入,并在请求发送给 DeepSeek 前调用任意 OpenAI 兼容视觉大模型将图片转译为文字。
本插件适合使用 DeepSeek V4 Pro 或普通纯文本 Flash 模型、又需要处理图片附件的 DSH 用户。工作流程为用户上传图片后,插件按配置顺序调用降级链上的视觉模型完成转译,转译结果以文字形式插入对话,最后依旧由原 DeepSeek 模型生成回复,全程保持 DeepSeek 作为对话的核心大脑。
本插件基于 MIT 许可证开源,要求 Node.js 版本不低于 22.19,支持通义千问百炼、智谱、OpenRouter 等多个厂商的视觉模型,也支持自动探测本地 Ollama 实现零配置本地图片转译(图片不离开本机)。转译结果会按图片内容哈希缓存在进程内,遇到限速或失败会自动跳过对应端点,不会卡住整轮对话。
请帮我安装这个 DSH 插件。安装前先完成【兼容性检查 + 安全性检查】,检查通过再动手。
插件:dsh-vision-proxy(Flyvhidbwo/dsh-vision-proxy)
仓库:https://github.com/Flyvhidbwo/dsh-vision-proxy
本站详情页:https://www.yhbd.top/plugins/flyvhidbwo-dsh-vision-proxy/
本站登记:类型 plugin · 归类 原生 DSH 插件 · 许可证 MIT · ⭐ 15 · 最近提交 2026-08-26 · 主语言 JavaScript
按下面顺序执行,每步先把结论告诉我,再进入下一步:
【1 兼容性检查】
① 我这边:DSH 版本、Node 版本、操作系统、当前 profile(web / desktop)。
② 读它的 README、package.json、插件 manifest,列出它要求的 DSH 版本 / Node 版本 / 操作系统 / 外部依赖 / 需要另外先装的运行时。
③ 逐条比对,结论只写「满足 / 不满足 / 未知」三种;不满足的给出可行替代方案。
④ 检查是否和我已装的插件冲突:命令名重复、skill / tool 重名、端口占用、重复注册的 MCP server。
【2 安全性检查】
① 仓库可信度:和上面「本站登记」是否一致;star / fork 数、创建时间、最近提交,是否归档或长期停更。
② 安装脚本:逐行看 package.json 的 preinstall / install / postinstall,以及 install.sh、setup.ps1 之类脚本。出现 curl|bash、下载后直接执行、混淆代码、访问与插件功能无关的域名,立刻停下来告诉我,不要继续装。
③ 依赖:列出新增依赖,标出无人维护、或与知名包拼写近似的可疑包(typosquatting)。
④ 权限与副作用:它会读写哪些目录、访问哪些域名、需要哪些 DSH 权限(filesystem / network / shell / clipboard 等),以及怎么卸载和回滚。
⑤ 如果它要求 sudo / 管理员权限,或权限明显超出功能所需,先停下来问我。
【3 安装】
上面两步没有「不满足」和「高危项」时才执行;用官方推荐方式安装,不要自行提权。
【4 汇报】
用表格输出:检查项 / 结论 / 依据 / 是否需要我决策。拿不准的一律写「未知」并说明要我怎么确认——不要猜,也不要替我决定。
Send this message to DSH in your current session: it verifies compatibility and security first (answering met / not met / unknown item by item) and only installs once everything checks out — it will stop and ask you if it finds a high-risk item. The box scrolls; the copy is the full prompt. CLI install commands may not be accurate across systems, so DSH is the safer route.把上面这条消息直接发给当前会话里的 DSH:它会先核对兼容性与安全性(逐条给「满足 / 不满足 / 未知」),确认没问题再安装,有高危项会停下来问你。框内可滚动,复制到的是完整提示词;安装命令不一定准确,发给 DSH 更稳。
- 15 stars - an early-stage project星标 15,属于早期项目
DSH walks through these 9 checksDSH 会逐条核对这 9 项
Compatibility兼容性
- DSH, Node, OS and profile requirementsDSH 版本 / Node 版本 / 操作系统 / profile 是否满足要求
- External dependencies and runtimes (Electron / Python / Docker, ...)外部依赖与运行时(Electron / Python / Docker 等)是否齐备
- Conflicts with installed plugins: command names, skill / tool names, ports, duplicate MCP registration与已装插件是否冲突:命令名、skill / tool 重名、端口占用、重复 MCP 注册
Security安全性
- Repo matches the facts registered here; archived or abandoned?仓库是否与页面登记一致,是否归档或长期停更
- Safety of preinstall / install / postinstall and install.sh / setup.ps1preinstall / install / postinstall 与 install.sh、setup.ps1 是否安全
- curl|bash, download-then-execute, obfuscation, unrelated domains → stop immediatelycurl|bash、下载即执行、混淆代码、无关域名 → 立刻停止
- Typosquatting or unmaintained packages among the new dependencies新增依赖里有没有 typosquatting 或无人维护的包
- Requested permissions vs. what the feature actually needs申请了哪些权限、是否超出功能所需(filesystem / network / shell / clipboard)
- Any sudo / admin requirement, plus uninstall and rollback是否要求 sudo / 管理员权限,以及卸载与回滚方式
Anything uncertain must be marked unknown with a note on how to confirm it. This site's signal screen is a static snapshot, not a security audit.拿不准的必须标「未知」并说明要我怎么确认。本站的信号筛查是静态快照,不能替代安全审计。
Or use CLI install (for developers)或使用命令行安装(适合开发者)
CLI Install命令行安装
dsh plugin --profile web add dsh-vision-proxy
把 Flyvhidbwo/dsh-vision-proxy 加入你的 DSH 配置(web profile)即可启用。
READMEREADME
dsh-vision-proxy
保持 DeepSeek 作为对话大脑,图片照样直接发。 为 DeepSeek Harness 打造:GUI 附加图片自动转译,纯文本 DeepSeek 也能识图。
⚠️ 兼容性与定位(2026-08):本插件已适配 dsh 0.1.1-rc.2(adapter prepareCall 接口)。dsh 0.1.1 起原生支持多模态(DeepSeek-V4-Flash-Vision-Exp 等官方视觉模型)——如果你用官方视觉模型,直接发图即可,不需要本插件。DeepSeek-V4-Pro / 普通 Flash 仍是纯文本模型,识图靠本插件转译桥接(官方仅 Flash-Vision-Exp 原生多模态)。插件适用于:Pro/文本模型识图、本地 Ollama(图片不出本机、免费)、自定义 OpenAI 兼容 VLM 场景。
为什么需要它
DeepSeek Harness 原生按模型声明的 inputModalities 决定是否放行图片附件。DeepSeek 的 chat-completions 线路是纯文本的,所以选中 DeepSeek 时附加图片会被原生拒绝。已有的视觉插件提供 view_image 等工具(适用于文件路径),但 GUI 图片附件对纯文本模型依然失败。
本插件补上这个缺口:注册一条新提供商路由(deepseek-vision),包装真正的 DeepSeek 适配器——对外声明支持图片输入(附件预检放行),并在请求流里把每张附加图片转译成文字后再委托给 DeepSeek。对话仍然由 DeepSeek 作答,识图只是附加能力。
用户附加图片 ──▶ deepseek-vision 路由 ──▶ 经 VLM 转译(OCR+版式+细节)
│ │
▼ ▼
DeepSeek 作答 ◀── 纯文本对话(图片已替换为 [图片转译] 文字)
特性
- 绝不卡死。匿名端点强制 20 秒超时上限(免费档挂起也拖不住整轮对话);匿名端点遇到 HTTP 429 立即失败(不做无意义的 Retry-After 等待);刚失败(429/超时)的端点进入 60 秒冷却并被跳过。
- 多模型、多厂商。任何 OpenAI 兼容 VLM 端点都行——百炼/Qwen、QwenCloud 国际站、智谱、OpenRouter、本地 Ollama、或你自己的端点。每条
fallbackModels都可以带各自独立的baseURL/model,一个安装即可串联多家。 - 零配置本地路径。
autoLocalOllama(默认开)启动时探测http://localhost:11434,检测到 Ollama 就自动加入降级链——图片不出本机,免 key 免注册。 - 快速且明确的失败。没有 key 也没有本地 Ollama 时,转译在几秒内失败并给出可操作指引(配置
VISION_API_KEY/DASHSCOPE_API_KEY或安装 Ollama)——绝不静默卡住。 - 有 key 自动提速。导出
VISION_API_KEY/DASHSCOPE_API_KEY后自动走你配置的付费端点(默认百炼qwen3.7-flash——快、便宜、不限速;百炼/QwenCloud/智谱/OpenRouter 或任意 OpenAI 兼容端点均可);没有 key 的条目会被跳过而不是失败。 - 安装时一问式确认。
postinstall询问你是否有 VLM API key。非交互环境自动跳过,安装永不卡死。启动时打印 PRIVACY NOTICE 标明当前使用的端点。 - 降级链 + 错误分类。
rate_limit/quota/auth/region/model_not_found/context_too_large/http分类给出可操作提示。 - 内容哈希缓存。转译结果按图片字节的 SHA-256 缓存(进程内,上限 200)——同一张图每个进程最多转译一次,重新附加或换对话也命中。
- 自动降采样(可选)。装有
sharp时,超过maxImagePixels的图片转译前自动缩小——大截图更快;没有 sharp 则优雅降级原图直发。 - 兼容
read_image。原生read_image工具在该路由下同样可用(它的能力门禁读取同一份模型信息)。
Showing the opening section of the README — the full document lives in the repository以上为 README 开头摘要,完整文档在仓库内 · View the full README on GitHub →在 GitHub 查看完整 README →
WNJXYK/dsh-codex-oauth
moduqishi/GrassVison
Totoro-qaq/dsh-plugin-bridge
william-jin-cmu/dsh-vision
rinDBeans/dsh-miraculous-standard
xiaoxianyu-office/dsh-router-flash
v587d/dsh-opencode-go-usage