Flyvhidbwo/dsh-vision-proxy 预览 preview

Flyvhidbwo/dsh-vision-proxy

DeepSeek Harness插件:DeepSeek大脑+自动识图。GUI附加图片自动经OpenAI兼容VLM转译成文字后交给DeepSeek作答;支持百炼/智谱/OpenRouter等任意OpenAI兼容端点(默认qwen3.7-flash),无key自动探测本地Ollama(图片不出本机);安装时有一问式确认。

Project Overview项目介绍

dsh-vision-proxy is a native plugin built exclusively for DeepSeek Harness (DSH). It enables image support for text-only DeepSeek models like V4 Pro and the base Flash model, which DSH natively blocks image attachments for. The plugin adds a new deepseek-vision provider route that declares image input support to pass DSH's attachment pre-check, then translates any attached images to text via a configured visual language model (VLM) before passing the request to the original DeepSeek adapter. DeepSeek remains the core model for generating all final responses.

This plugin is targeted at DSH users who prefer to use DeepSeek's text-only flagship models while still needing to handle occasional image attachments. It supports multiple VLM providers including Alibaba Cloud Dashscope, Tongyi Qwen, Zhipu AI, OpenRouter, and any OpenAI-compatible VLM endpoint. It can also automatically detect a local running Ollama instance and add it to the fallback chain, enabling zero-configuration local image translation that never sends image data off your machine. Each endpoint in the fallback chain can point to a different provider for flexible configuration.

dsh-vision-proxy is released under the permissive MIT open-source license, and requires Node.js version 22.19 or newer to run. It includes multiple reliability-focused features: 20-second hard timeouts for anonymous endpoints, automatic 60-second cooldown for rate-limited or failed endpoints, and in-process caching of translation results by image content hash to avoid redundant re-translation. It defaults to inserting a clear placeholder text with error details when all VLM endpoints fail for an image, rather than aborting the entire conversation, which keeps your dialog flowing even when most VLM services are currently unavailable or unresponsive.

dsh-vision-proxy 是专为 DeepSeek Harness(DSH)开发的原生插件,核心作用是为纯文本版本的 DeepSeek 模型添加图片识别能力。DSH 原生会根据模型声明的输入模态拦截图片附件,本插件通过新增 deepseek-vision 路由,允许图片附件传入,并在请求发送给 DeepSeek 前调用任意 OpenAI 兼容视觉大模型将图片转译为文字。

本插件适合使用 DeepSeek V4 Pro 或普通纯文本 Flash 模型、又需要处理图片附件的 DSH 用户。工作流程为用户上传图片后,插件按配置顺序调用降级链上的视觉模型完成转译,转译结果以文字形式插入对话,最后依旧由原 DeepSeek 模型生成回复,全程保持 DeepSeek 作为对话的核心大脑。

本插件基于 MIT 许可证开源,要求 Node.js 版本不低于 22.19,支持通义千问百炼、智谱、OpenRouter 等多个厂商的视觉模型,也支持自动探测本地 Ollama 实现零配置本地图片转译(图片不离开本机)。转译结果会按图片内容哈希缓存在进程内,遇到限速或失败会自动跳过对应端点,不会卡住整轮对话。

Pre-install check安装前体检Compatibility · Security兼容性 · 安全性 1 note1 项提示
  • 15 stars - an early-stage project星标 15,属于早期项目
DSH walks through these 9 checksDSH 会逐条核对这 9 项

Compatibility兼容性

  • DSH, Node, OS and profile requirementsDSH 版本 / Node 版本 / 操作系统 / profile 是否满足要求
  • External dependencies and runtimes (Electron / Python / Docker, ...)外部依赖与运行时(Electron / Python / Docker 等)是否齐备
  • Conflicts with installed plugins: command names, skill / tool names, ports, duplicate MCP registration与已装插件是否冲突:命令名、skill / tool 重名、端口占用、重复 MCP 注册

Security安全性

  • Repo matches the facts registered here; archived or abandoned?仓库是否与页面登记一致,是否归档或长期停更
  • Safety of preinstall / install / postinstall and install.sh / setup.ps1preinstall / install / postinstall 与 install.sh、setup.ps1 是否安全
  • curl|bash, download-then-execute, obfuscation, unrelated domains → stop immediatelycurl|bash、下载即执行、混淆代码、无关域名 → 立刻停止
  • Typosquatting or unmaintained packages among the new dependencies新增依赖里有没有 typosquatting 或无人维护的包
  • Requested permissions vs. what the feature actually needs申请了哪些权限、是否超出功能所需(filesystem / network / shell / clipboard)
  • Any sudo / admin requirement, plus uninstall and rollback是否要求 sudo / 管理员权限,以及卸载与回滚方式

Anything uncertain must be marked unknown with a note on how to confirm it. This site's signal screen is a static snapshot, not a security audit.拿不准的必须标「未知」并说明要我怎么确认。本站的信号筛查是静态快照,不能替代安全审计。

Or use CLI install (for developers)或使用命令行安装(适合开发者)

CLI Install命令行安装

dsh plugin --profile web add dsh-vision-proxy

把 Flyvhidbwo/dsh-vision-proxy 加入你的 DSH 配置(web profile)即可启用。

READMEREADME

dsh-vision-proxy

English | 简体中文

保持 DeepSeek 作为对话大脑,图片照样直接发。 为 DeepSeek Harness 打造:GUI 附加图片自动转译,纯文本 DeepSeek 也能识图。

awesome · DSH plugin npm version CI (Node 22/24) 14 项测试通过 MIT Node >=22.19 GitHub stars

⚠️ 兼容性与定位(2026-08):本插件已适配 dsh 0.1.1-rc.2(adapter prepareCall 接口)。dsh 0.1.1 起原生支持多模态(DeepSeek-V4-Flash-Vision-Exp 等官方视觉模型)——如果你用官方视觉模型,直接发图即可,不需要本插件。DeepSeek-V4-Pro / 普通 Flash 仍是纯文本模型,识图靠本插件转译桥接(官方仅 Flash-Vision-Exp 原生多模态)。插件适用于:Pro/文本模型识图、本地 Ollama(图片不出本机、免费)、自定义 OpenAI 兼容 VLM 场景。

为什么需要它

DeepSeek Harness 原生按模型声明的 inputModalities 决定是否放行图片附件。DeepSeek 的 chat-completions 线路是纯文本的,所以选中 DeepSeek 时附加图片会被原生拒绝。已有的视觉插件提供 view_image 等工具(适用于文件路径),但 GUI 图片附件对纯文本模型依然失败。

本插件补上这个缺口:注册一条新提供商路由(deepseek-vision),包装真正的 DeepSeek 适配器——对外声明支持图片输入(附件预检放行),并在请求流里把每张附加图片转译成文字后再委托给 DeepSeek。对话仍然由 DeepSeek 作答,识图只是附加能力。

用户附加图片 ──▶ deepseek-vision 路由 ──▶ 经 VLM 转译(OCR+版式+细节)
                   │                        │
                   ▼                        ▼
            DeepSeek 作答 ◀── 纯文本对话(图片已替换为 [图片转译] 文字)

特性

  • 绝不卡死。匿名端点强制 20 秒超时上限(免费档挂起也拖不住整轮对话);匿名端点遇到 HTTP 429 立即失败(不做无意义的 Retry-After 等待);刚失败(429/超时)的端点进入 60 秒冷却并被跳过。
  • 多模型、多厂商。任何 OpenAI 兼容 VLM 端点都行——百炼/Qwen、QwenCloud 国际站、智谱、OpenRouter、本地 Ollama、或你自己的端点。每条 fallbackModels 都可以带各自独立的 baseURL/model,一个安装即可串联多家。
  • 零配置本地路径。autoLocalOllama(默认开)启动时探测 http://localhost:11434,检测到 Ollama 就自动加入降级链——图片不出本机,免 key 免注册。
  • 快速且明确的失败。没有 key 也没有本地 Ollama 时,转译在几秒内失败并给出可操作指引(配置 VISION_API_KEY / DASHSCOPE_API_KEY 或安装 Ollama)——绝不静默卡住。
  • 有 key 自动提速。导出 VISION_API_KEY / DASHSCOPE_API_KEY 后自动走你配置的付费端点(默认百炼 qwen3.7-flash——快、便宜、不限速;百炼/QwenCloud/智谱/OpenRouter 或任意 OpenAI 兼容端点均可);没有 key 的条目会被跳过而不是失败。
  • 安装时一问式确认。postinstall 询问你是否有 VLM API key。非交互环境自动跳过,安装永不卡死。启动时打印 PRIVACY NOTICE 标明当前使用的端点。
  • 降级链 + 错误分类。rate_limit / quota / auth / region / model_not_found / context_too_large / http 分类给出可操作提示。
  • 内容哈希缓存。转译结果按图片字节的 SHA-256 缓存(进程内,上限 200)——同一张图每个进程最多转译一次,重新附加或换对话也命中。
  • 自动降采样(可选)。装有 sharp 时,超过 maxImagePixels 的图片转译前自动缩小——大截图更快;没有 sharp 则优雅降级原图直发。
  • 兼容 read_image。原生 read_image 工具在该路由下同样可用(它的能力门禁读取同一份模型信息)。

Showing the opening section of the README — the full document lives in the repository以上为 README 开头摘要,完整文档在仓库内 · View the full README on GitHub →在 GitHub 查看完整 README →

← 上一个 Prev fnos-dsh 下一个 Next dsh-ponytail →