xiaoxianyu-office/dsh-image-tools

DSH捆绑插件:聊天图像桥接+拒绝读取图像+为纯文本主模型提供对话式图像识别 | 纯文本主模型图像识别桥接与识别工具

Project Overview项目介绍

This is a native plugin built exclusively for DeepSeek Harness (DSH) that enables image recognition capabilities for pure-text large language models like deepseek-v4-pro and deepseek-v4-flash. It automatically saves uploaded images to the working directory, dynamically disables the native read_image tool to prevent 400 errors when pure-text models receive image blocks, and provides a conversational image_recognize tool that offloads vision tasks to a dedicated vision sub-agent. The plugin works globally across all DSH sessions and supports four preset configurations to match different user workflows, from general use to code-focused work and minimal setups.

This plugin uses a dynamic bridging logic instead of hardcoding a fixed routing list. It determines whether to intercept image requests based on the actual capability declaration of the current session’s model. If a model’s image capability comes from an explicit declaration in your DSH settings.yaml, the plugin enables bridging for that model. If the model is a native multimodal model already included in the DSH model catalog, it uses the native multimodal flow without any extra processing or interception. To disable bridging for a model, you simply remove the image capability declaration from your settings, making it easy to toggle the feature on or off for different models.

Before installing the plugin, you need three prerequisites: a global installation of pnpm via npm, a preconfigured model route for your vision sub-agent in your DSH settings, and a valid API key for your vision sub-agent stored in your DSH credentials file. After the initial installation, you must restart your DSH web service for changes to take effect. The first time you run DSH after installing the plugin, a configuration wizard will automatically pop up to check your routing, API key, and model connectivity. If you encounter issues while using the plugin, it will automatically open a troubleshooting panel to help you diagnose and fix common problems. The plugin is released under the MIT license, with no costs for use.

这是一款专为DeepSeek Harness开发的原生插件,核心作用是为纯文本大模型(如deepseek-v4-pro、deepseek-v4-flash)添加识图能力。它具备聊天发图自动落盘、原生read_image工具动态禁用、对话式image_recognize识图三个核心功能,识图任务会委派给用户指定的视觉子Agent完成。插件全局生效,支持standard、code、minimal、cordis四种预设会话配置,适配不同使用场景。

它采用动态桥接逻辑,不会预先写死路由列表,而是根据当前会话模型的能力声明动态判断是否拦截。如果模型的图片能力来自用户settings.yaml设置中的显式声明,就会启用桥接;如果是目录自带的原生多模态模型,则直接走原生链路,不做额外拦截处理。用户如果要关闭某模型的桥接,只需移除设置中的图片能力声明即可,操作灵活,适合需要用纯文本模型处理图片的DSH用户使用。

安装该插件需要提前配置pnpm、模型路由和识图子Agent的API密钥,通过DSH的CLI命令安装,安装后需要重启DSH Web服务才能生效。首次安装后会自动弹出配置向导,帮助用户检查路由、API密钥和模型连通性,使用中出现故障也会自动弹出自检面板。插件采用MIT开源许可证,无使用成本,卸载后除用户设置外无残留文件。

Pre-install check安装前体检Compatibility · Security兼容性 · 安全性 1 warning1 项注意
  • Only 2 stars - very few users, little community feedback星标只有 2,几乎没人在用,遇到问题缺少社区反馈
DSH walks through these 9 checksDSH 会逐条核对这 9 项

Compatibility兼容性

  • DSH, Node, OS and profile requirementsDSH 版本 / Node 版本 / 操作系统 / profile 是否满足要求
  • External dependencies and runtimes (Electron / Python / Docker, ...)外部依赖与运行时(Electron / Python / Docker 等)是否齐备
  • Conflicts with installed plugins: command names, skill / tool names, ports, duplicate MCP registration与已装插件是否冲突:命令名、skill / tool 重名、端口占用、重复 MCP 注册

Security安全性

  • Repo matches the facts registered here; archived or abandoned?仓库是否与页面登记一致,是否归档或长期停更
  • Safety of preinstall / install / postinstall and install.sh / setup.ps1preinstall / install / postinstall 与 install.sh、setup.ps1 是否安全
  • curl|bash, download-then-execute, obfuscation, unrelated domains → stop immediatelycurl|bash、下载即执行、混淆代码、无关域名 → 立刻停止
  • Typosquatting or unmaintained packages among the new dependencies新增依赖里有没有 typosquatting 或无人维护的包
  • Requested permissions vs. what the feature actually needs申请了哪些权限、是否超出功能所需(filesystem / network / shell / clipboard)
  • Any sudo / admin requirement, plus uninstall and rollback是否要求 sudo / 管理员权限,以及卸载与回滚方式

Anything uncertain must be marked unknown with a note on how to confirm it. This site's signal screen is a static snapshot, not a security audit.拿不准的必须标「未知」并说明要我怎么确认。本站的信号筛查是静态快照,不能替代安全审计。

Or use CLI install (for developers)或使用命令行安装(适合开发者)

CLI Install命令行安装

dsh plugin --profile web add -w github:xiaoxianyu-office/dsh-image-tools#v0.3.7

把 xiaoxianyu-office/dsh-image-tools 加入你的 DSH 配置(web profile)即可启用。

READMEREADME

dsh-image-tools

让纯文本模型(deepseek-v4-pro / flash)具备识图能力的 DSH 插件包: 聊天发图自动落盘 + 原生 read_image 动态禁用 + 对话式 image_recognize 识图工具(委派视觉子 agent,如 xiaomi/mimo-v2.5)。

宿主层全局生效,四个 preset(standard / code / minimal / cordis)的会话通用。

动态桥接原理(v0.2.0)

不再配置任何路由列表。插件按会话当前模型的真实声明动态判定是否拦截:

  • 桥接条件:模型解析出 image 能力,且该能力来自你在 settings.yaml 里的显式声明 (路由 defaultInput、模型条目 input、或 modelOverrides 的 input 含 image)。 手写 image 声明的唯一用途就是放行上传准入,所以声明即桥接意图。
  • 原生条件:模型能力来自目录(catalog,如 qwen3.6-plus / grok-4.5 / mimo-v2.5), 无需任何声明 → 原生多模态链路,不转存、不拦截 read_image。

规则:真多模态模型不要手写 input 声明(目录已提供);移除声明即关闭该模型的桥接 (恢复纯文本,上传会被准入拒绝)。

安装

前置条件:

  1. dsh plugin 需要 pnpm:npm i -g pnpm
  2. 模型路由(插件只挂载插件行,路由在设置层,需已存在):
# ~/.dsh/settings.yaml
llm-pi-ai:
  providers:
    xiaomi:
      displayName: 识图模型(MiMo)   # 识图模型分组(目录原生多模态)
      apiKeyEnv: XIAOMI_API_KEY      # key 可自由更换
      baseURL: https://opencode.ai/zen/go/v1
      models:
        - id: mimo-v2.5
          name: MiMo-V2.5
    opencode-go:
      apiKeyEnv: OPENCODE_GO_API_KEY
      modelOverrides:               # 只覆写这两个模型,其余目录模型(qwen/grok)保持原生
        deepseek-v4-pro:
          input: [ text, image ]    # 桥接声明:仅用于放行上传准入
        deepseek-v4-flash:
          input: [ text, image ]
agent-default-model:
  provider: opencode-go
  model: deepseek-v4-flash
  1. 识图 token:~/.dsh/.credentials.yaml 中 XIAOMI_API_KEY

安装(始终使用最新发布 tag,见仓库 Releases;示例为当前最新 v0.3.7):

dsh plugin --profile web add -w github:xiaoxianyu-office/dsh-image-tools#v0.3.7

安装后重启 dsh web 服务生效(插件代码在进程内)。

升级(始终切到最新 tag)

升级 = 重复 add 并指定最新的 tag,不要用 update 选择 Git 引用:

dsh plugin --profile web add -w github:xiaoxianyu-office/dsh-image-tools#v0.3.7

卸载

dsh plugin --profile web remove @dsh-external/dsh-image-tools

卸载后重启服务。插件层(依赖、node_modules、组合行)无残留; settings.yaml 里的路由与默认模型属于设置层,需手动还原(见上「前置条件」反向操作)。 另外 ~/.dsh/image-tools-state.json(向导完成标记)为可选清理项。

行为

  • read_image 动态禁用:桥接模型(如 deepseek)调用原生 read_image 直接返回 「已禁用,请改用 image_recognize」——防止图片块进入纯文本端点请求导致 400; 目录原生多模态模型(qwen/grok/mimo)不受影响,可正常使用 read_image;
  • 发图桥接:桥接模型会话上传图片自动落盘 <工作区>/uploads/,消息中显示 [图片] 文件名; 原生多模态会话的图片直接进入模型,不做任何处理;
  • image_recognize:必须传针对性读取任务(想从图中获得什么);同一图片路径再次调用 自动衔接此前问答,可持续追问;
  • 视觉子 agent(xiaomi/mimo-v2.5,无声明)不受 read_image 禁用影响;
  • 识图输出严格规范(v0.3.2):识图子 agent 只回答任务问题,位置给像素坐标或明确方位、颜色给 #RRGGBB 色值,禁止模糊词(偏上/大概/差不多/看起来等),图中没有的内容回答「图中未出现」,不确定回答「无法从图中确认」并说明原因。

Showing the opening section of the README — the full document lives in the repository以上为 README 开头摘要,完整文档在仓库内 · View the full README on GitHub →在 GitHub 查看完整 README →

← 上一个 Prev dsh-usage-dashboard 下一个 Next dsh-update-checker →