shinjiyu/dsh-plugin-multimodal 预览 preview

shinjiyu/dsh-plugin-multimodal

Plugin插件 Native原生 ⭐ 3 MIT Vision & Media视觉与多媒体

DeepSeek Harness 的视觉辅助工具:在纯文本模型上接受粘贴的图像。

Project Overview项目介绍

This repository holds a native plugin built exclusively for DeepSeek Harness (DSH), addressing a common pain point for DSH users who work with images in the official web interface. The official DSH release only supports text-only models, so any image pasted into the chat will be rejected with an error saying the current model does not support images. This plugin enables image acceptance by leveraging a separate vision-capable sidecar model to process images before sending content to the main text model. Installing the plugin is straightforward: you can run the dsh plugin command from GitHub or a local directory, then restart your dsh web instance to apply changes.

To configure the plugin, you need to set several environment variables that store your vision API base URL, API key, and the name of a vision-capable model that can process images. If you do not set the DSH-specific vision environment variables, the plugin will automatically fall back to any existing OpenAI environment variables you have already configured for your DSH instance. You cannot use a text-only model as the sidecar model, because it will not be able to extract text or describe content from the input image. You can also add configuration directly to your profile’s cordis.patch.yml file if you prefer that over setting environment variables.

After you finish installing and configuring the plugin, you can run a quick 30-second test to verify everything works as expected. Start by keeping your main model as the official text-only DeepSeek model, then confirm your sidecar vision model is properly configured. Paste an error screenshot into the DSH web chat and ask the model what the red text in the image says. If you do not get the "current model does not support images" error and the model answers your question correctly, the plugin is working properly. The plugin is released under the open source MIT license, and there is a community WeChat group for users to discuss issues and share feedback.

这是一个专为DeepSeek Harness开发的原生插件,解决了官方DSH网页界面粘贴图片会提示「当前模型不支持图片」从而被拒绝的问题。它通过额外调用一个具备视觉能力的sidecar模型,先将图片内容转换为文字再传递给主文本模型,实现图片内容的正常处理。用户可以通过dsh插件命令直接从GitHub或本地路径安装,安装后需要重启dsh web才能生效。

配置插件需要设置环境变量,包括视觉接口地址、API密钥以及具备视觉能力的模型名称,还可以自定义转换提示词。如果没有配置DSH专属的视觉变量,插件会自动回退使用已配置的OpenAI环境变量。需要注意,纯文本模型不能用作sidecar模型,否则无法完成图片文字提取工作,配置也可以写在profile的cordis.patch.yml文件中。

完成安装配置后,用户可以通过简单步骤快速验收功能是否正常:使用官方纯文本主模型,配置好能处理图片的sidecar模型,在网页中粘贴一张报错截图,询问模型图中红字内容,如果不再弹出不支持图片的提示,说明功能正常。插件基于MIT许可证开源,用户还可以加入社区微信群讨论使用问题。

Pre-install check安装前体检Compatibility · Security兼容性 · 安全性 1 warning1 项注意
  • Only 3 stars - very few users, little community feedback星标只有 3,几乎没人在用,遇到问题缺少社区反馈
DSH walks through these 9 checksDSH 会逐条核对这 9 项

Compatibility兼容性

  • DSH, Node, OS and profile requirementsDSH 版本 / Node 版本 / 操作系统 / profile 是否满足要求
  • External dependencies and runtimes (Electron / Python / Docker, ...)外部依赖与运行时(Electron / Python / Docker 等)是否齐备
  • Conflicts with installed plugins: command names, skill / tool names, ports, duplicate MCP registration与已装插件是否冲突:命令名、skill / tool 重名、端口占用、重复 MCP 注册

Security安全性

  • Repo matches the facts registered here; archived or abandoned?仓库是否与页面登记一致,是否归档或长期停更
  • Safety of preinstall / install / postinstall and install.sh / setup.ps1preinstall / install / postinstall 与 install.sh、setup.ps1 是否安全
  • curl|bash, download-then-execute, obfuscation, unrelated domains → stop immediatelycurl|bash、下载即执行、混淆代码、无关域名 → 立刻停止
  • Typosquatting or unmaintained packages among the new dependencies新增依赖里有没有 typosquatting 或无人维护的包
  • Requested permissions vs. what the feature actually needs申请了哪些权限、是否超出功能所需(filesystem / network / shell / clipboard)
  • Any sudo / admin requirement, plus uninstall and rollback是否要求 sudo / 管理员权限,以及卸载与回滚方式

Anything uncertain must be marked unknown with a note on how to confirm it. This site's signal screen is a static snapshot, not a security audit.拿不准的必须标「未知」并说明要我怎么确认。本站的信号筛查是静态快照,不能替代安全审计。

Or use CLI install (for developers)或使用命令行安装(适合开发者)

CLI Install命令行安装

dsh plugin --profile web add github:shinjiyu/dsh-plugin-multimodal

把 shinjiyu/dsh-plugin-multimodal 加入你的 DSH 配置(web profile)即可启用。

READMEREADME

dsh-plugin-multimodal

Awesome DSH Plugin

DeepSeek Harness 官方线路是纯文本。Web 里贴图会被拒:当前模型不支持图片。

这个插件补的是贴图准入,不是视觉工具箱。

  1. 主模型本身收图(Claude / GPT / 自建视觉网关)→ 原样把图交给模型,不转文字
  2. 主模型是纯文本(官方 DeepSeek)→ GUI 先收下图,sidecar 转成文字再发给主模型
  3. see_image 给磁盘上的截图用

官方 PI adapter 已经会做第 1 步,不会做第 2 步:不支持就 UNSUPPORTED_CONTENT。Anionex 的 dsh-vision-toolkit 是另一条路:给 agent 一堆 vision_* 工具,要模型自己去调。本插件让粘贴不被拒。

对应需求:Discussions #588 · 介绍帖:Show and tell #1709

安装

dsh plugin --profile web add github:shinjiyu/dsh-plugin-multimodal

本地路径:

dsh plugin --profile web add D:\tempWorkspace\dsh-plugin-multimodal

然后重启 dsh web。旧会话的工具表不会更新。

GitHub topic:dsh-plugin

配置

环境变量(不要把 key 写进仓库):

变量 作用
DSH_VISION_BASE_URL 视觉接口,例如 https://api.example/v1
DSH_VISION_API_KEY 该接口的 key
DSH_VISION_MODEL 必须是真能看图的模型,例如 glm-4.5v
DSH_VISION_PROMPT 可选。默认 OCR + 描述界面

没设 DSH_VISION_* 时回退 OPENAI_BASE_URL / OPENAI_API_KEY。GLM-5.2-FP8 这类文本模型不能当 sidecar。

也可以在 profile 的 cordis.patch.yml 里写:

- id: dsh-plugin-multimodal
  name: dsh-plugin-multimodal
  inject: [llm, tools, attachments, systemPrompt]
  config:
    model: glm-4.5v
    apiKeyEnv: DSH_VISION_API_KEY

30 秒验收

  1. 主模型仍用官方纯文本(或 PocketCity 文本模型)
  2. 配好一个真能看图的 sidecar
  3. 在 Web 里粘贴一张报错截图,问「红字说了什么」
  4. 过关:不再弹「当前模型不支持图片」,主模型能引用图上的字

磁盘文件:让它调 see_image。

不要做

  • 不要改 runtime/ 或官方 DeepSeek adapter
  • 不要把 sidecar 指到 deepseek-v4-flash / 其它纯文本模型
  • 不要把它宣传成「视觉工具箱」——那是 toolkit 的词

社区

微信群:deepseek harness 讨论群(2群)

扫码进群,用来分享和讨论插件。不是官方群。

微信群二维码

微信群码大约 7 天过期(本张到 2026-08-23)。过期后把新码覆盖 docs/community/wechat-group-2.jpg 即可。

宣发稿(Discussions / V2EX / 群公告)在 docs/PROMO.md。

License

MIT

← 上一个 Prev dsh-usage-meter 下一个 Next dsh-cot-summary →