GOU-GEE/deepseek-vision

Plugin插件 ⭐ 4 MIT Vision & Media视觉与多媒体

Project Overview项目介绍

This repository hosts an open-source Model Context Protocol (MCP) server that adds image recognition capability to pure-text large language models like DeepSeek, and it ships with a pre-built native plugin bundle for DeepSeek Harness. The core function is exposed through the analyze_image tool, which routes input images to OpenAI-compatible third-party vision model APIs such as Zhipu GLM-4.6V, Qwen2.5-VL, and Tongyi Qianwen VL, then passes the text recognition result back to the calling main model. It works with DSH natively, and also supports other MCP-compatible AI agents including Claude Code, Codex, and OpenCode across all major desktop platforms.

It is built primarily for DSH users who want to add visual capability to their pure-text DeepSeek models, without having to switch to a different multimodal base model. When paired with its included Skill file, the main DeepSeek model will automatically trigger the image analysis tool whenever it encounters an image input, so users do not need to manually write tool call commands. For DSH desktop users, the project supports one-session automated deployment, and after installation, users can configure their API key through a built-in visual settings page that protects key confidentiality. Keys are stored in DSH’s official credential manager instead of plaintext configuration files.

The project is released under the permissive MIT open-source license, and it is free for anyone to use, modify, and redistribute, with no paid features locked behind a paywall. It requires Python 3.10 or newer to run the MCP server, and DSH native deployment additionally requires Node.js 22.19 or newer. Users must obtain their own API key from a supported third-party vision provider, and any image you analyze will be sent to your selected provider, so you should never upload sensitive or classified images to this service. The default recommended Zhipu GLM-4.6V-flash model offers a free tier for new users, so you can test the service without upfront payment.

这是一款为纯文本大模型添加图像识别能力的开源MCP服务器,同时提供了预构建的DeepSeek Harness原生插件包。它对外暴露analyze_image工具,可将图片发送给兼容OpenAI格式的第三方视觉大模型API(如智谱GLM-4.6V、通义千问Qwen-VL等)完成识别,再将识别后的文本结果返回给调用它的主模型。它不仅支持DSH,也可接入Claude Code、Codex、OpenCode等其他兼容MCP协议的AI客户端。

它主要面向需要让纯文本DeepSeek模型获得识图能力的DSH用户,配合自带的Skill文件,主模型遇到图片时会自动触发工具调用,无需用户手动输入指令。在DSH桌面版中,支持一键自动化部署,安装后用户可直接在可视化配置页面填写API密钥,密钥由DSH官方凭据系统托管,不会明文暴露或存储在公共配置文件中。普通MCP客户端也可按照通用步骤手动配置部署。

本项目采用MIT许可证开源,完全免费使用,用户只需自行申请第三方视觉API的密钥即可,默认推荐使用智谱GLM-4.6V-flash,该模型提供免费额度。项目要求Python版本不低于3.10,DSH部署还要求Node版本不低于22.19。使用时请注意,图片会发送到你选择的第三方服务商,请勿传入敏感涉密图片,项目本身仅在本地运行,不会收集你的密钥或图片数据。

Pre-install check安装前体检Compatibility · Security兼容性 · 安全性 2 warnings2 项注意
  • Only 4 stars - very few users, little community feedback星标只有 4,几乎没人在用,遇到问题缺少社区反馈
  • No DSH plugin manifest detected - it may only carry the dsh-plugin topic, so the install method must be confirmed on the spot未检测到 DSH 插件清单:可能只是打了 dsh-plugin 话题,安装方式要现场确认
  • Not DSH-native: a multi-platform tool that may require Node / Electron or another runtime first非 DSH 原生,是多平台兼容工具:可能要先装 Node / Electron 等运行时
DSH walks through these 9 checksDSH 会逐条核对这 9 项

Compatibility兼容性

  • DSH, Node, OS and profile requirementsDSH 版本 / Node 版本 / 操作系统 / profile 是否满足要求
  • External dependencies and runtimes (Electron / Python / Docker, ...)外部依赖与运行时(Electron / Python / Docker 等)是否齐备
  • Conflicts with installed plugins: command names, skill / tool names, ports, duplicate MCP registration与已装插件是否冲突:命令名、skill / tool 重名、端口占用、重复 MCP 注册

Security安全性

  • Repo matches the facts registered here; archived or abandoned?仓库是否与页面登记一致,是否归档或长期停更
  • Safety of preinstall / install / postinstall and install.sh / setup.ps1preinstall / install / postinstall 与 install.sh、setup.ps1 是否安全
  • curl|bash, download-then-execute, obfuscation, unrelated domains → stop immediatelycurl|bash、下载即执行、混淆代码、无关域名 → 立刻停止
  • Typosquatting or unmaintained packages among the new dependencies新增依赖里有没有 typosquatting 或无人维护的包
  • Requested permissions vs. what the feature actually needs申请了哪些权限、是否超出功能所需(filesystem / network / shell / clipboard)
  • Any sudo / admin requirement, plus uninstall and rollback是否要求 sudo / 管理员权限,以及卸载与回滚方式

Anything uncertain must be marked unknown with a note on how to confirm it. This site's signal screen is a static snapshot, not a security audit.拿不准的必须标「未知」并说明要我怎么确认。本站的信号筛查是静态快照,不能替代安全审计。

Or use CLI install (for developers)或使用命令行安装(适合开发者)

CLI Install命令行安装

npx -y @deepseek-ai/dsh@0.1.0-rc.6 plugin --profile web add dsh-plugin-deepseek-vision@0.4.2

把 GOU-GEE/deepseek-vision 加入你的 DSH 配置(web profile)即可启用。

READMEREADME

deepseek-vision-mcp

CI Publish PyPI version npm version Python 3.10+ License: MIT Awesome DSH Plugin

给 DeepSeek(及其他纯文本大模型)装上「眼睛」 的开源 MCP Server。

DeepSeek 系列模型是纯文本模型,无法直接识别图片。本项目的思路是: 通过 MCP Server 暴露一个 analyze_image 工具,把图片交给第三方 OpenAI 兼容的视觉模型 API(智谱 GLM-4.6V、硅基流动 Qwen2.5-VL、 通义千问 qwen-vl-plus 等)识别,再把识别文本返回给主模型。 配合项目自带的 Skill 文件,DeepSeek 主模型在遇到图片时会自动调用 该工具——对用户来说,就像 DeepSeek 突然会「看图」了。

视觉识别能力由第三方模型提供,图片会被发送到对应服务商的 API。 请阅读文末的隐私说明。


🚀 一句话安装(复制给 AI 助手,免手动操作)

主要在 DeepSeek Harness 中使用时,复制下面的推荐提示词给助手。它会从全新克隆构建 本仓库的 DSH Bundle、安装到桌面版共用的 Web profile,并引导你在 DSH 可视化页面中 安全填写 Key,最后完成自动托管 Python 与真实图片验收。 其他 MCP 客户端请使用后面的通用版。

关于 API Key(请先读):

  • ✅ 默认方案使用智谱 glm-4.6v-flash:安装后在 DSH 的 设置 → 插件 → 插件配置 → DeepSeek Vision 中填写你在 https://open.bigmodel.cn 申请的 Key。
  • ⚠️ Key、模型和 Base URL 必须属于同一家服务商。硅基流动或通义千问的 Key 也能用, 但必须同时把对应的 VISION_MODEL 与 VISION_BASE_URL 告诉助手,不能套用智谱默认值。
  • 🔒 免费或付费 Key 都不必发给 AI 助手:可视化页面不会回显 Key,凭据由 DSH 官方凭据存储保存,不会写进仓库或普通配置文件。

DeepSeek Harness 推荐版(一会话从零部署并验收):

请在一个会话内从零安装并验收 https://github.com/GOU-GEE/deepseek-vision 的 DSH
插件,让 macOS DeepSeek Harness 桌面版中的文本模型获得视觉能力(默认智谱免费模型
glm-4.6v-flash):
1. 在全新的、不复用旧仓库或虚拟环境的目录克隆 main,确认 Node >=22.19、Python >=3.10,
   并运行 corepack enable、corepack prepare pnpm@11.7.0 --activate。
2. 进入 plugins/dsh-plugin-deepseek-vision,运行 npm ci --ignore-scripts、npm test,
   再以可用的 Python 设置 VISION_BUILD_PYTHON 并运行 npm pack;检查 tarball 内含 LICENSE
   和 runtime/deepseek_vision_mcp-<pyproject 中的版本>-py3-none-any.whl
   (当前 main 为 deepseek_vision_mcp-0.4.2-py3-none-any.whl)。
3. 确认 `/Applications/DeepSeek Harness.app` 存在并读取它的实际 DSH 版本;先让我用
   Cmd+Q 完全退出桌面版,再使用 App 内置的 DSH CLI,把源码构建的 tarball 合并安装到
   默认 `web` profile。不要尝试安装尚未发布的 npm 版本,不覆盖整个 profile。用同一
   CLI 的 `--profile web --dump-config` 确认 deepseek-vision-host、deepseek-vision-mcp
   和插件内置 launcher.js 均已加载;若不是 macOS 桌面版,再回退到同版本 npx dsh 命令。
4. 让我重新打开 DeepSeek Harness,进入
   `设置 → 插件 → 插件配置 → DeepSeek Vision`,选择“智谱 GLM(推荐免费)”,确认模型
   自动为 glm-4.6v-flash、Base URL 自动为 https://open.bigmodel.cn/api/paas/v4;让我亲自
   在密码框填写 Key,点击“保存”和“测试连接”。不得向我索取或读取 Key,测试连接失败时
   展示页面错误。保存后按页面提示 Cmd+Q 并重新打开,使 MCP 使用新配置。
5. 从仓库根目录运行 python scripts/verify_dsh_plugin.py,必须通过真实 MCP stdio 握手、
   4 个工具和 vision_status;然后在重新打开的桌面版中选择文本版 DeepSeek,把
   examples/test_image.jpg 粘贴或拖入输入框,确认输入框上方出现不超过输入框宽度的
   缩略图卡带、输入框内不显示工具指令长文本;不输入文字直接发送时,确认自动注入
   analyze_image 指令且工具实际调用成功、返回 model=glm-4.6v-flash,发送后缩略图卡带
   自动关闭。再发同图同问题,确认 cached=true。
6. 每一步失败立即停止并展示完整错误;配置只做合并,不覆盖用户原有 DSH profile,
   不删除旧安装。最后汇报克隆路径、tarball、DSH profile、托管运行时路径及全部验收结果。

Showing the opening section of the README — the full document lives in the repository以上为 README 开头摘要,完整文档在仓库内 · View the full README on GitHub →在 GitHub 查看完整 README →

← 上一个 Prev dsh-git-forge 下一个 Next dsh-notify-on-complete →