tdf1995/dsh-plugin-vision

DeepSeek Harness(DSH)中纯文本LLM的视觉能力:通过免费的Gemini和GLM视觉API实现图像描述/OCR/VQA

Project Overview项目介绍

This is a native plugin built exclusively for DeepSeek Harness (DSH) that adds vision capabilities to text-only large language models running inside the DSH environment. It leverages the free public vision APIs from Google Gemini and Zhipu GLM to enable image description, optical character recognition, and visual question answering directly within DSH chat sessions. The plugin follows DSH's official plugin specification, supports installation via npm as an overlay patch for temporary testing or permanent merge into a user profile, and never stores user API keys in the public repository code. To verify a successful install, simply restart your DSH session and ask the model to analyze an image path to confirm the see_image tool is called correctly.

This plugin is designed for users who run DSH with a text-only DeepSeek LLM and regularly need to process image-based tasks like screenshot analysis, document OCR, and visual question answering. After you install the plugin and configure at least one valid API key, you can just describe your request naturally in the chat window. For example, you can ask the model what product is in an order screenshot, and it will automatically call the vision tool to get the answer. The plugin also supports pasting or dragging and dropping images directly in the browser-based DSH web UI, which creates a native-style attachment card for easy management.

The plugin requires Node.js version 20 or higher to run, and users must obtain at least one free API key from either Google AI Studio for Gemini or BigModel for Zhipu GLM. Gemini requires a proxy for access from within mainland China, while Zhipu GLM offers free direct access for domestic Chinese users. The plugin is released under the permissive MIT open source license, and there is no cost to use the plugin itself. All vision calls use the free tier of the external APIs, so typical usage has negligible cost for most daily tasks.

这是一款专为 DeepSeek Harness 开发的原生视觉插件,为 DSH 中的纯文本大模型补充图像理解能力。它调用 Gemini 和智谱 GLM 的免费视觉 API,支持在 DSH 会话内直接完成图像描述、OCR 识别和视觉问答任务。本插件遵循 DSH 官方插件格式发布,可通过 npm 安装或 overlay patch 挂载,不会将用户 API Key 存入仓库代码。

适合使用 DeepSeek Harness 搭配纯文本 DeepSeek 模型,有日常图像分析、OCR 识别需求的用户使用。安装配置 API Key 后,用户只需在对话中自然描述需求,比如询问截图内容、提取订单信息,模型就会自动调用 see_image 工具完成处理。插件还支持浏览器端直接粘贴或拖入图片,生成原生风格附件卡片,操作流程简洁流畅。

本插件要求 Node.js 版本不低于 20,用户需要至少申请一个 Gemini 或智谱 GLM 的免费 API Key。Gemini 国内访问需要代理,GLM 可国内直连免费使用。插件采用 MIT 许可证开源,自身无使用成本,仅依赖外部 API 的免费额度,日常使用开销可忽略。

Pre-install check安装前体检Compatibility · Security兼容性 · 安全性 1 warning1 项注意
  • Only 4 stars - very few users, little community feedback星标只有 4,几乎没人在用,遇到问题缺少社区反馈
DSH walks through these 9 checksDSH 会逐条核对这 9 项

Compatibility兼容性

  • DSH, Node, OS and profile requirementsDSH 版本 / Node 版本 / 操作系统 / profile 是否满足要求
  • External dependencies and runtimes (Electron / Python / Docker, ...)外部依赖与运行时(Electron / Python / Docker 等)是否齐备
  • Conflicts with installed plugins: command names, skill / tool names, ports, duplicate MCP registration与已装插件是否冲突:命令名、skill / tool 重名、端口占用、重复 MCP 注册

Security安全性

  • Repo matches the facts registered here; archived or abandoned?仓库是否与页面登记一致,是否归档或长期停更
  • Safety of preinstall / install / postinstall and install.sh / setup.ps1preinstall / install / postinstall 与 install.sh、setup.ps1 是否安全
  • curl|bash, download-then-execute, obfuscation, unrelated domains → stop immediatelycurl|bash、下载即执行、混淆代码、无关域名 → 立刻停止
  • Typosquatting or unmaintained packages among the new dependencies新增依赖里有没有 typosquatting 或无人维护的包
  • Requested permissions vs. what the feature actually needs申请了哪些权限、是否超出功能所需(filesystem / network / shell / clipboard)
  • Any sudo / admin requirement, plus uninstall and rollback是否要求 sudo / 管理员权限,以及卸载与回滚方式

Anything uncertain must be marked unknown with a note on how to confirm it. This site's signal screen is a static snapshot, not a security audit.拿不准的必须标「未知」并说明要我怎么确认。本站的信号筛查是静态快照,不能替代安全审计。

Or use CLI install (for developers)或使用命令行安装(适合开发者)

CLI Install命令行安装

dsh plugin --profile web add github:tdf1995/dsh-plugin-vision

把 tdf1995/dsh-plugin-vision 加入你的 DSH 配置(web profile)即可启用。

READMEREADME

dsh-plugin-vision

为 DeepSeek Harness 中的纯文本大模型提供视觉能力——通过 Gemini / GLM 免费视觉 API 完成图像描述、OCR 与视觉问答。

Vision for text-only LLMs inside DeepSeek Harness (DSH) — describe images, OCR, and answer visual questions through the free Gemini / GLM vision APIs.

License Node Platform Vision


目录 Table of Contents


简介 Introduction

DeepSeek 等主流文本模型不支持图片输入。本插件通过 Gemini(gemini-3.6-flash)与 智谱 GLM(glm-4.6v-flash,完全免费、国内直连)等外部视觉 API,把「看图」能力封装为 DSH 会话内可直接调用的工具,无需切换模型即可完成截图解读、文档 OCR、图表分析等任务。

本插件遵循 DSH 官方插件格式发布(dsh.bundle.patch + cordis.patch.yml),可作为 npm 包安装,或以 overlay patch 挂载。任何 API Key 都不会进入代码或仓库。

功能特性 Features

特性 说明
🖼️ see_image 工具 分析本地图片(png / jpg / jpeg / webp / gif,≤ 20MB),支持自定义提问
🔀 双提供商 Gemini 与 GLM 双通道,auto 模式自动选择可用 Key 的提供商
⚡ 粘性路由 自动记住上次成功的提供商,减少无效请求等待
🔄 故障转移 网络失败或限流时自动切换至另一提供商
⏱️ 限流重试 429 / 访问量过大自动退避重试(默认 3 次)
🗜️ 大图压缩 超过 4MB 的图片自动压缩至 1920px(JPEG 质量 85)后上传
🔑 vision_set_key 会话内保存 API Key(写入 DSH 凭据库,立即生效)
📊 vision_status 查看各提供商 Key 配置状态(不泄露 Key 本身)
📋 粘贴/拖入图片(内置) 浏览器端 Ctrl+V / 拖入 / 🖼️ 按钮 → 原生风格附件卡片,随包集成、重启不丢

安装 Installation

前置要求 Prerequisites

  • Node.js ≥ 20
  • DeepSeek Harness (DSH) 已部署
  • 至少一个视觉 API Key(Gemini 或智谱 GLM,均可免费申请)

方式一:overlay patch(临时试用)

Showing the opening section of the README — the full document lives in the repository以上为 README 开头摘要,完整文档在仓库内 · View the full README on GitHub →在 GitHub 查看完整 README →

← 上一个 Prev plugin-switch 下一个 Next dsh-soul →