woyeshishen/dsh-vision-plugin

Plugin插件 Native原生 ⭐ 2 Apache-2.0 Vision & Media视觉与多媒体

Project Overview项目介绍

dsh-vision-plugin adds image understanding to DeepSeek Harness by wiring up an OpenAI-compatible external vision model. A text-only main model can call the describe_image tool to hand off an image and receive a plain-text description for screenshots, photos, charts, OCR, and UIs. Images flow only to the secondary model; the main model always works with text. The plugin auto-loads at DSH startup and ships with a GUI settings page for URL, API key, and model selection, with credentials stored encrypted. Use it when conversations need to reason over image content. Caveat: an external /chat/completions endpoint with image support is required.

为 DeepSeek Harness 引入图像理解能力:接入 OpenAI 兼容的外部视觉模型后,纯文本主模型(如 deepseek)即可调用 describe_image 工具,将截图、照片、图表等交给外部模型并获得纯文本描述。架构上图像仅送往辅助模型,主模型始终处理文本。安装后插件自动加载,提供 GUI 配置页填写 URL、API Key 和模型名,密钥加密保存不再回显。适用于需要识别 UI、OCR 或图表内容的对话场景。注意:需自备支持图像的 OpenAI 兼容 /chat/completions 端点。

Pre-install check安装前体检Compatibility · Security兼容性 · 安全性 1 warning1 项注意
  • Only 2 stars - very few users, little community feedback星标只有 2,几乎没人在用,遇到问题缺少社区反馈
DSH walks through these 9 checksDSH 会逐条核对这 9 项

Compatibility兼容性

  • DSH, Node, OS and profile requirementsDSH 版本 / Node 版本 / 操作系统 / profile 是否满足要求
  • External dependencies and runtimes (Electron / Python / Docker, ...)外部依赖与运行时(Electron / Python / Docker 等)是否齐备
  • Conflicts with installed plugins: command names, skill / tool names, ports, duplicate MCP registration与已装插件是否冲突:命令名、skill / tool 重名、端口占用、重复 MCP 注册

Security安全性

  • Repo matches the facts registered here; archived or abandoned?仓库是否与页面登记一致,是否归档或长期停更
  • Safety of preinstall / install / postinstall and install.sh / setup.ps1preinstall / install / postinstall 与 install.sh、setup.ps1 是否安全
  • curl|bash, download-then-execute, obfuscation, unrelated domains → stop immediatelycurl|bash、下载即执行、混淆代码、无关域名 → 立刻停止
  • Typosquatting or unmaintained packages among the new dependencies新增依赖里有没有 typosquatting 或无人维护的包
  • Requested permissions vs. what the feature actually needs申请了哪些权限、是否超出功能所需(filesystem / network / shell / clipboard)
  • Any sudo / admin requirement, plus uninstall and rollback是否要求 sudo / 管理员权限,以及卸载与回滚方式

Anything uncertain must be marked unknown with a note on how to confirm it. This site's signal screen is a static snapshot, not a security audit.拿不准的必须标「未知」并说明要我怎么确认。本站的信号筛查是静态快照,不能替代安全审计。

Or use CLI install (for developers)或使用命令行安装(适合开发者)

CLI Install命令行安装

dsh plugin --profile web add @woyeshishen/dsh-vision-plugin

把 woyeshishen/dsh-vision-plugin 加入你的 DSH 配置(web profile)即可启用。

READMEREADME

dsh-vision-plugin

npm version license

中文 | English

Adds image understanding to DeepSeek Harness (DSH): wire up an OpenAI-compatible external vision model, and a text-only main model (e.g. deepseek) can call the describe_image tool to hand an image to it and get a plain-text description — understanding screenshots, photos, charts, OCR, UIs, and more.

Design: images go only to the secondary model (external vision model); the main model always deals with text.

Features

🖼️ Image understanding The main model calls describe_image and gets a plain-text description
⚙️ GUI configuration Fill in URL / API key / model on a settings page — no config files to edit
🔒 Secure credentials API key stored in the credential store, never echoed
📦 Install once, keep working Auto-loads at DSH startup, survives restarts

Install

One-liner (recommended)

Windows (PowerShell)

irm https://raw.githubusercontent.com/woyeshishen/dsh-vision-plugin/main/scripts/install.ps1 | iex

macOS / Linux

bash <(curl -fsSL https://raw.githubusercontent.com/woyeshishen/dsh-vision-plugin/main/scripts/install.sh)

dsh plugin command

From npm

dsh plugin --profile web add @woyeshishen/dsh-vision-plugin

From GitHub

dsh plugin --profile web add github:woyeshishen/dsh-vision-plugin

After install, the plugin auto-mounts into the profile; restart DSH (or hot-reload) to activate.

Usage

Step 1: Configure the external vision model

Open Settings → Multimodal Vision:

Field Description
URL (Base URL) OpenAI-compatible endpoint, e.g. https://api.example.com/v1
API key Secret for the external model (stored encrypted, never echoed)
Model Click "Load models" to fetch and pick from the endpoint

Click Save.

Step 2: Ask the main model to look at an image

In a conversation, say:

Take a look at D:\path\to\image.png and describe what's in it.

The main model calls describe_image, sends the image to the external vision model, and continues reasoning from the returned description.

Tool

describe_image

Parameter Required Description
path ✅ Image file path; supports png / jpg / jpeg / webp / gif
prompt ❌ Specific question about the image; defaults to "describe the image in detail"

Requirements

  • DeepSeek Harness
  • An OpenAI-compatible (/chat/completions), image-capable external vision model

License

Apache-2.0

← 上一个 Prev dsh-file-ref 下一个 Next dsh-skill-eval →