sala003/dsh-tool-describe-image
Project Overview项目介绍
dsh-tool-describe-image is a native DSH plugin built specifically for the DeepSeek Harness platform. It adds image description capabilities to DSH, allowing pure text models that do not natively support image input to "see" images by converting them to text descriptions or structured HTML using any OpenAI-compatible visual API endpoint. Supported endpoints include Alibaba Cloud Bailian Qwen-VL, Zhipu GLM-4V, OpenRouter, and any self-hosted visual model endpoint that follows OpenAI’s API format. The plugin also adds an interactive animated whale pet widget to the bottom right of DSH’s web interface, which supports dragging, custom interactions, and automatic linkage to DSH Agent status.
There are two main workflows to use the image description feature. For web users, you can simply copy an image to your clipboard and paste it directly into DSH’s input box, and the plugin will automatically process the image, convert it to text, and insert the result into your input for you to send to the model. You can also ask the model to call the built-in describe_image tool directly by passing a local file path to the image you want analyzed. Double-clicking the whale pet opens a floating settings panel where you can configure all plugin settings, including API credentials, model selection, output format, and DeepSeek balance checks.
To install the plugin, you first need to have DSH version 0.1.0-rc.6 or newer installed on your system. You can install the plugin via npm with two common steps: first run npm install -g dsh-tool-describe-image to install it globally, then run dsh plugin --profile web add dsh-tool-describe-image to add it to your DSH web profile. Version 0.5.0 had a known packaging bug that caused tool calls to crash, so if you installed that version you must uninstall it first before installing the fixed 0.5.1 version, then restart DSH and refresh your browser. The plugin is open source under the MIT license, free to use and modify.
这是一款专为 DeepSeek Harness(DSH)开发的原生插件,核心功能是让纯文本大模型也能处理图片,通过任意兼容 OpenAI 格式的视觉端点(包括百炼千问、智谱 GLM-4V、OpenRouter 以及自建端点),把图片转换为文字描述或者结构化 HTML,供纯文本模型理解。插件自带一个可交互的鲸鱼娘桌宠,放置在 DSH Web 界面右下角,支持拖动、表情切换,还能和 Agent 的运行状态联动。
普通用户可以通过两种方式使用图片描述功能:在 Web 界面直接粘贴截图,插件会自动识别转换后把结果填入输入框;也可以让模型调用 describe_image 工具,按指定路径读取图片后生成描述。双击桌宠就能打开悬浮设置面板,配置 API 端点、模型参数,还能一键查询 DeepSeek 账户余额,查看各币种剩余额度以及赠送/充值的拆分情况。
该插件要求用户预先安装 DSH 0.1.0-rc.6 或更高版本,可通过 npm 全局安装后用 dsh plugin add 命令完成安装,所有配置都可以在插件面板内完成,零配置也能启动。旧版 0.5.0 存在打包缺陷,如果遇到工具调用崩溃的问题,需要卸载后重新安装最新版 0.5.1。项目采用 MIT 许可证开源,允许自由修改和分发。
请帮我安装这个 DSH 插件。安装前先完成【兼容性检查 + 安全性检查】,检查通过再动手。
插件:dsh-tool-describe-image(sala003/dsh-tool-describe-image)
仓库:https://github.com/sala003/dsh-tool-describe-image
本站详情页:https://www.yhbd.top/plugins/sala003-dsh-tool-describe-image/
本站登记:类型 plugin · 归类 原生 DSH 插件 · 许可证 MIT · ⭐ 4 · 最近提交 2026-08-25 · 主语言 TypeScript
按下面顺序执行,每步先把结论告诉我,再进入下一步:
【1 兼容性检查】
① 我这边:DSH 版本、Node 版本、操作系统、当前 profile(web / desktop)。
② 读它的 README、package.json、插件 manifest,列出它要求的 DSH 版本 / Node 版本 / 操作系统 / 外部依赖 / 需要另外先装的运行时。
③ 逐条比对,结论只写「满足 / 不满足 / 未知」三种;不满足的给出可行替代方案。
④ 检查是否和我已装的插件冲突:命令名重复、skill / tool 重名、端口占用、重复注册的 MCP server。
【2 安全性检查】
① 仓库可信度:和上面「本站登记」是否一致;star / fork 数、创建时间、最近提交,是否归档或长期停更。
② 安装脚本:逐行看 package.json 的 preinstall / install / postinstall,以及 install.sh、setup.ps1 之类脚本。出现 curl|bash、下载后直接执行、混淆代码、访问与插件功能无关的域名,立刻停下来告诉我,不要继续装。
③ 依赖:列出新增依赖,标出无人维护、或与知名包拼写近似的可疑包(typosquatting)。
④ 权限与副作用:它会读写哪些目录、访问哪些域名、需要哪些 DSH 权限(filesystem / network / shell / clipboard 等),以及怎么卸载和回滚。
⑤ 如果它要求 sudo / 管理员权限,或权限明显超出功能所需,先停下来问我。
【3 安装】
上面两步没有「不满足」和「高危项」时才执行;用官方推荐方式安装,不要自行提权。
【4 汇报】
用表格输出:检查项 / 结论 / 依据 / 是否需要我决策。拿不准的一律写「未知」并说明要我怎么确认——不要猜,也不要替我决定。
Send this message to DSH in your current session: it verifies compatibility and security first (answering met / not met / unknown item by item) and only installs once everything checks out — it will stop and ask you if it finds a high-risk item. The box scrolls; the copy is the full prompt. CLI install commands may not be accurate across systems, so DSH is the safer route.把上面这条消息直接发给当前会话里的 DSH:它会先核对兼容性与安全性(逐条给「满足 / 不满足 / 未知」),确认没问题再安装,有高危项会停下来问你。框内可滚动,复制到的是完整提示词;安装命令不一定准确,发给 DSH 更稳。
- Only 4 stars - very few users, little community feedback星标只有 4,几乎没人在用,遇到问题缺少社区反馈
DSH walks through these 9 checksDSH 会逐条核对这 9 项
Compatibility兼容性
- DSH, Node, OS and profile requirementsDSH 版本 / Node 版本 / 操作系统 / profile 是否满足要求
- External dependencies and runtimes (Electron / Python / Docker, ...)外部依赖与运行时(Electron / Python / Docker 等)是否齐备
- Conflicts with installed plugins: command names, skill / tool names, ports, duplicate MCP registration与已装插件是否冲突:命令名、skill / tool 重名、端口占用、重复 MCP 注册
Security安全性
- Repo matches the facts registered here; archived or abandoned?仓库是否与页面登记一致,是否归档或长期停更
- Safety of preinstall / install / postinstall and install.sh / setup.ps1preinstall / install / postinstall 与 install.sh、setup.ps1 是否安全
- curl|bash, download-then-execute, obfuscation, unrelated domains → stop immediatelycurl|bash、下载即执行、混淆代码、无关域名 → 立刻停止
- Typosquatting or unmaintained packages among the new dependencies新增依赖里有没有 typosquatting 或无人维护的包
- Requested permissions vs. what the feature actually needs申请了哪些权限、是否超出功能所需(filesystem / network / shell / clipboard)
- Any sudo / admin requirement, plus uninstall and rollback是否要求 sudo / 管理员权限,以及卸载与回滚方式
Anything uncertain must be marked unknown with a note on how to confirm it. This site's signal screen is a static snapshot, not a security audit.拿不准的必须标「未知」并说明要我怎么确认。本站的信号筛查是静态快照,不能替代安全审计。
Or use CLI install (for developers)或使用命令行安装(适合开发者)
CLI Install命令行安装
dsh plugin --profile web add dsh-tool-describe-image
把 sala003/dsh-tool-describe-image 加入你的 DSH 配置(web profile)即可启用。
READMEREADME
dsh-tool-describe-image(大肥鱼看世界)
让 DeepSeek 看懂图片的 DSH 插件 —— 通过任意 OpenAI 兼容视觉端点(百炼 qwen-vl、智谱 GLM-4V、OpenRouter、自建代理等),把图片转成文字或结构化 HTML,纯文本模型也能"看见"图片。
项目特点
- 悬浮鲸鱼娘桌宠:Web 右下角的动画桌宠,支持拖动、表情切换、状态气泡、Agent 状态联动——是"大肥鱼看世界"的守护精灵
- Agent 状态联动:模型工作时鲸鱼娘切"专注"表情 + 气泡"我在认真工作呢…",空闲时回待机 + "等你来聊~"
- 零配置可进 DSH:不配 API key 也能正常启动,视觉功能随时可从设置窗补配
- 通用视觉通道:baseURL + API key + 模型名自由填写,兼容任何 OpenAI 兼容视觉 API
- 自定义 Prompt + 输出格式:Prompt 指令可编辑,输出可选文本描述或结构化 HTML
- 粘贴即识别(Web):
Ctrl+V粘贴图片,自动识别成文字/HTML 填入输入框 describe_image工具:模型按路径读取图片转成文字,任何 surface(Web / headless)都可用- DeepSeek 余额检测:
check_balance工具 + 鲸鱼娘设置面板「查余额」按钮,实时查询 DeepSeek 账户余额(账户可用状态 + 各币种剩余额度,含赠送/充值拆分),余额不足时第一时间知道 - 零源码改动:完全通过 DSH 公开扩展点实现(工具注册、RPC 通道、凭据服务、设置服务、客户端槽位),不改 Harness 任何代码
- 单包全包:一个 npm 包 = host 半区 + 浏览器半区(鲸鱼娘桌宠 + 粘贴识别),
dsh plugin add即装即用 - 双语文案:跟随 Web 界面语言自动切换中文 / 英文
鲸鱼娘桌宠交互
| 操作 | 行为 |
|---|---|
| 单击 | 随机表情反应 + 气泡"嘿嘿,戳我干嘛~" |
| 双击 | 打开跟随鲸鱼娘的悬浮设置面板(非全屏) |
| 按住拖动 | 拖动鲸鱼娘,松手落位 + 位置记忆(刷新后回到原位) |
| Agent 工作 | 切"专注"表情 + 气泡"我在认真工作呢…" |
| Agent 空闲 | 回待机(眨眼循环)+ 气泡"等你来聊~" |
表情动画用交叉淡化 + 淡入覆盖实现(无重影、无变淡),眨眼瞬时换帧保持利落。
工作原理
DeepSeek 官方适配器是纯文本路由,DSH 内置的 read_image 工具会在模型不支持 image 输入时拒绝。本插件绕开这道限制:
粘贴图片 / 指定路径
→ 图片字节交给用户配置的 OpenAI 兼容视觉端点
→ 返回文字描述或结构化 HTML(作为 tool result 或 draft 文本)
→ DeepSeek 基于文字/HTML 理解图片
- 工具模式:
describe_image(file_path, question?)注册到ctx.tools,schema 自动流入系统提示词 - 余额检测:
check_balance注册到ctx.tools,复用 Harness 的标准DEEPSEEK_API_KEY凭据,调用 DeepSeek 官方GET /user/balance返回账户可用状态与各币种剩余额度 - 粘贴模式:浏览器半区在 document 上挂 capture 阶段 paste 监听,拦截图片 → base64 → host RPC(
/dsh-describe-image)→ 文字填入输入框,带"正在识别…"提示与失败 toast - 鲸鱼娘桌宠:挂
shell.overlay(框架级悬浮层槽),表情帧由 host 静态路由/plugins-assets/...按需加载,Agent 状态经ctx.sessions订阅
用户配置(悬浮窗内全部可配)
| 配置项 | 默认值 | 说明 |
|---|---|---|
| API Base URL | https://dashscope.aliyuncs.com/compatible-mode/v1 |
任意 OpenAI 兼容端点 |
| API Key | 无 | 存 DSH 凭据库;优先级:config > 环境变量 > 凭据库 > .env |
| 模型名称 | qwen-vl-plus |
自由输入 |
| 输出格式 | text |
文本描述 / HTML 结构化 |
| Prompt 指令 | 默认详细描述指令 | 用户可自定义 |
| 鲸鱼娘位置 | 右下角 | 拖动后自动记忆(petX/petY) |
| DeepSeek 余额 | — | 「查余额」按钮,查询 DeepSeek 账户剩余额度(key 复用 DEEPSEEK_API_KEY) |
Showing the opening section of the README — the full document lives in the repository以上为 README 开头摘要,完整文档在仓库内 · View the full README on GitHub →在 GitHub 查看完整 README →
nexu-io/open-design
Devin-AXIS/iPolloWork
liustack/modlens
Alisa0808/vox-director
EthanYoQ/AI-Novel-Writer
Anionex/agent-vision-toolkit
ysr666/dsh-vision-router
tong-io/tongflow