yuqingsh/dsh-image-subagent 预览 preview

yuqingsh/dsh-image-subagent

Project Overview项目介绍

This is a native plugin built exclusively for DeepSeek Harness (DSH) that enables pure text base models like deepseek-v4-pro to handle image attachments in chat sessions. It converts uploaded images into explicit text placeholders that carry the full attachment ID, then lets the base model delegate the image reading task to a configured vision sub-agent. The plugin matches different DSH versions with separate version branches, and can be installed either from a local checkout or a tagged Git release. To install via the CLI, you run the command dsh plugin --profile web add <source> with either the local path or the GitHub tag.

Before using the plugin, you need to have a pre-configured vision sub-agent in your DSH profile, that sub-agent must use a model that supports image input and include the read_image tool in its toolset. Once everything is set up, your workflow follows a simple pattern: you paste an image and add your question, the plugin modifies the route declaration to pass DSH's image access gate, the base model gets the rich placeholder with full attachment info. The base model then delegates the reading task to the vision sub-agent, which returns a detailed description, and the base model finally answers your original question. This plugin is built for DSH users who want to use pure-text base models to answer image-related questions.

Installing this plugin requires that you have the pnpm package manager installed on your system first; if you do not have it, you can install it via Homebrew with brew install pnpm on macOS. After installation completes, you need to restart the dsh web process and refresh your browser page to activate the plugin. You can also send a pre-written prompt to your DSH agent to let it handle the entire installation process for you. To confirm the plugin is working correctly, you can send a test POST request to a local API endpoint and check the response matches the expected output for your version. The plugin is released under the open-source MIT license, with no costs or restrictions on use.

这是一个专为DeepSeek Harness(DSH)开发的原生插件,核心功能是让不支持图片输入的纯文本主模型也能处理会话中的图片附件。它会将会话内的图片转换为带有完整附件ID的富文本占位符,由主模型委托已配置的视觉子代理读取图片内容后返回描述,最终主模型基于描述作答。它适配不同版本的DSH,可通过本地检出或Git标签两种方式完成安装。

本插件的典型工作流程为:用户上传图片并提出问题后,插件修改路由声明通过DSH的图片准入门禁,主模型收到带完整附件信息的占位符,随后委托预设的视觉子代理调用read_image工具读取图片,返回详细描述后主模型给出回答。它适合需要用纯文本大模型处理带图片问题的DSH用户,完全不影响原本支持图片输入的视觉路由正常工作。

安装本插件需要系统预先安装pnpm包管理器,安装完成后需要重启dsh web进程并刷新浏览器页面,也可以直接让DSH代理代为完成整个安装流程。你可以通过调用特定本地API接口验证插件是否安装成功,不同插件版本对应不同的预期返回结果。本插件采用MIT许可证开源,可供免费使用。

Pre-install check安装前体检Compatibility · Security兼容性 · 安全性 1 warning1 项注意
  • Only 4 stars - very few users, little community feedback星标只有 4,几乎没人在用,遇到问题缺少社区反馈
DSH walks through these 9 checksDSH 会逐条核对这 9 项

Compatibility兼容性

  • DSH, Node, OS and profile requirementsDSH 版本 / Node 版本 / 操作系统 / profile 是否满足要求
  • External dependencies and runtimes (Electron / Python / Docker, ...)外部依赖与运行时(Electron / Python / Docker 等)是否齐备
  • Conflicts with installed plugins: command names, skill / tool names, ports, duplicate MCP registration与已装插件是否冲突:命令名、skill / tool 重名、端口占用、重复 MCP 注册

Security安全性

  • Repo matches the facts registered here; archived or abandoned?仓库是否与页面登记一致,是否归档或长期停更
  • Safety of preinstall / install / postinstall and install.sh / setup.ps1preinstall / install / postinstall 与 install.sh、setup.ps1 是否安全
  • curl|bash, download-then-execute, obfuscation, unrelated domains → stop immediatelycurl|bash、下载即执行、混淆代码、无关域名 → 立刻停止
  • Typosquatting or unmaintained packages among the new dependencies新增依赖里有没有 typosquatting 或无人维护的包
  • Requested permissions vs. what the feature actually needs申请了哪些权限、是否超出功能所需(filesystem / network / shell / clipboard)
  • Any sudo / admin requirement, plus uninstall and rollback是否要求 sudo / 管理员权限,以及卸载与回滚方式

Anything uncertain must be marked unknown with a note on how to confirm it. This site's signal screen is a static snapshot, not a security audit.拿不准的必须标「未知」并说明要我怎么确认。本站的信号筛查是静态快照,不能替代安全审计。

Or use CLI install (for developers)或使用命令行安装(适合开发者)

CLI Install命令行安装

dsh plugin --profile web add github:yuqingsh/dsh-image-subagent#v0.2.0

把 yuqingsh/dsh-image-subagent 加入你的 DSH 配置(web profile)即可启用。

READMEREADME

dsh-image-subagent

让纯文本主模型(如 deepseek-v4-pro)也能接收图片附件:图片进入会话后投影为携带完整 attachmentId 的显式文本占位符,由主模型委托视觉子代理读取。

版本与 dsh 的对应关系:

  • v0.1.x —— dsh 0.1.0-rc.x(rc.6 时代,通过"补报 image 能力"放行门禁);
  • v0.2.0 —— dsh ≥ 0.1.1-rc.2(0.1.1 适配版,通过"抹除模态声明"放行门禁,详见下文"适配策略")。

适配策略(v0.2.0 为什么这样改)

0.1.1 的图片准入门禁只拒绝显式声明 inputModalities 且不含 image 的路由,声明省略(undefined)视为负能力直接放行;同时官方适配器对纯文本路由已有原生占位投影(只含 sha256 前 8 位摘要)。因此本版把旧版的"补报 image 能力"改为"抹除纯文本路由的 inputModalities 声明":

  • 门禁:undefined → 放行,贴图不再被拒绝;
  • 适配器:仍按纯文本路由处理 → 本插件在主循环路径投影为富占位(完整 id + 委托提示),漏网路径由官方原生占位兜底——不存在把图片当真发给纯文本端点的失败模式;
  • read_image 工具门禁:主模型路由 undefined → 仍拒绝(主模型不能自己读图,委托流保持不变);
  • 视觉路由(如 deepseek-v4-flash-vision-exp)完全不受影响:声明含 image,插件旁路,原生直读。

安装

# 本地 checkout(推荐,方便跟进改动)
dsh plugin --profile web add ./dsh-image-subagent

# 或按 git 标签
dsh plugin --profile web add github:yuqingsh/dsh-image-subagent#v0.2.0

需要 pnpm(brew install pnpm)。安装后重启 dsh web 并刷新页面。

用 Prompt 安装

把下面这段发给你的 DSH 会话,让 Agent 代装:

请安装 dsh-image-subagent 插件:在 /Users/kane/Documents/OpenCode 下运行 dsh plugin --profile web add ./dsh-image-subagent。完成后重启 dsh web 进程,刷新浏览器页面。

使用

前置条件

  • 预设里有一个视觉子代理(如 subagent_observer),其模型声明 image 输入(如 deepseek-v4-flash-vision-exp 或 MiniMax-M3);
  • 该子代理的工具集包含 read_image;
  • 贴图进入会话后,主模型会收到富占位文本(含 id=sha256:<64位hex>),委托时把占位符里给出的附件对象路径传给子代理。

日常使用

贴图并附上问题 → 门禁放行 → 主模型收到富占位符(含完整 id=sha256:… 和附件对象路径)→ 委托 subagent_observer 用 read_image 读该路径 → 返回像素级描述 → 主模型作答。

附件对象路径规则:~/.dsh/attachments/v1/objects/<id 前 2 位 hex>/<64 位 hex>。

验证安装

curl -s -X POST http://127.0.0.1:3080/image-subagent/status \
  -H 'content-type: application/json' \
  -d '{"type":"client-request","rpcId":"s1","method":"status","payload":{}}'

预期(v0.2.0):bridged 为 "(omitted — paste gate passes)",real 为 ["text"](路由真实声明)。旧版 v0.1.x 的预期是 bridged 含 "image"。

License

MIT

← 上一个 Prev DeepSeek-splash-animation 下一个 Next dsh-auth-everying →