haiziyao/dsh-vision-mix 预览 preview

haiziyao/dsh-vision-mix

用于DeepSeek Harness的视觉路由与图像生成,通过固定Mix模型实现。

Project Overview项目介绍

This is a native vision enhancement plugin built exclusively for DeepSeek Harness. To install it, you can run the command pnpm dsh plugin --profile web add dsh-vision-mix from your DSH repository root, then restart the DSH web server to activate the plugin. It combines your pre-configured text models, vision models, and optional image generation APIs to register a single Mix model entry, adding multimodal capabilities to text-only DSH agents without requiring you to re-enter your API credentials. All credentials are pulled directly from DSH’s native credential manager, so the plugin never stores or replicates your sensitive API keys.

After installation, you will complete a short setup flow in the DSH settings panel. First, test and enable the image input capability for your vision model, then select your base text model, vision model, and optional image generation provider from your pre-configured providers. This plugin is designed for DSH users that need to work with images, including developers who want to generate code from UI design mockups, analyze web page screenshot errors, or generate and edit custom images directly from an agent conversation. It supports cross-turn image follow-up questions and automatic detection of images returned by DSH agent tools.

This plugin is released under the permissive MIT open source license, so it is free to use, modify, and distribute for any purpose. It requires DSH version between 0.1.2-rc.1 and 0.2.0, and Node.js version 22.19.0 or higher, or 24.0.0 or higher, and only supports the DSH web profile. The optional sidebar history view for vision and generation calls requires the separate dsh-better-sidebar plugin, but core functionality works fully without it. All early alpha versions of the plugin are marked as incompatible due to reported dependency issues, so only the verified release should be used.

dsh-vision-mix 是专为 DeepSeek Harness 开发的原生视觉增强插件,它通过组合用户已经在 DSH 中配置好的文本模型、视觉模型和可选图片生成 API,输出一个名为 Mix 的混合模型入口,让原本不支持图片输入的 DSH Agent 获得图片识别、网页截图分析、跨轮图片追问以及图片生成编辑能力。所有模型凭据沿用 DSH 原生凭据系统管理,插件本身不存储或复制用户 API 密钥。

用户安装完成后,只需在 DSH 设置中完成三步配置:先检测并启用自定义视觉模型的图片能力,再分别选择基础对话模型、识图模型和可选的生图 API 提供者,保存路由后即可在新建对话中选择 Mix 模型开始使用。它适合需要让 DSH 处理图片输入、根据 UI 设计稿生成代码、分析网页布局问题的开发者与普通用户,支持模糊意图判断和多轮跨轮图片追问。

该插件要求 DSH 版本在 0.1.2-rc.1 到 0.2.0 之间,Node.js 版本需为 22.19.0 以上或 24.0.0 以上,仅支持 DSH 的 web 配置文件。插件采用 MIT 开源许可,完全免费使用,侧边栏历史记录功能需要可选插件 dsh-better-sidebar 支持,不安装也不影响核心功能使用,早期测试版存在部分兼容性问题未标记为可用。

Pre-install check安装前体检Compatibility · Security兼容性 · 安全性 1 warning1 项注意
  • Only 3 stars - very few users, little community feedback星标只有 3,几乎没人在用,遇到问题缺少社区反馈
DSH walks through these 9 checksDSH 会逐条核对这 9 项

Compatibility兼容性

  • DSH, Node, OS and profile requirementsDSH 版本 / Node 版本 / 操作系统 / profile 是否满足要求
  • External dependencies and runtimes (Electron / Python / Docker, ...)外部依赖与运行时(Electron / Python / Docker 等)是否齐备
  • Conflicts with installed plugins: command names, skill / tool names, ports, duplicate MCP registration与已装插件是否冲突:命令名、skill / tool 重名、端口占用、重复 MCP 注册

Security安全性

  • Repo matches the facts registered here; archived or abandoned?仓库是否与页面登记一致,是否归档或长期停更
  • Safety of preinstall / install / postinstall and install.sh / setup.ps1preinstall / install / postinstall 与 install.sh、setup.ps1 是否安全
  • curl|bash, download-then-execute, obfuscation, unrelated domains → stop immediatelycurl|bash、下载即执行、混淆代码、无关域名 → 立刻停止
  • Typosquatting or unmaintained packages among the new dependencies新增依赖里有没有 typosquatting 或无人维护的包
  • Requested permissions vs. what the feature actually needs申请了哪些权限、是否超出功能所需(filesystem / network / shell / clipboard)
  • Any sudo / admin requirement, plus uninstall and rollback是否要求 sudo / 管理员权限,以及卸载与回滚方式

Anything uncertain must be marked unknown with a note on how to confirm it. This site's signal screen is a static snapshot, not a security audit.拿不准的必须标「未知」并说明要我怎么确认。本站的信号筛查是静态快照,不能替代安全审计。

Or use CLI install (for developers)或使用命令行安装(适合开发者)

CLI Install命令行安装

dsh plugin --profile web add github:haiziyao/dsh-vision-mix

把 haiziyao/dsh-vision-mix 加入你的 DSH 配置(web profile)即可启用。

READMEREADME

Vision Mix

npm License DeepSeek Harness

Vision Mix 是 DeepSeek Harness 的视觉增强插件。它把文本模型、视觉模型和可选的图片生成 API 组合成一个 Mix 模型,让原本不支持图片的 Agent 能识别上传图片、理解网页截图、持续追问图片内容,并生成或编辑图片。

模型的 Base URL、协议和 API Key 仍由 DSH 的“设置 → 模型”统一管理。Vision Mix 只引用已经配置好的 Provider,不复制凭据。

三分钟上手

1. 安装

在 DeepSeek Harness 仓库目录执行:

pnpm dsh plugin --profile web add dsh-vision-mix

启动 Web:

pnpm dsh web

安装后的正常启动不需要 --patch。

兼容版本

  • DSH:>=0.1.2-rc.1 <0.2.0
  • Node.js:^22.19.0 || >=24.0.0
  • Profile:web

0.1.2-rc.1 已在 Windows x64 的一次性 Profile 中完成本地 tarball 安装、配置加载、Web 冷启动、HTTP 访问和卸载验证。0.1.3-alpha.1 没有可获取的 npm 发行物;0.1.3-alpha.2 的官方 CLI 在本机被自身的 fs-ext 原生依赖阻断,因此这两个版本均保守标记为 unknown。完整命令和边界见 兼容性验证记录。

2. 配置准备使用的模型

打开“设置 → 模型”,配置至少两个模型:

用途 要求 示例
基础模型 能处理文本、推理和工具调用 deepseek-chat、deepseek-v4-pro
识图模型 中转站实际支持 OpenAI/Anthropic 图片消息 gpt-5.6-sol、其他视觉模型

识图和基础模型可以来自同一个 Provider,也可以分别使用不同中转站。

如果还要生成或编辑图片,再配置一个支持 OpenAI-compatible /images/generations 和 /images/edits 的 Provider。它可以使用与识图模型完全不同的 Base URL 和 API Key。

3. 检测并启用模型的图片能力

打开“设置 → Vision Mix”,页面最上方是“基础设置”,下方的“模型图片能力”专门用于检测和声明能力:

  1. 从“待检测模型”选择刚才配置的视觉模型。
  2. 点击“测试并启用 image”。
  3. Vision Mix 会临时授予该模型 image 能力,发送一张红色测试图。
  4. 模型正确识别红色后,Vision Mix 才会保存能力声明;失败会自动回滚。
  5. 回到上方“基础设置”,由你自行在“图片模型”中选择该模型。

Vision Mix 基础路由与模型图片能力设置

测试会产生一次很小的模型调用,但不会出现在聊天会话中。

4. 选择基础模型和可选能力

继续在“设置 → Vision Mix”完成:

  1. 选择基础模型。
  2. 在图片模型列表中选择刚才测试通过的模型。
  3. 根据需要选择意图识别模型,用于判断“再详细一点”等模糊的跨轮图片追问。
  4. 根据需要开启“自动识别 Agent 工具返回的截图或图片”。
  5. 点击“保存路由”。

5. 开始使用

新建对话,在模型选择器中选择 Mix,然后尝试:

上传一张图片:这张图里是什么?
下一轮追问:右下角还有什么?
让 Agent 截图:检查当前网页有没有布局问题。

Mix 是 Vision Mix 注册的固定模型入口,不是一个需要单独填写 API Key 的模型服务。

中转站模型图片能力说明

很多中转站的 /models 只返回模型 id,不会声明模型能否接受图片。DSH 会安全地把这类自定义模型视为纯文本模型;如果直接发送图片,请求会在本地被拒绝,并不会到达中转站。

Vision Mix 提供三种接入方式:

操作 是否调用 API 结果
测试图片能力 是 临时声明 image,发送红色测试图,结束后始终恢复原配置
强制启用 image 否 直接保存 input: [text, image],适合已经自行验证过的中转站
测试并启用 image 是 正确识别红色才保存能力声明;失败自动回滚,推荐使用

Showing the opening section of the README — the full document lives in the repository以上为 README 开头摘要,完整文档在仓库内 · View the full README on GitHub →在 GitHub 查看完整 README →

← 上一个 Prev dsh-friendly-steps 下一个 Next dsh-PaddleOCR-Skills →