Anionex/dsh-vision-toolkit 预览 preview

Anionex/dsh-vision-toolkit

[dsh] 为纯文本模型打造更强大的视觉工具箱:免费安装、粘贴图片即可识别、支持多图问答、截图还原为前端UI等|DeepSeek Harness 原生集成 agent-vision-toolkit:图像问答、长截图OCR、UI还原、目标定位、像素对比、Artifacts及Web UI。

Project Overview项目介绍

This is a native DSH plugin built exclusively for DeepSeek Harness, wrapping the upstream general-purpose agent-vision-toolkit to add visual processing capabilities to text-only large language models running on the DSH platform. It supports a range of visual tasks including image Q&A, long screenshot OCR, UI restoration, and general GUI visual work, and works with both DSH Web and Headless Profiles. To install, users run a single one-line command dsh plugin --profile web add @anionex/dsh-vision-toolkit via the DSH CLI, then configure their preferred vision model provider in the DSH settings menu after installation completes.

When using the plugin on DSH Web, pasting an image automatically switches the text-only model to a Vision Toolkit-adapted variant, so users do not need to manually adjust file paths or switch models manually. All native DSH features like thumbnails, session history, and workspace paths remain fully intact after the switch. The plugin also lets users preview generated artifacts directly in the DSH Web interface. It is designed for any DSH user that needs their text-only agent to handle image-based tasks, whether for development or general personal use.

The plugin is released under the permissive MIT open-source license, and it is free for all personal and commercial use. Each visual task only sends the necessary intent and image to the multimodal model, and no extra context is accumulated across calls, so additional usage costs stay very low. Users can further cut costs by using a locally deployed small multimodal side model like Gemma 4 or Qwen 3.5/3.6. First-time setup may require downloading a standalone Python runtime, so users need to ensure they have stable network access and sufficient available disk space to complete the initial setup.

这是一个专为 DeepSeek Harness (DSH) 打造的原生视觉能力插件,基于上游通用 agent-vision-toolkit 封装,为 DSH 中的纯文本大模型添加图像理解能力,支持图像问答、长截图 OCR、界面还原、GUI 视觉任务等多种场景。用户可通过 DSH 命令行一键安装,安装完成后只需在设置面板配置好视觉模型提供商即可开始使用。

在 DSH Web 端使用时,粘贴图像后系统会自动切换为 Vision Toolkit 适配的模型变体,无需手动修改文件路径或切换模型,同时完整保留原生缩略图、会话历史和工作区路径,还支持预览插件生成的结果文件。该工具适合所有需要让纯文本 DSH 代理处理各类图像相关任务的开发者和普通用户。

本插件采用 MIT 许可证开源,完全免费使用,每次视觉任务仅向多模态模型发送必要的意图和图像,不会累积上下文,因此额外产生的成本很低。它支持使用本地部署的小型多模态侧模型来进一步降低成本,首次运行时可能需要下载 Python 运行环境,需保证网络连接和可用磁盘空间。

Pre-install check安装前体检Compatibility · Security兼容性 · 安全性 no risk signal found未发现风险信号
  • This site's static screen found no obvious risk signal (stars, license, activity, manifest)本站静态筛查没发现明显风险信号(星标、许可证、更新活跃度、清单完整度)
DSH walks through these 9 checksDSH 会逐条核对这 9 项

Compatibility兼容性

  • DSH, Node, OS and profile requirementsDSH 版本 / Node 版本 / 操作系统 / profile 是否满足要求
  • External dependencies and runtimes (Electron / Python / Docker, ...)外部依赖与运行时(Electron / Python / Docker 等)是否齐备
  • Conflicts with installed plugins: command names, skill / tool names, ports, duplicate MCP registration与已装插件是否冲突:命令名、skill / tool 重名、端口占用、重复 MCP 注册

Security安全性

  • Repo matches the facts registered here; archived or abandoned?仓库是否与页面登记一致,是否归档或长期停更
  • Safety of preinstall / install / postinstall and install.sh / setup.ps1preinstall / install / postinstall 与 install.sh、setup.ps1 是否安全
  • curl|bash, download-then-execute, obfuscation, unrelated domains → stop immediatelycurl|bash、下载即执行、混淆代码、无关域名 → 立刻停止
  • Typosquatting or unmaintained packages among the new dependencies新增依赖里有没有 typosquatting 或无人维护的包
  • Requested permissions vs. what the feature actually needs申请了哪些权限、是否超出功能所需(filesystem / network / shell / clipboard)
  • Any sudo / admin requirement, plus uninstall and rollback是否要求 sudo / 管理员权限,以及卸载与回滚方式

Anything uncertain must be marked unknown with a note on how to confirm it. This site's signal screen is a static snapshot, not a security audit.拿不准的必须标「未知」并说明要我怎么确认。本站的信号筛查是静态快照,不能替代安全审计。

Or use CLI install (for developers)或使用命令行安装(适合开发者)

CLI Install命令行安装

dsh plugin --profile web add @anionex/dsh-vision-toolkit

把 Anionex/dsh-vision-toolkit 加入你的 DSH 配置(web profile)即可启用。

READMEREADME

DSH Vision Toolkit helps text-only DeepSeek Harness agents understand images and complete visual tasks

DSH Vision Toolkit

Anionex%2Fdsh-vision-toolkit | Trendshift

Recommended by dshfind dshfind score: 94 — highest-rated plugin agentic leaderboard

npm MIT DSH

A more powerful vision toolkit—give text-only models in DeepSeek Harness eyes: image Q&A, long-screenshot OCR, UI restoration, and GUI visual tasks in one toolkit and Skill.

🚀 Paste an image and ask directly | Install with one command | Broad use cases

Highlights | Quick start | Toolbox | Configuration and limits | Troubleshooting | Community

🌐 English | 中文

🏆 This project is the first comprehensive vision-tool plugin in the DeepSeek Harness ecosystem: it was initiated before internal beta and built during the beta with reference to agent-vision-toolkit.

Showing the opening section of the README — the full document lives in the repository以上为 README 开头摘要,完整文档在仓库内 · View the full README on GitHub →在 GitHub 查看完整 README →

← 上一个 Prev agent-qa 下一个 Next Wegent →