54xkeee/dsh-youreyes

DeepSeek Harness 上纯文本 DeepSeek 的视觉扩展:模型可调用的视觉工具 + 包装适配器 + 通用 VLM 通道(兼容 OpenAI / Gemini / 本地 Ollama)

Project Overview项目介绍

dsh-youreyes is a native vision plugin built exclusively for DeepSeek Harness (DSH), designed to add image recognition capabilities to the text-only DeepSeek model. It lets users paste images, screenshots or pass local file paths directly in DSH conversations, then calls a visual model from multiple supported backends to process the image, and returns structured recognition results back to DeepSeek for answering. Supported backends include the default Antigravity IDE channel, local Ollama instances, Gemini API, and any OpenAI-compatible endpoints, so it can fit different user preferences and network environments.

The basic workflow for using dsh-youreyes is straightforward: after installing the plugin via the DSH CLI, users just paste their image and type their question directly into the chat. The plugin automatically converts the image input to a text placeholder to avoid DeepSeek rejecting the request, then DeepSeek triggers the vision tool to run recognition, and the plugin inserts the structured text evidence back into the conversation context for DeepSeek to generate a final answer. This plugin is built for any DSH user that needs to work with images, from casual question asking to professional UI debugging and document OCR.

To run dsh-youreyes, you need Node.js version 20 or higher, and the plugin is released under the permissive MIT license, so it is free to use, modify, and distribute. There are a few usage limits: single images can not exceed 8MB in size, you can process a maximum of 8 images per request, and the in-memory LRU cache holds up to 64 recognition entries. For privacy, when using local Ollama, images never leave your machine, and API keys are only stored in your local config file and never logged, so your sensitive information stays protected.

dsh-youreyes 是专为 DeepSeek Harness 开发的原生视觉识别插件,核心功能是给纯文本 DeepSeek 增加视觉能力,支持在 DSH 中粘贴图片、截图或者传入本地文件路径,调用多后端视觉模型识别图片内容,再把结构化识别结果返回给 DeepSeek 让它回答用户问题。它支持多种视觉后端,包括反重力 IDE 默认通道、本地 Ollama、Gemini API、任意 OpenAI 兼容端点等,能满足不同用户的配置和使用需求。

用户使用时只需要粘贴图片并提问,插件会自动把图片转换为文本占位符,避免纯文本 DeepSeek 直接拒绝图片输入报错,然后 DeepSeek 调用视觉工具完成识别,插件将识别得到的结构化文字证据回填到会话中,DeepSeek 再基于识别结果回答用户问题。这款插件面向所有需要在 DeepSeek Harness 中使用视觉能力的用户,解决了纯文本 DeepSeek 无法处理图片输入的核心痛点。

该插件要求 Node.js 版本不低于 20,采用 MIT 许可证开源,免费供用户使用和修改。它对单张图片大小限制为 8MB,单次最多支持 8 张图片识别,进程内 LRU 缓存最多保存 64 条识别记录。本地 Ollama 模式下图片不会流出本机,API 密钥仅保存在本地配置文件中,不会被记录到日志中,隐私性有保障。

Pre-install check安装前体检Compatibility · Security兼容性 · 安全性 1 warning1 项注意
  • Only 2 stars - very few users, little community feedback星标只有 2,几乎没人在用,遇到问题缺少社区反馈
DSH walks through these 9 checksDSH 会逐条核对这 9 项

Compatibility兼容性

  • DSH, Node, OS and profile requirementsDSH 版本 / Node 版本 / 操作系统 / profile 是否满足要求
  • External dependencies and runtimes (Electron / Python / Docker, ...)外部依赖与运行时(Electron / Python / Docker 等)是否齐备
  • Conflicts with installed plugins: command names, skill / tool names, ports, duplicate MCP registration与已装插件是否冲突:命令名、skill / tool 重名、端口占用、重复 MCP 注册

Security安全性

  • Repo matches the facts registered here; archived or abandoned?仓库是否与页面登记一致,是否归档或长期停更
  • Safety of preinstall / install / postinstall and install.sh / setup.ps1preinstall / install / postinstall 与 install.sh、setup.ps1 是否安全
  • curl|bash, download-then-execute, obfuscation, unrelated domains → stop immediatelycurl|bash、下载即执行、混淆代码、无关域名 → 立刻停止
  • Typosquatting or unmaintained packages among the new dependencies新增依赖里有没有 typosquatting 或无人维护的包
  • Requested permissions vs. what the feature actually needs申请了哪些权限、是否超出功能所需(filesystem / network / shell / clipboard)
  • Any sudo / admin requirement, plus uninstall and rollback是否要求 sudo / 管理员权限,以及卸载与回滚方式

Anything uncertain must be marked unknown with a note on how to confirm it. This site's signal screen is a static snapshot, not a security audit.拿不准的必须标「未知」并说明要我怎么确认。本站的信号筛查是静态快照,不能替代安全审计。

Or use CLI install (for developers)或使用命令行安装(适合开发者)

CLI Install命令行安装

dsh plugin --profile web add dsh-youreyes

把 54xkeee/dsh-youreyes 加入你的 DSH 配置(web profile)即可启用。

READMEREADME

👁️ dsh-youreyes

English | 简体中文

给纯文本 DeepSeek 一双眼睛。 在 DeepSeek Harness 里粘贴图片、截图、文件路径——模型就能"看见"并回答与图片相关的问题。DeepSeek 依然是大脑,识图只是眼睛。

awesome · DSH plugin npm version CI MIT Node >=20 GitHub stars

一句话:DeepSeek 不会看图?装上它就会了——粘贴、识图、回答,三步完成。

✨ 为什么值得用

痛点 dsh-youreyes 的解法
DeepSeek 纯文本,粘贴图片直接被拒 包装适配器声明图片输入,图片自动转文本占位,不再报错
别的识图插件只认自家 API 反重力(默认)+ 任意 OpenAI 兼容端点 + Gemini + 本地 Ollama,你的 key 都能用
配置麻烦、要注册一堆东西 本地 Ollama 零配置自动检测;有 Gemini 免费 key 一行配置即可
模型看到图但"忘了" 视觉证据记忆:识图结果写入会话,后续轮次可复用,压缩后自动恢复
同一张图反复花钱识别 内容哈希缓存:同图同问进程内只识别一次
复杂画面识别不准 auto 档位分流:先标准检查,画面复杂自动升级深度检查
WSL / 被墙环境连不上 API winCurl 自动降级:fetch 失败自动走 Windows curl 重试

🎯 真实效果(2026-08-15 实测,全链路真实调用)

输入:一张橘猫照片 + 这是什么动物?它在做什么?请用中文回答。

输出(Gemini 通道):

这是一只橘色虎斑猫(橘猫)。它正四脚朝天、肚皮朝上地仰卧在黑色床单/毯子上安稳地睡觉,姿态非常放松惬意。

输出(OpenAI 兼容通道 · 通义 qwen3.7-flash):

这是一只橘猫(或者叫橘色虎斑猫)。它正四脚朝天地仰面躺在深色的床单(或毯子)上。它闭着眼睛,看起来睡得很沉或很香;阳光照在它身上形成了明显的光影;四肢完全伸展,呈现出一种非常舒展、毫无防备的姿态。

你粘贴图片 + 提问
  → dsh-youreyes 把图片转成占位符,DeepSeek(大脑)看到
  → DeepSeek 调用 vision 工具 → 通用视觉通道识图
  → 文字证据回填 → DeepSeek 继续回答

🧠 识图引擎——复杂识图到底怎么工作

天真的"看图→描述"循环在密集截图、表格、UI 原型和多图对比面前会崩。dsh-youreyes 把识图做成了结构化、自动升级、带记忆的流水线:

1. 档位自动升级——它知道一遍不够

detail: auto 走两遍策略:

第一遍(standard + 预判)──▶ complexity == "simple" ──▶ 完成
                            └─▶ complexity == "complex" ──▶ 第二遍(deep)──▶ escalated 结果

Showing the opening section of the README — the full document lives in the repository以上为 README 开头摘要,完整文档在仓库内 · View the full README on GitHub →在 GitHub 查看完整 README →

← 上一个 Prev deepseek-harness-plugin-mcp 下一个 Next dsh-prism-plugin →