jing-hy/picturereader 预览 preview

jing-hy/picturereader

DSH插件:面向纯文本模型的像素到文本图像读取。包含image_scan/image_ocr/image_sample工具及图像读取技能(基于34张图像训练的方法论)。纯本地运行,可选PaddleOCR。

Project Overview项目介绍

picturereader is a natively developed plugin built exclusively for DeepSeek Harness (DSH). It adds full-stack image reading, OCR, document conversion, and local image editing capabilities to text-only DeepSeek large language models that lack native visual encoders. It uses a visual twin adapter to wrap text-only models in-place, enabling native DSH thumbnail rendering for pasted images and allowing image blocks to enter the conversation seamlessly. DSH users can install it directly via the DSH plugin marketplace with the command dsh plugin add picturereader.

The plugin ships with three distinct routing modes to match different use cases. Privacy mode blocks all external API calls, so all image processing stays on the local device, making it ideal for sensitive images like ID cards and contracts, as well as fully offline working environments. Smart mode is enabled by default: it prioritizes local processing and only calls external VLM APIs when the image content is complex enough to require semantic understanding. Strict mode prioritizes accuracy over speed, using cross-verification across multiple sources for high-stakes tasks like auditing and proofreading.

picturereader is released under the permissive MIT open source license. It supports four OCR engines, with platform-native engines for Windows and macOS, plus cross-platform PaddleOCR and RapidOCR, and allows optional configuration for an external VLM bridge to enhance semantic understanding of image content. The current v3.3.3 release is compatible with multiple DSH version lines from 0.1.1-rc.2 up to 0.1.3-alpha.1. Users running DSH 0.1.3 or newer must use v3.3.2 or later to avoid full plugin tree loading failures.

picturereader 是专为 DeepSeek Harness (DSH) 开发的原生插件,为纯文本大模型提供看图、读文档、OCR 文字识别和本地修图的全栈能力。它通过视觉孪生适配器原位包装纯文本模型,让模型支持图片输入,解锁 DSH 原生缩略图渲染、图片块进入会话的能力,还内置三模式路由策略和完整的本地像素级工具链。

插件提供三种路由模式满足不同场景需求:隐私模式强制禁止调用任何外部视觉 API,所有分析都在本地完成,适合敏感图片和离线场景;智能模式默认开启,优先用本地工具处理,仅在内容复杂需要语义理解时才调用外部 VLM;严谨模式优先保证准确率,结合多源证据交叉验证,适合需要高可靠结论的审图、校对场景。

本插件采用 MIT 许可证开源,支持四种 OCR 引擎,包括平台原生的 Windows、macOS 引擎,以及跨平台的 PaddleOCR 和 RapidOCR,还可配置可选的外部 VLM 桥接增强能力。当前 v3.3.3 版本兼容 DSH 0.1.1-rc.2 到 0.1.3-alpha.1 多个版本线,提醒 DSH 0.1.3 及以上用户必须使用 v3.3.2 及以上版本,避免插件加载失败。

Pre-install check安装前体检Compatibility · Security兼容性 · 安全性 1 note1 项提示
  • 37 stars - an early-stage project星标 37,属于早期项目
DSH walks through these 9 checksDSH 会逐条核对这 9 项

Compatibility兼容性

  • DSH, Node, OS and profile requirementsDSH 版本 / Node 版本 / 操作系统 / profile 是否满足要求
  • External dependencies and runtimes (Electron / Python / Docker, ...)外部依赖与运行时(Electron / Python / Docker 等)是否齐备
  • Conflicts with installed plugins: command names, skill / tool names, ports, duplicate MCP registration与已装插件是否冲突:命令名、skill / tool 重名、端口占用、重复 MCP 注册

Security安全性

  • Repo matches the facts registered here; archived or abandoned?仓库是否与页面登记一致,是否归档或长期停更
  • Safety of preinstall / install / postinstall and install.sh / setup.ps1preinstall / install / postinstall 与 install.sh、setup.ps1 是否安全
  • curl|bash, download-then-execute, obfuscation, unrelated domains → stop immediatelycurl|bash、下载即执行、混淆代码、无关域名 → 立刻停止
  • Typosquatting or unmaintained packages among the new dependencies新增依赖里有没有 typosquatting 或无人维护的包
  • Requested permissions vs. what the feature actually needs申请了哪些权限、是否超出功能所需(filesystem / network / shell / clipboard)
  • Any sudo / admin requirement, plus uninstall and rollback是否要求 sudo / 管理员权限,以及卸载与回滚方式

Anything uncertain must be marked unknown with a note on how to confirm it. This site's signal screen is a static snapshot, not a security audit.拿不准的必须标「未知」并说明要我怎么确认。本站的信号筛查是静态快照,不能替代安全审计。

Or use CLI install (for developers)或使用命令行安装(适合开发者)

CLI Install命令行安装

dsh plugin --profile web add picturereader

把 jing-hy/picturereader 加入你的 DSH 配置(web profile)即可启用。

READMEREADME

picturereader

v3.4.0 — 给纯文本模型(DeepSeek / text-only)的全能「看图 / 读文档 / 修图」能力。 融合 视觉孪生 adapter(把任意文本模型原位包装成「支持图片」→ DSH 原生缩略图 + 图片块自动分析)、三模式路由、本地像素级工具链(scan / OCR×4 引擎 / crop / palette / compare / batch)、文档转图片(pdf / word / excel / ppt)、本地修图工具 image_edit(Pillow/OpenCV 纯 CPU:缩放 / 旋转 / 滤镜 / 合成 / 水印 / 去背景 / 超分等)与可选外部 VLM 桥。一个插件全包。

v3.4.0 新增:原生识图能力感知。插件现在从模型能力元数据(inputModalities,与 DSH 内核的图片门控同源)判断当前会话的模型是否自己就能看图:命中时在系统提示词里明确说明「不要用 image_scan / image_sample 这类"伪多模态"替代路径,但 image_ocr 仍应正常使用」,并让粘贴/读取的图片直通该模型、不再降级成文本引导;纯文本模型与元数据缺失(未知)时行为与 v3.3.3 完全一致。判定走未包装的原始 adapter,因此不会被本插件自己的视觉孪生误判成原生识图。

v3.3.0 新增:macOS 原生 Vision OCR 引擎(engine="macos",scripts/setup-macos.mjs 一键编译,PR #4 合入);OCR 引擎选项按平台条件显示(macos 仅 macOS、windows 仅 Windows,paddle/rapid 跨平台始终显示);修复 PaddleOCR 新环境首次调用三个缺陷(stdout 污染 / w/h→width/height / 缓存路径写死,issue #2);设置卡 UI 重做(settings-panel 设计语言);调试日志门控(llm/stream 桥不再刷屏);peerDependencies 兼容 DSH 0.1.1-rc.2(issue #3)。

dsh-plugin dsh.so security dsh.so install


定位

DeepSeek 等纯文本模型没有视觉编码器,无法直接看图片;DSH 原生缩略图也需要模型被声明为「支持图片」才会渲染。

⭐ 已支持外部视觉 API(OpenAI 兼容端点 / LM Studio / 云端 VLM),由 LLM 自行按需调用:配置好端点后,模型会在智能/严谨模式下自主判断"这张图值不值得外呼视觉模型",需要时用 vision_analyze 调外部 API 做语义理解,简单内容则本地像素/OCR 搞定——外部 API 是即插即用的增强能力,不是必须依赖。

picturereader 现在解决三件事:

  1. 把「看图/读文档」翻译成纯文本模型能理解的结构化证据(像素级 hue/结构/材质分析 + OCR 实读 + 可选 VLM 语义描述),并沉淀为读图方法论 skill。
  2. 通过「视觉孪生 adapter」让纯文本模型在 DSH 里获得原生缩略图体验:勾选模型即生成「(视觉)」变体,粘贴图片显示原生缩略图、图片块进会话、并被自动分析成文本路径 + 本地证据再交给模型。即使上游先降级为 attachment sha256 文本,图片桥也会仅从本机附件对象库恢复经过文件头校验的图片并注入本地工具路径;模型拿到的永远是纯文本,不会触发 UNSUPPORTED_CONTENT。
  3. 本地直接修图 / 批量处理图片:image_edit 提供缩放、旋转、滤镜、合成、水印、去背景、拼接、透视校正等纯 CPU 动作,图片不出本机。

版本兼容性:本版本专门兼容 dsh 0.1.3-alpha.1 / dsheac 5.4.0(并实测兼容 dsh 0.2.0-rc.2),同时向后兼容 dsh 0.1.1-rc.2 与 dsheac 5.1.0,以及 DeepSeek Harness EAC 4.2.0 与 @deepseek-ai/dsh-client-ui-workspace rc.7。peerDependencies 采用 ^0.1.0-rc.6 || ^0.1.1-rc.2 || ^0.2.0-rc.2 || ^0.1.3-alpha.1 联合区间,覆盖 0.1.x 三条发布线并延伸到 0.2.x。

⚠️ 升级到 dsh 0.1.3 必须使用 v3.3.2 之后的版本:内核 0.1.3 起 @deepseek-ai/dsh-settings 不再导出 settingsNamespace() 品牌函数(命名空间校验收进 register() 内部)。ESM 静态导入不存在的符号会让模块整体加载失败,进而拖垮整棵插件树、触发 EAC 的 guard 安全模式(表现为「一对话就报错」+ 插件大范围消失)。v3.3.3 已适配。

Showing the opening section of the README — the full document lives in the repository以上为 README 开头摘要,完整文档在仓库内 · View the full README on GitHub →在 GitHub 查看完整 README →

← 上一个 Prev dsh-trading 下一个 Next mnemosyne →