liustack/modlens 预览 preview

liustack/modlens

DeepSeek Harness 首款视觉插件,也是所有纯文本编码智能体的视觉桥梁。粘贴图片,即可获取结构化 JSON 证据(OCR、版面、语义)。 | 全网最强 DeepSeek Harness 外挂视觉插件,为 DeepSeek、GLM 等纯文本模型外挂视觉能力,粘贴图片即得结构化 JSON 证据(OCR、版面、语义)。

Project Overview项目介绍

ModLens is a vision plugin built natively for DeepSeek Harness (DSH) that adds image reading capabilities to text-only large language models. It lets users paste images directly into chat conversations without requiring users to first save images to a local file and pass the file path manually. It can be installed on DSH with a single terminal command, and it also works across Claude Code, Codex, OpenCode, and Pi AI agents. It adds no invasive changes to existing host configurations, so users can remove it any time without breaking their base setup.

ModLens is designed for any user who needs to work with images using text-only DeepSeek, GLM, or MiMo Pro models. It offers two different paste interaction modes to fit different user preferences. The first mode automatically generates a temporary file path for the pasted image and triggers the image reading tool right away. The second mode adds a dedicated “(modlens vision)” entry to the model selector, keeps thumbnails visible in chat, and processes the image when the model requests it. It auto-detects text-only models and adds the vision capability without altering native vision models.

ModLens is released under the permissive MIT open-source license, so it is free to use, modify, and distribute for both personal and commercial purposes. It does not require any local proxy daemon, extra hooks, or configuration changes to the host agent, and uninstalling only requires deleting the plugin folder. It can reuse existing multimodal model setups that users already have configured, and also supports free access via Antigravity CLI or a personal Gemini API key. Users are responsible for following the terms of service of any upstream APIs they use with ModLens. First-time users can start using it immediately after installation with zero extra configuration.

ModLens 是面向 DeepSeek Harness(DSH)的原生视觉插件,核心功能是为纯文本大模型添加图像识别能力,支持用户直接粘贴图片到对话,无需预先保存图片文件再传递文件路径。它支持 DSH 原生一键安装,同时也兼容 Claude Code、Codex、OpenCode、Pi 等多个主流 AI 代理平台,用户可通过终端一行命令快速完成安装配置。

对于需要让纯文本大模型处理图片内容的开发者和普通用户来说,ModLens 提供两种粘贴交互模式:一种是粘贴后自动生成临时文件路径调用工具识别,另一种是在模型选择器添加带 ModLens 视觉的模型条目,粘贴后保留缩略图,更接近原生多模态的使用体验。它会自动识别已配置的文本模型并完成适配,不会干扰原生多模态模型的原有功能。

ModLens 采用 MIT 开源许可证,完全免费使用,对原有宿主配置无侵入修改,不会添加多余的钩子、代理或守护进程,卸载仅需删除对应插件文件夹即可。它复用用户已有的多模态模型配置,也支持免费接入通道,首次安装后无需额外配置即可直接使用图片识别功能。

Pre-install check安装前体检Compatibility · Security兼容性 · 安全性 no risk signal found未发现风险信号
  • This site's static screen found no obvious risk signal (stars, license, activity, manifest)本站静态筛查没发现明显风险信号(星标、许可证、更新活跃度、清单完整度)
DSH walks through these 9 checksDSH 会逐条核对这 9 项

Compatibility兼容性

  • DSH, Node, OS and profile requirementsDSH 版本 / Node 版本 / 操作系统 / profile 是否满足要求
  • External dependencies and runtimes (Electron / Python / Docker, ...)外部依赖与运行时(Electron / Python / Docker 等)是否齐备
  • Conflicts with installed plugins: command names, skill / tool names, ports, duplicate MCP registration与已装插件是否冲突:命令名、skill / tool 重名、端口占用、重复 MCP 注册

Security安全性

  • Repo matches the facts registered here; archived or abandoned?仓库是否与页面登记一致,是否归档或长期停更
  • Safety of preinstall / install / postinstall and install.sh / setup.ps1preinstall / install / postinstall 与 install.sh、setup.ps1 是否安全
  • curl|bash, download-then-execute, obfuscation, unrelated domains → stop immediatelycurl|bash、下载即执行、混淆代码、无关域名 → 立刻停止
  • Typosquatting or unmaintained packages among the new dependencies新增依赖里有没有 typosquatting 或无人维护的包
  • Requested permissions vs. what the feature actually needs申请了哪些权限、是否超出功能所需(filesystem / network / shell / clipboard)
  • Any sudo / admin requirement, plus uninstall and rollback是否要求 sudo / 管理员权限,以及卸载与回滚方式

Anything uncertain must be marked unknown with a note on how to confirm it. This site's signal screen is a static snapshot, not a security audit.拿不准的必须标「未知」并说明要我怎么确认。本站的信号筛查是静态快照,不能替代安全审计。

Or use CLI install (for developers)或使用命令行安装(适合开发者)

CLI Install命令行安装

npx -y @deepseek-ai/dsh plugin --profile web add @liustack/modsearch@latest

把 liustack/modlens 加入你的 DSH 配置(web profile)即可启用。

READMEREADME

ModLens

ModLens

Give a text-only model sight, and just paste the image.

🥇 The most capable vision plugin for DeepSeek Harness (dsh) 🥇

简体中文 · Configuration · Troubleshooting · Security · 🔍 ModSearch (the best free web search plugin for DSH)

Follow @liustack on X npm Node.js License Not backed by Y Combinator Users unknown

DeepSeek's flagship chat models, and GLM-5.3 itself, are text-only and cannot read images. GLM-5.3-Flash is native multimodal. ModLens is a plug-in vision engine that gives a text-only model sight. ModLens reads images pasted straight into the chat, no saving to a file and passing a path first.

Talk to us

Issues are welcome any time: open one. Follow the liustack WeChat official account, and come find me on X: @liustack. What you built with it, which harness you are on, and what should come next are all shared on WeChat and X. A proper community space is on the way.

Showing the opening section of the README — the full document lives in the repository以上为 README 开头摘要,完整文档在仓库内 · View the full README on GitHub →在 GitHub 查看完整 README →

← 上一个 Prev petdex 下一个 Next treg →