2672243194/dsh-read-url 预览 preview

2672243194/dsh-read-url

DeepSeek Harness URL读取器:抓取任意页面并返回干净的正文内容/Markdown。自动字符集识别(GBK/GB2312/UTF-8/Big5),高效利用token(6000字符上限、缓存、偏移量),零依赖,无需API密钥。网页一键读全文 → 干净正文 / 结构化Markdown

Project Overview项目介绍

This is a native plugin built exclusively for DeepSeek Harness (DSH) that enables DSH AI agents to read content from arbitrary public URLs. It supports multiple content types including HTML web pages, plain text, JSON APIs, CSV data, and RSS/Atom feeds, and automatically handles character encoding detection for common formats like UTF-8, GBK/GB2312, UTF-16 BOM, Big5, and Shift-JIS. To install the plugin, you can add it via the DSH CLI, and it requires zero external runtime dependencies beyond the built-in Node 20+ APIs, so no extra packages or external API keys are needed to run it.

The core use case for this plugin is when a DSH agent finds a link during a web search and needs to retrieve the full content of the linked page for further processing. Unlike DSH's official web_fetch tool that returns the entire page including navigation, ads, and sidebars that bloat token usage, this plugin cleans the content to only return what the AI model actually needs to complete its task. It is designed for AI agents working on coding, research, documentation writing, and any task that requires pulling structured or unstructured content from external web sources.

dsh-read-url is released under the permissive MIT license, is 100% open source, and performs all processing locally with no external data collection. It has built-in session caching with a 5-minute TTL to speed up repeated requests, and it follows all DSH architecture guidelines for capability seams, effect cleanup, and cooperative timeouts. The project entered maintenance mode at v1.0.0, so only bug fixes are released moving forward, and new feature updates are rare.

这是一款专门为DeepSeek Harness(DSH)构建的原生URL读取插件。它可以从任意URL获取网页、JSON API、RSS/Atom订阅源内容,自动检测字符编码,提取干净的正文内容,自动拼接分页文章,输出节省token的紧凑文本或结构化Markdown。它零运行时依赖,仅需Node 20及以上版本,不需要API密钥或外部服务器,安装后即可直接使用。

DSH代理本身支持搜索获取链接和摘要,但缺少读取URL并整理为干净正文的能力,这款插件填补了这个空白。对比DSH官方的web_fetch工具,它默认只返回模型需要的清理后正文和必要元数据,默认输出上限更低,不会造成token浪费,适合AI代理需要读取外部网页内容进行分析、写作、调研的工作流,目标用户是使用DSH的开发人员和研究者。

它采用MIT许可,完全开源免费,没有数据收集,所有处理都在本地完成。它支持处理多种编码格式,包括UTF-16 BOM、GBK、GB2312等常见中文编码,能解决乱码问题,默认对输出内容做段落对齐截断,还内置5分钟TTL的会话级缓存。当前插件已进入维护模式,仅进行错误修复,更新频率较低。

Pre-install check安装前体检Compatibility · Security兼容性 · 安全性 1 note1 项提示
  • 19 stars - an early-stage project星标 19,属于早期项目
DSH walks through these 9 checksDSH 会逐条核对这 9 项

Compatibility兼容性

  • DSH, Node, OS and profile requirementsDSH 版本 / Node 版本 / 操作系统 / profile 是否满足要求
  • External dependencies and runtimes (Electron / Python / Docker, ...)外部依赖与运行时(Electron / Python / Docker 等)是否齐备
  • Conflicts with installed plugins: command names, skill / tool names, ports, duplicate MCP registration与已装插件是否冲突:命令名、skill / tool 重名、端口占用、重复 MCP 注册

Security安全性

  • Repo matches the facts registered here; archived or abandoned?仓库是否与页面登记一致,是否归档或长期停更
  • Safety of preinstall / install / postinstall and install.sh / setup.ps1preinstall / install / postinstall 与 install.sh、setup.ps1 是否安全
  • curl|bash, download-then-execute, obfuscation, unrelated domains → stop immediatelycurl|bash、下载即执行、混淆代码、无关域名 → 立刻停止
  • Typosquatting or unmaintained packages among the new dependencies新增依赖里有没有 typosquatting 或无人维护的包
  • Requested permissions vs. what the feature actually needs申请了哪些权限、是否超出功能所需(filesystem / network / shell / clipboard)
  • Any sudo / admin requirement, plus uninstall and rollback是否要求 sudo / 管理员权限,以及卸载与回滚方式

Anything uncertain must be marked unknown with a note on how to confirm it. This site's signal screen is a static snapshot, not a security audit.拿不准的必须标「未知」并说明要我怎么确认。本站的信号筛查是静态快照,不能替代安全审计。

Or use CLI install (for developers)或使用命令行安装(适合开发者)

CLI Install命令行安装

npx @deepseek-ai/dsh plugin --profile web add github:2672243194/dsh-read-url

把 2672243194/dsh-read-url 加入你的 DSH 配置(web profile)即可启用。

READMEREADME

dsh-read-url

🌐 English | 中文

dsh-read-url

npm License

URL reader plugin for DeepSeek Harness: fetch any URL — webpages (HTML), JSON APIs, RSS/Atom feeds — auto-detect encoding (UTF-16 BOM / GBK/GB2312 / UTF-8 / Big5 / Shift-JIS), extract the clean main content (auto-joining paginated articles), and return token-efficient compact text or structured Markdown.

Zero runtime dependencies (Node 20+ built-ins handle fetch/decode/extract), no API key, no server side — install and use.

Supported content

  • HTML technical documentation, forums, nested articles and custom elements, with visible code examples and metadata preserved.
  • Plain text, Markdown, CSV, JSON and vendor +json APIs.
  • RSS 1.0/2.0 and Atom, including namespace prefixes, relative article links and xml:base.
  • UTF-16 BOM text without a Content-Type header.

Continue at charsStart + charsReturned; the returned charsStart may exceed the requested offset when paragraph separators are skipped. Markdown tables convert at most the first 25 rows and report omitted rows.

Why

DSH agents can search (getting links and snippets) but lack the step of "reading a URL into clean body text". The official tool-web web_fetch does a whole-page turndown conversion (nav/ads/sidebars all preserved) with a default cap of 200,000 characters — a token black hole. This plugin returns only what the model actually needs: cleaned body + essential metadata, truncated by default.

Competitor comparison (measured from source/docs, 2026-08-15)

Capability Official tool-web web_fetch dsh-webfetch dsh-scrape-webpage dsh-read-url
Body cleaning (container-level) ❌ whole page ⚠️ tag-level, nav/footer leak in ⚠️ custom, noisy ✅ article/main containers + noise stripping
Default output cap 200,000 chars 50,000 chars 30,000 chars 6,000 chars + paragraph-aligned truncation
Chinese GBK/GB2312 provider-dependent ⚠️ not normalized, GB2312 garbles ❌ not handled ✅ normalized + mojibake fallback + UTF-16 BOM + Shift-JIS
JSON / RSS / Atom URLs ❌ HTML only ❌ ❌ ✅ native compact rendering
Paginated articles ❌ manual per-page calls ❌ ❌ ✅ auto-joined (default 3 pages)
Session-level cache ❌ ❌ ❌ ✅ 5-min TTL
ctx.web seam ✅ (official core) ❌ global fetch ❌ ✅ seam-first, fallback included
ctx.effect unload cleanup ✅ ❌ ❌ ✅
Cooperative timeout (hidden from model) ✅ ⚠️ self-managed ⚠️ self-managed ✅ timeoutMs + exec.signal
Model-facing output whole-page Markdown compact text 15-field JSON compact text (no JSON parsing)
Dependencies official TypeScript build zero deps zero deps (JS ESM, drop-in)
Anti-bot / degraded responses (UA & TLS fingerprint) ⚠️ Node default UA; measured: https intercepted by middlebox TLS fingerprinting, Baidu returns a degraded page without trending topics ❓ not disclosed ❓ not disclosed ✅ full browser UA; measured: full page fetched (Baidu trending topics intact)

Showing the opening section of the README — the full document lives in the repository以上为 README 开头摘要,完整文档在仓库内 · View the full README on GitHub →在 GitHub 查看完整 README →

← 上一个 Prev dsh-minecraft-ui 下一个 Next deepseek-harness-action →