TC635807/better-crawler4agent

Plugin插件 Native原生 ⭐ 2 MIT Search & Web Access搜索与联网

When I was using agents, I found that most agents rely on curl commands to scrape web pages, which has a very low success rate and is slow, so I made this.

Project Overview项目介绍

better-crawler-4-agent is a stdio Model Context Protocol server that turns arbitrary web URLs into clean article text for AI coding agents, using a real headless Chromium via Playwright so that JavaScript-rendered pages and Chinese anti-bot sites such as Zhihu, CSDN, Juejin, Douban, and Tencent Cloud Community are handled correctly. Its core capability is a three-tier fallback (real browser, trafilatura, httpx with full browser fingerprint), plus font honeypot detection, fake-success detection that catches 200 responses which are actually login walls or "article deleted" illustrations, automatic skeleton-page retries, SSRF protection, and per-tier diagnostics returned with every response. Integration is explicitly cross-platform: the generic .mcp.json works in Claude Code, Cursor, Windsurf, Cline, Codex, and ZCode, while the repository root ships a package.json declaring dsh.bundle.patch plus a dsh-plugin/ directory with cordis.patch.yml and install.sh, so DeepSeek Harness can install the whole repo as a bundle via dsh plugin --profile web add github:TC635807/better-crawler4agent or the local ./dsh-plugin/install.sh script, registering three MCP tools (mcp__better-crawler__fetch_url, mcp__better-crawler__fetch_urls, mcp__better-crawler__crawler_status) through @deepseek-ai/dsh-mcp-client.

The typical workflow is straightforward and aimed at developers and researchers who need a coding agent to actually read pages inside a session. A single tool call passes one URL or up to twenty concurrent URLs (the README recommends batching no more than eight at a time because most MCP clients impose a 30-second call timeout); the server warms or reuses a singleton Chromium instance, performs tiered fetching, detects skeleton pages whose HTML exceeds 30 KB but whose body text is under 1500 characters, retries them once without retrying deterministic failures like 404 and 403, and finally returns the longest successful payload along with the hit tier, character count, and an optional warning. URL discovery is deliberately out of scope, so callers must pick the pages themselves, and an accompanying vendor-neutral skill at skills/web-to-text/SKILL.md documents when to fetch, which URLs to avoid (tag pages, user pages, video pages), and how to judge whether a result is trustworthy, which agents that support skill directories can simply copy.

Dependencies, limits, and first-run caveats are spelled out clearly. The project requires Python 3.10 or newer and locks playwright==1.61.0 because Playwright binds strictly to browser revisions (1.61 → Chromium revision 1228, 1.62 → revision 1234), so upgrading Playwright without rerunning python -m playwright install chromium will silently break browser reuse. The bundled python scripts/install.py creates an in-repo .venv, probes the system for an existing matching browser to avoid re-downloading roughly 700 MB, prints ready-to-paste MCP configuration for every supported client, and offers flags such as --print-config, --reuse-env, --browsers-path, and --no-browser; verification is done via python scripts/selfcheck.py (15 online checks including real fetches) or python scripts/selfcheck.py --offline (12 offline checks), and unit tests cover SSRF, honeypot recognition, body extraction, and error-page identification. All tuning goes through environment variables including BETTER_CRAWLER_TIMEOUT (25 s), BETTER_CRAWLER_MAX_CHARS (50000), BETTER_CRAWLER_ALLOW_PRIVATE, BETTER_CRAWLER_PYTHON, BETTER_CRAWLER_BROWSERS_PATH, and BETTER_CRAWLER_LOG_LEVEL; SSRF blocks file://, raw protocols, loopback, private RFC1918 ranges, and cloud metadata 169.254.169.254 by default and only opens them when BETTER_CRAWLER_ALLOW_PRIVATE=1 is set explicitly for local debugging. The license is MIT, and the README reminds users to respect each target site's robots.txt and terms of service, warning against using the tool for high-frequency bulk scraping.

better-crawler-4-agent 是一个面向 AI 编码代理的 URL 转正文 stdio MCP server,使用 Playwright 真实 Chromium 渲染,能够处理 JS 渲染页与知乎、CSDN、掘金等中文站点。核心能力包含三层降级(浏览器→trafilatura→httpx)、字体蜜罐识别、假成功识别、SSRF 防护以及每层诊断信息。集成方式上,它同时支持 Claude Code、Cursor、Windsurf、Cline、Codex、ZCode 等通用 MCP 客户端,以及 DeepSeek Harness(dsh)插件形态——仓库根目录的 package.json 声明了 dsh.bundle.patch,dsh-plugin/ 目录下含 cordis.patch.yml 与 install.sh。

典型工作流:用户给定一条或多条 URL(最多 20 条并发),工具启动或复用已有 Chromium(锁定 Playwright 1.61.0),通过真实浏览器渲染后调用 extract.py 提取正文并自动重试一次(确定性 404/403 不重试),返回结果附带命中层级、字符数与 warning 字段。fetch_urls 建议一次不超过 8 条,因多数 MCP 客户端 30 秒调用超时。适合需要在代理工作流里直接读文章、做研究引用或解析文档站点的开发者与研究用户,不包含 URL 发现能力,需要调用方自己选 URL。

依赖与限制:Python 3.10+,Playwright 1.61.0 与 Chromium rev 1228 强绑定,升级须同步 python -m playwright install chromium,否则浏览器复用会静默失败;首次安装通过 python scripts/install.py 在仓库内建 .venv 并探测已有浏览器;SSRF 默认开启,仅允许公网 http/https,本地调试需 BETTER_CRAWLER_ALLOW_PRIVATE=1;授权 MIT;合规上要求遵守目标站点 robots.txt,禁止高频批量抓取。

Pre-install check安装前体检Compatibility · Security兼容性 · 安全性 1 warning1 项注意
  • Only 2 stars - very few users, little community feedback星标只有 2,几乎没人在用,遇到问题缺少社区反馈
DSH walks through these 9 checksDSH 会逐条核对这 9 项

Compatibility兼容性

  • DSH, Node, OS and profile requirementsDSH 版本 / Node 版本 / 操作系统 / profile 是否满足要求
  • External dependencies and runtimes (Electron / Python / Docker, ...)外部依赖与运行时(Electron / Python / Docker 等)是否齐备
  • Conflicts with installed plugins: command names, skill / tool names, ports, duplicate MCP registration与已装插件是否冲突:命令名、skill / tool 重名、端口占用、重复 MCP 注册

Security安全性

  • Repo matches the facts registered here; archived or abandoned?仓库是否与页面登记一致,是否归档或长期停更
  • Safety of preinstall / install / postinstall and install.sh / setup.ps1preinstall / install / postinstall 与 install.sh、setup.ps1 是否安全
  • curl|bash, download-then-execute, obfuscation, unrelated domains → stop immediatelycurl|bash、下载即执行、混淆代码、无关域名 → 立刻停止
  • Typosquatting or unmaintained packages among the new dependencies新增依赖里有没有 typosquatting 或无人维护的包
  • Requested permissions vs. what the feature actually needs申请了哪些权限、是否超出功能所需(filesystem / network / shell / clipboard)
  • Any sudo / admin requirement, plus uninstall and rollback是否要求 sudo / 管理员权限,以及卸载与回滚方式

Anything uncertain must be marked unknown with a note on how to confirm it. This site's signal screen is a static snapshot, not a security audit.拿不准的必须标「未知」并说明要我怎么确认。本站的信号筛查是静态快照,不能替代安全审计。

Or use CLI install (for developers)或使用命令行安装(适合开发者)

CLI Install命令行安装

dsh plugin --profile web add github:TC635807/better-crawler4agent

把 TC635807/better-crawler4agent 加入你的 DSH 配置(web profile)即可启用。

READMEREADME

better-crawler-4-agent

License: MIT Python 3.10+ MCP Playwright

把 URL 变成干净正文的 MCP server。用真实浏览器渲染,知乎、CSDN、掘金这些"HTTP 抓回来是空壳"的站点也能拿到内容。

一个标准 stdio MCP server,不绑定任何客户端——Claude Code、Cursor、Windsurf、Cline、Codex、ZCode 都能用。

English summary

A vendor-neutral MCP server that turns URLs into clean article text using a real headless Chromium. Built for sites where plain HTTP fetching fails: JS-rendered pages and Chinese anti-bot sites like Zhihu, CSDN, and Juejin.

Scope: URL → text only. URL discovery is left to the caller, so the tool stays single-purpose and needs no search API keys.

Key features: real browser rendering, tiered fallback (browser → trafilatura → httpx+full browser fingerprint), font honeypot detection, fake-success detection (200 responses that are actually error pages), SSRF protection, and per-tier diagnostics in every response.


为什么需要它

大多数 Agent 自带的网页抓取是 HTTP 客户端 + 正文提取。对静态站点够用,但遇到下面三种情况会失败:

问题 表现
JS 渲染 返回空壳页面,只有导航和页脚
反爬拦截 知乎返回 403,CSDN 返回 521
提取退化 页面上有内容,提取出来却是零散片段

本项目用真实 Chromium 渲染,并补齐完整浏览器指纹(UA / Sec-Ch-Ua / Sec-Fetch-* 全套一致)+ 导航器覆盖。

实测对比

同一组 6 个 URL,开浏览器 vs 关浏览器(本项目 Fetcher,2026-09 实测):

成功率 知乎 CSDN 字符数 博客园字符数
有浏览器 6/6 = 100% 27069 57218 15360
无浏览器(仅静态路径) 5/6 = 83% 失败 13845(−76%) 4627(−70%)

反爬站点直接失败,其他站点正文完整度也大幅下降——没有 JS 渲染只能拿到首屏片段。

已实测可抓取的站点

知乎专栏、CSDN、博客园、掘金、豆瓣读书、少数派、腾讯云社区、React 官方文档、Python 官方文档等。单页耗时 2–5 秒(浏览器暖启动后)。


快速开始

git clone git@github.com:TC635807/better-crawler4agent.git
cd better-crawler4agent
python scripts/install.py

install.py 会自动:

  1. 在仓库内建 .venv 并安装依赖
  2. 查找并复用系统中已有的 Playwright 浏览器(核对版本,避免重复下载 ~700MB)
  3. 打印各 agent 的接入配置,复制粘贴即可

验证安装:

Showing the opening section of the README — the full document lives in the repository以上为 README 开头摘要,完整文档在仓库内 · View the full README on GitHub →在 GitHub 查看完整 README →

← 上一个 Prev new-project-init 下一个 Next dsh-eyes →