sunshine-lang/dsh-pdf 预览 preview

sunshine-lang/dsh-pdf

Plugin插件 Native原生 ⭐ 7 MIT Docs & Writing文档与写作

DeepSeek Harness 的 PDF 工具箱:通过 pdfjs-dist 提取文本、元数据和页面范围(本地运行,无需 API 密钥)

Project Overview项目介绍

This is a native plugin built exclusively for DeepSeek Harness, designed to handle local PDF file processing for DSH agents. It uses PDF.js to parse PDF files entirely on your local machine, so no external API key or internet connection is required to extract text, metadata, and selected page content from PDFs. It exposes a pdf_read tool that returns text with page number markers, supports selecting custom page ranges to handle large documents in chunks, and enforces built-in limits to avoid oversized outputs. It reads files via DSH’s native file system interface, so it automatically inherits DSH’s permission and sandbox policies.

After you install the plugin and restart your DSH profile, you can start using it right away by instructing the DSH model to work with your local PDF files. For example, you can ask the model to summarize the first three pages of a research paper, or pull the exact content of a specific page from a legal contract. When the output hits the configured character or page limit, the plugin truncates the result cleanly at a page boundary and prompts the model to request the remaining content in a follow-up call. This structured workflow makes it easy to work with even large PDF documents without overwhelming the model’s context window.

The plugin is distributed with pre-built output code, so you don’t need any local build permissions or tools to install it from GitHub, npm, or a local source directory. A standard installation via the DSH plugin command automatically installs all required dependencies, including the PDF.js distribution library. You can customize the default limits for maximum file size, pages per call, and characters per call by adding a patch configuration to your DSH profile. The plugin is open source under the permissive MIT license, so you can modify it freely to fit your specific use case.

这是一个专为 DeepSeek Harness 开发的原生插件,提供 PDF 文件处理工具,可以在本地提取 PDF 的文本内容、文档元数据和指定页码范围的内容,基于 PDF.js 实现,无需 API 密钥,也不需要联网。插件提供 pdf_read 工具,可逐页提取带页码标记的文本,支持指定页码范围读取大文档,自带多种边界控制规则。

安装完成后,用户可以在 DSH 启动后向模型提问,让模型读取本地 PDF 的指定页码内容,比如总结论文前 3 页,或是查看合同第 7 页的具体内容。当单次读取结果超过长度限制时,插件会在页边界处截断结果并提示模型,可以通过指定新的页码范围继续读取剩余内容,适合需要用 DSH 代理处理本地 PDF 文档的用户。

插件已经预构建完成,可以直接通过 DSH 的插件命令从 GitHub 或 npm 安装,安装后无需额外构建即可使用。用户可以通过配置文件自定义最大文件字节数、单次解析最大页数和单次调用最大字符数等限制,适配不同场景需求。标准安装会自动处理所有依赖,插件以 MIT 开源许可证发布。

Pre-install check安装前体检Compatibility · Security兼容性 · 安全性 1 warning1 项注意
  • Only 7 stars - very few users, little community feedback星标只有 7,几乎没人在用,遇到问题缺少社区反馈
DSH walks through these 9 checksDSH 会逐条核对这 9 项

Compatibility兼容性

  • DSH, Node, OS and profile requirementsDSH 版本 / Node 版本 / 操作系统 / profile 是否满足要求
  • External dependencies and runtimes (Electron / Python / Docker, ...)外部依赖与运行时(Electron / Python / Docker 等)是否齐备
  • Conflicts with installed plugins: command names, skill / tool names, ports, duplicate MCP registration与已装插件是否冲突:命令名、skill / tool 重名、端口占用、重复 MCP 注册

Security安全性

  • Repo matches the facts registered here; archived or abandoned?仓库是否与页面登记一致,是否归档或长期停更
  • Safety of preinstall / install / postinstall and install.sh / setup.ps1preinstall / install / postinstall 与 install.sh、setup.ps1 是否安全
  • curl|bash, download-then-execute, obfuscation, unrelated domains → stop immediatelycurl|bash、下载即执行、混淆代码、无关域名 → 立刻停止
  • Typosquatting or unmaintained packages among the new dependencies新增依赖里有没有 typosquatting 或无人维护的包
  • Requested permissions vs. what the feature actually needs申请了哪些权限、是否超出功能所需(filesystem / network / shell / clipboard)
  • Any sudo / admin requirement, plus uninstall and rollback是否要求 sudo / 管理员权限,以及卸载与回滚方式

Anything uncertain must be marked unknown with a note on how to confirm it. This site's signal screen is a static snapshot, not a security audit.拿不准的必须标「未知」并说明要我怎么确认。本站的信号筛查是静态快照,不能替代安全审计。

Or use CLI install (for developers)或使用命令行安装(适合开发者)

CLI Install命令行安装

dsh plugin --profile web add dsh-pdf

把 sunshine-lang/dsh-pdf 加入你的 DSH 配置(web profile)即可启用。

READMEREADME

dsh-pdf

English | 中文

DeepSeek Harness PDF 工具箱:从 PDF 文件中提取文本、元数据与页码范围。本地解析(基于 PDF.js / pdfjs-dist)——无需 API key、无需网络。

功能特性

  • pdf_read 工具:逐页提取文本,带页码标记。
  • 页码选择:"1-3,5" 或 "all"——大文档可分块读取。
  • 文档元数据(标题)与总页数。
  • 内置边界控制:文件字节上限、单次解析页数上限、单次调用字符上限——截断行为明确,并提示模型如何继续。
  • 通过 harness 文件系统接缝(ctx.fs)读取,部署的权限与沙箱策略自动生效。

安装

从 GitHub 安装

dsh plugin --profile web add "github:sunshine-lang/dsh-pdf"

然后重启 dsh --profile web。lib/ 已预构建并提交,安装无需构建权限。

从 npm 安装

dsh plugin --profile web add dsh-pdf

从本地源码安装(开发)

dsh plugin --profile web add ./dsh-pdf

DSH 官方运行组件由宿主提供,插件不会另装一套旧版核心。开发时在本仓库运行 npm install --ignore-scripts 安装构建依赖;本地 link: 安装不会自动安装插件自身的 pdfjs-dist,请先在插件目录安装依赖。

兼容性

0.1.1 已在 DSH 0.1.5-rc.2 和 0.1.6-alpha.2 的一次性最小 Profile 中通过打包安装、启动、PDF 读取与卸载验收。0.1.6-alpha.1 保留为 unknown。Node.js 要求 >=22.19;实际测试环境为 macOS arm64、Node.js 22.23.1。

完整范围、复现命令与限制见 兼容性验证记录。这些结果不代表其他系统、Web UI 或模型调用已验收。

使用方法

启动 Web UI 后,向模型提问,例如:

读取 paper.pdf 的前 3 页并总结。

contract.pdf 第 7 页写了什么?

模型会调用 pdf_read:参数 path(必填),可选 pages("1-3,5" 或 "all")。输出达到上限时会在页边界截断并给出提示,模型会用页码范围继续读取。

配置

可通过 cordis.patch.yml 或 profile 的 patch 层覆盖任意配置项:

- patch:
    - id: dsh-pdf
      config:
        maxFileBytes: 52428800
        maxPages: 200
        maxCharsPerCall: 20000
配置项 默认值 含义
maxFileBytes 20971520(20 MiB) 整个 PDF 的字节上限(含);超过直接报错
maxPages 500 单次调用最多解析的页数;超出则截断并提示
maxCharsPerCall 12000 单次调用返回的最大字符数;结果在页边界截断

配置无效时插件加载会直接失败,并给出可操作的错误信息。

开发

npm install        # 或 pnpm install
npm run build      # tsc → lib/

在 DeepSeek Harness 仓库内构建(类型解析指向工作区源码)时,改用 tsconfig.local.json:tsc -p tsconfig.local.json。

测试:tests/fixtures/sample.pdf 由 make-test-pdf.mjs 生成(无依赖);w3-dummy.pdf 为 W3C 官方测试文件。集成测试:在 harness 仓库内运行 node --import tsx/esm test-integration.ts。

同作者更多插件

该作者的全部 DeepSeek Harness 插件(统一入口):dsh-plugins

许可证

MIT。

← 上一个 Prev dsh-plugin-anydoc 下一个 Next dsh-docs-panel →