zyh20041227/improved_vision_for_deepseek 预览 preview

zyh20041227/improved_vision_for_deepseek

Full-coverage image tiling for DeepSeek Harness vision models, dense-text OCR, and document AI

项目介绍Project Overview

DSH Vision Tiler 是面向 DeepSeek Harness 视觉模型的图像分块插件。它把高分辨率图生成总览、重叠覆盖块和可选密集区域裁剪,并分批返回坐标与覆盖信息。适用于收据、表格、图表、长截图等小字易被缩放丢失的 OCR/文档识别场景。注意:它只保证几何像素覆盖,不保证模型能正确识别每个字符。

DSH Vision Tiler is a DeepSeek Harness plugin for vision models. It converts a high-resolution image into a downscaled overview, overlapping full-coverage tiles, and optional dense-region crops, returned in model-safe batches with coordinates and coverage metadata. Use it for receipts, tables, diagrams, screenshots, and document OCR where small text may be lost to image scaling. It guarantees geometric pixel coverage, not perfect semantic recognition.

或使用命令行安装(适合开发者)Or use CLI install (for developers)

命令行安装CLI Install

dsh plugin --profile web add github:zyh20041227/improved_vision_for_deepseek#v0.2.3

zyh20041227/improved_vision_for_deepseek 加入你的 DSH 配置(web profile)即可启用。

READMEREADME

DSH Vision Tiler

中文说明 · Technical report · Dense-text benchmark

Full-coverage image tiling for DeepSeek Harness (DSH) vision models. The plugin turns one high-resolution image into a global overview, overlapping coverage tiles, and an optional dense-region detail crop before the model reads it.

Dense-text benchmark comparison

Why it exists

DeepSeek documents a maximum of 384 vision tokens per image and scales large images before inference. That budget is often enough for ordinary photos, but it can remove small characters from receipts, tables, diagrams, and long screenshots. DSH Vision Tiler gives each local region its own image budget while preserving a complete, auditable view of the source.

  • 100% geometric coverage: coverage tiles are audited against every source pixel.
  • Overlapping seams: text and shapes that cross a tile edge remain visible in a neighbour.
  • Content-aware cuts: document mode moves horizontal seams toward low-ink areas.
  • Dense-region review: an optional 512×512 detail crop supplements, never replaces, base coverage.
  • Bounded batches: large tile sets are returned in model-safe batches.
  • Traceable output: each tile carries source coordinates, role, batch state, coverage, and a conservative token cap.
  • No native image build: v0.2.3 uses pure JavaScript plus bundled WebAssembly, avoiding sharp/libvips conflicts inside DSH Web.

Install

Requirements: DeepSeek Harness, a vision-capable DSH model profile, and Node.js 22 or later.

Pinned GitHub release:

dsh plugin --profile web add github:zyh20041227/improved_vision_for_deepseek#v0.2.3
dsh --profile web --dump-config

No allow-build entry is required. Runtime packages are declared in package.json and locked in package-lock.json; these files are the Node.js equivalent of Python's requirements.txt. An npm registry release is planned but is not yet published, so use the pinned GitHub command above.

DSH Web showing Vision Tiler enabled

Use

Ask the model to call the registered segment_image tool and continue through every returned batch:

Call segment_image for D:\images\document.png with mode=document and batch_index=0.
If remaining_batch_indices is not empty, read every remaining batch before answering.
Report uncertain_regions and cite the tile IDs used.
Argument Meaning
path Absolute path, or a path relative to the DSH process directory
mode auto, document, diagram, or photo
strategy adaptive (default) or uniform (control mode)
batch_index Zero-based output batch

The DSH profile must use a model that accepts image attachments. A text-only route can run the tiler, but it cannot pass the resulting images to the model.

Measured results

The controlled dense-text benchmark contains four synthetic pages with 100 unique eight-character codes each. Every model arm read each page independently three times: 12 calls and 1,200 exact-code decisions per arm.

Configuration Exact-code F1 Median latency Mean total tokens/call Estimated cost/call
GPT-5.5 99.25% 27.95 s Not exposed Codex subscription; not convertible
GPT-5.6 Terra 98.67% 25.17 s Not exposed Codex subscription; not convertible
DeepSeek + plugin 96.44% 6.08 s 4,970.8 ¥0.004521 observed-cache estimate
GPT-5.6 Luna 94.99% 27.48 s Not exposed Codex subscription; not convertible
DeepSeek direct image 19.68% 6.87 s 1,135.5 ¥0.001830 estimate

For this task, tiling increased the estimated DeepSeek charge per call by about 2.47×, but reduced estimated cost per 100 correct codes by about 51%. DeepSeek's experimental vision model has no separate public price row, so these values use the published V4 Flash rates and are estimates, not invoices. Codex does not expose per-task vision tokens or billable API cost here, so GPT prices are intentionally not guessed.

Does the 384-token cap reduce reading quality?

It can, especially when a large image contains small, low-contrast, or tightly packed text. In the controlled test, the model and prompt stayed the same while the input changed from one scaled image to complete local tiles; F1 rose from 19.68% to 96.44%. This is strong engineering evidence for that workload, not a claim that every image needs tiling.

How it works

  1. Decode PNG/JPEG/BMP/GIF/TIFF with Jimp and WebP with bundled WASM.
  2. Apply EXIF orientation and reject images above the 100-million-pixel safety limit.
  3. Generate a downscaled overview.
  4. Plan overlapping tiles whose union covers the full oriented source.
  5. In adaptive document mode, move seams toward low-density rows and select optional dense details.
  6. Audit coverage, encode PNG attachments, and return bounded batches with coordinates.

The plugin guarantees geometric pixel coverage. It cannot guarantee that a model semantically recognises every visible character; blurred input, compression artefacts, unusual fonts, and model errors still require review.

Evidence and reproducibility

The public repository contains aggregate results and reproducible generators, but never API keys or local caches.

Development

npm install
npm test
npm pack

The test suite covers exact geometric coverage, seam overlap, safety caps, deterministic batching, adaptive detail selection, WebP/WASM decoding, EXIF orientation, and DSH tool rendering. See CONTRIBUTING.md and SECURITY.md.

License

MIT

上一个 Prev dsh--prompt--enhance 下一个 Next dsh-HoldThatBigBlueFatFish