DDDFXYqiming/dsh-ocr1-memory

使用DeepSeek-OCR(OCR1)的光学压缩记忆

Project Overview项目介绍

This is a native DSH optical memory plugin, built following the Context Optical Compression (OCR1) approach introduced in the DeepSeek-OCR paper. The plugin stores text memories by rendering each paragraph into a numbered SoM image, and reduces the resolution of older memories over time. When retrieving relevant memories, it first filters candidates based on text and OCR evidence, and if an optical locator is configured, it further narrows down relevant paragraphs before returning the original verbatim text. It also includes a dedicated tool to output the actual token compression ratio, letting users see how many tokens are saved by the optical compression approach.

The plugin comes with more than 10 built-in tools that cover all core memory management workflows. These tools include functions to store, retrieve, update, and delete memories, check configuration status, view compression metrics, calibrate token baselines, and test the rendering pipeline. The default workflow cuts input text into paragraphs based on blank lines and length limits, then renders each paragraph into a SoM image. Memory resolution has three levels that decay as the memory ages, and if a low-resolution memory is retrieved, it is automatically restored to high resolution before being returned. When an optical locator is configured, the model outputs relevance labels for each paragraph, and the plugin selects paragraphs based on a threshold and Top-K rules.

The plugin is open-sourced under the BSD-3-Clause license, and can be installed directly via the DSH plugin command. The installation command is dsh plugin --profile web add github:DDDFXYqiming/dsh-ocr1-memory. After installation, users need to configure options such as storage directory, OCR backend URL, and whether to enable the optical locator in the profile’s cordis.patch.yml file. The plugin supports falling back to environment variables for configuration when the corresponding config fields are left empty. It relies on a llama-server with an OpenAI-compatible interface as the OCR backend, and can automatically start and manage the backend process when configured to do so. All 88 test cases currently pass, and tests that require a backend are automatically skipped when no backend is available.

这是一款专为DeepSeek Harness打造的原生光学记忆插件,设计思路来自DeepSeek-OCR论文的上下文光学压缩(OCR1)方法。插件将文本记忆按段落渲染为带SoM编号的图像,对老旧记忆按年龄降低分辨率,检索默认依据文本和OCR证据筛选,配置光学定位器后可进一步筛选相关段落,最终返回原始准确文本。插件还提供了查看文本token、视觉token压缩比的工具,能让用户直观看到压缩节省的token量,明确效果。

插件内置十余个工具,涵盖存储、检索、更新、删除记忆,查看状态、统计压缩指标、校准基线、测试渲染管线等功能。工作时先将文本按空行和长度切分为段落,再渲染为SoM图像,记忆分辨率分清晰、正常、模糊三级随年龄衰减,命中低清记忆会自动恢复高清后返回。配置光学定位器后,由模型输出段落相关性标签,插件按阈值和Top-K规则筛选段落,最终仅返回选中的原文段落。

插件采用BSD-3-Clause许可证开源,可通过dsh插件命令直接安装,安装命令为dsh plugin --profile web add github:DDDFXYqiming/dsh-ocr1-memory。用户需要在profile的cordis.patch.yml配置存储路径、OCR后端地址、光学定位器开关等选项,支持环境变量回退配置。插件依赖兼容OpenAI接口的llama-server作为OCR后端,可自动管理后端进程,当配置autoStartOcrServer为true时会自动启动后端,只清理自身启动的进程,所有测试用例当前全部通过,无后端时相关测试会自动跳过。

Pre-install check安装前体检Compatibility · Security兼容性 · 安全性 1 warning1 项注意
  • Only 2 stars - very few users, little community feedback星标只有 2,几乎没人在用,遇到问题缺少社区反馈
DSH walks through these 9 checksDSH 会逐条核对这 9 项

Compatibility兼容性

  • DSH, Node, OS and profile requirementsDSH 版本 / Node 版本 / 操作系统 / profile 是否满足要求
  • External dependencies and runtimes (Electron / Python / Docker, ...)外部依赖与运行时(Electron / Python / Docker 等)是否齐备
  • Conflicts with installed plugins: command names, skill / tool names, ports, duplicate MCP registration与已装插件是否冲突:命令名、skill / tool 重名、端口占用、重复 MCP 注册

Security安全性

  • Repo matches the facts registered here; archived or abandoned?仓库是否与页面登记一致,是否归档或长期停更
  • Safety of preinstall / install / postinstall and install.sh / setup.ps1preinstall / install / postinstall 与 install.sh、setup.ps1 是否安全
  • curl|bash, download-then-execute, obfuscation, unrelated domains → stop immediatelycurl|bash、下载即执行、混淆代码、无关域名 → 立刻停止
  • Typosquatting or unmaintained packages among the new dependencies新增依赖里有没有 typosquatting 或无人维护的包
  • Requested permissions vs. what the feature actually needs申请了哪些权限、是否超出功能所需(filesystem / network / shell / clipboard)
  • Any sudo / admin requirement, plus uninstall and rollback是否要求 sudo / 管理员权限,以及卸载与回滚方式

Anything uncertain must be marked unknown with a note on how to confirm it. This site's signal screen is a static snapshot, not a security audit.拿不准的必须标「未知」并说明要我怎么确认。本站的信号筛查是静态快照,不能替代安全审计。

Or use CLI install (for developers)或使用命令行安装(适合开发者)

CLI Install命令行安装

dsh plugin --profile web add github:DDDFXYqiming/dsh-ocr1-memory

把 DDDFXYqiming/dsh-ocr1-memory 加入你的 DSH 配置(web profile)即可启用。

READMEREADME

简体中文 | English

@dsh-external/dsh-ocr1-memory

A DSH optical-memory plugin built on the idea in the DeepSeek-OCR paper (Contexts Optical Compression, OCR1).

Text memories are split into paragraphs and rendered as SoM-numbered images. Older memories move through lower-resolution tiers as they age. Retrieval defaults to text and OCR evidence, an optical locator can optionally pick the relevant segments first, and the original verbatim text comes back deterministically at the end.

Rendering memory as images borrows the paper's approach to optical context compression. To see what that actually saves, ocr1_mem_metrics reports text tokens, visual tokens, and the compression ratio.

Capabilities

Tool Purpose
ocr1_mem_status Inspect storage, renderer, and OCR status
ocr1_mem_store Segment, render, and store a memory
ocr1_mem_update Replace a memory and reset its freshness
ocr1_mem_retrieve Retrieve, OCR read back, and active-recall memories
ocr1_mem_list List entries and hit counts
ocr1_mem_metrics Inspect text/visual tokens and compression ratios
ocr1_mem_calibrate Calibrate the text-token baseline
ocr1_mem_forget Delete a memory and its optical artifacts
ocr1_mem_render_test Test the rendering pipeline
ocr1_mem_embed_test Test visual embeddings
memory_read / memory_retrieve Read governed memory and use OCR1-backed retrieval
memory_write / memory_update Evidence-backed writes and updates
memory_search / memory_promote Full-text search and cross-namespace promotion
memory_pending / memory_accept Review and accept distilled candidates
memory_maintain / memory_stats Deduplicate, compact L1, and inspect state
memory_index / memory_archive / memory_rollback Rebuild the index, archive, and restore history
memory_expand / memory_activate Expand DSH provenance events and activate governance

How it works

  1. Text is split by blank lines and length into paragraphs, then rendered into SoM images.
  2. Resolution decays with age through the vivid → normal → fuzzy tiers. When a lookup hits a low-resolution memory, active recall restores it to high resolution first.
  3. Retrieval uses text overlap and optional OCR evidence by default. With an optical locator configured, the model emits K-bit 0/1 labels and the plugin selects segments using threshold and Top-K rules.
  4. Fetch reads the selected segments from the persisted source text and returns them verbatim, with no substitute text in between.
  5. Visual embeddings and hit-frequency decay are optional. Per-turn context injection is on by default, with index mode injecting L1 and optical metadata, and contextMode: snapshot keeping body snapshots.
  6. Governance tools share the same plugin instance. Maintenance is bounded by maintenanceBatchSize, can be cancelled, runs single-flight per namespace, and drains on disposal.

Showing the opening section of the README — the full document lives in the repository以上为 README 开头摘要,完整文档在仓库内 · View the full README on GitHub →在 GitHub 查看完整 README →

← 上一个 Prev ds-web-ui 下一个 Next dsh-lan-proxy →