BiBoyang/dsh-eval-harness 预览 preview

BiBoyang/dsh-eval-harness

DSH插件评测工具:基于YAML用例驱动的真实代理回归评测,并支持与基线对比的PASS/WARN/FAIL门禁检查|DeepSeek Harness插件的回归评测框架

Project Overview项目介绍

dsh-eval-harness is a native plugin built exclusively for DeepSeek Harness (DSH) plugin and skill creators. It adds a full end-to-end regression testing workflow that can be integrated directly into continuous integration pipelines to catch breaking changes before they are merged. The workflow starts with developers writing test cases in simple YAML format, then the tool runs the actual DSH agent in headless mode, parses saved session traces, runs custom assertions, and outputs a structured report for CI gates to process.

Installation is done directly via the DSH plugin command, pulling the source code straight from the GitHub repository. The original npm distribution channel is deprecated and no longer receives updates, so users are strongly advised to avoid installing from npm. The plugin ships three core commands: eval_run to execute all test cases and generate reports, eval_gate to compare results against a baseline and output a pass/fail verdict, and eval_judge_validate to calibrate LLM-based judges. It also includes a helper skill called eval that teaches DSH how to write properly formatted test cases for new plugins.

Each test case runs in an isolated workspace, so concurrent execution of multiple test cases does not cause interference between runs. The tool automatically detects both compressed zstd session traces and plain uncompressed JSONL traces, so no extra configuration is needed to read existing DSH session logs. Test cases support a wide range of assertions, from checking tool call sequences and output content to running LLM-based semantic judgments for open-ended requirements. The entire workflow can be configured to run automatically on GitHub Actions, with scheduled full regression runs and manual baseline updates after review.

这是一款专为 DeepSeek Harness (DSH) 插件和技能开发者打造的原生 DSH 插件,可为 DSH 插件开发提供可集成进 CI 的回归评测门禁工具。用户使用 YAML 格式编写测试用例,工具会以后台无头模式驱动真实 DSH 代理运行,解析会话轨迹,执行断言,再与基准报告对比输出 PASS/FAIL 结果和退出码,供 CI 拦截不合格的回归。

安装方式是通过 DSH 官方的插件命令从 GitHub 源码拉取,原 npm 渠道已停止更新不再推荐使用。工具提供三个核心能力:eval_run 运行所有测试用例并生成报告,eval_gate 对比基准输出门禁判定,eval_judge_validate 用于校准 LLM 评审模型。它还附带一个 eval 技能,可帮助开发者学习编写规范的测试用例。

测试用例支持多种断言类型,包括工具调用序列检查、输出内容匹配、正则匹配、LLM 语义评审等,还支持标签筛选、重试、可靠性测试等配置。每个测试用例运行在独立隔离的工作区,支持并发运行互不干扰,工具可直接集成进 GitHub Actions 等常见 CI 流程,方便开发者做持续回归测试。

Pre-install check安装前体检Compatibility · Security兼容性 · 安全性 1 warning1 项注意
  • No license declared - all rights reserved by default; ask the author before commercial use or redistribution未声明开源许可证 —— 默认「保留所有权利」,商用或再分发前先问作者
  • 13 stars - an early-stage project星标 13,属于早期项目
DSH walks through these 9 checksDSH 会逐条核对这 9 项

Compatibility兼容性

  • DSH, Node, OS and profile requirementsDSH 版本 / Node 版本 / 操作系统 / profile 是否满足要求
  • External dependencies and runtimes (Electron / Python / Docker, ...)外部依赖与运行时(Electron / Python / Docker 等)是否齐备
  • Conflicts with installed plugins: command names, skill / tool names, ports, duplicate MCP registration与已装插件是否冲突:命令名、skill / tool 重名、端口占用、重复 MCP 注册

Security安全性

  • Repo matches the facts registered here; archived or abandoned?仓库是否与页面登记一致,是否归档或长期停更
  • Safety of preinstall / install / postinstall and install.sh / setup.ps1preinstall / install / postinstall 与 install.sh、setup.ps1 是否安全
  • curl|bash, download-then-execute, obfuscation, unrelated domains → stop immediatelycurl|bash、下载即执行、混淆代码、无关域名 → 立刻停止
  • Typosquatting or unmaintained packages among the new dependencies新增依赖里有没有 typosquatting 或无人维护的包
  • Requested permissions vs. what the feature actually needs申请了哪些权限、是否超出功能所需(filesystem / network / shell / clipboard)
  • Any sudo / admin requirement, plus uninstall and rollback是否要求 sudo / 管理员权限,以及卸载与回滚方式

Anything uncertain must be marked unknown with a note on how to confirm it. This site's signal screen is a static snapshot, not a security audit.拿不准的必须标「未知」并说明要我怎么确认。本站的信号筛查是静态快照,不能替代安全审计。

Or use CLI install (for developers)或使用命令行安装(适合开发者)

CLI Install命令行安装

dsh plugin --profile headless add github:BiBoyang/dsh-eval-harness

把 BiBoyang/dsh-eval-harness 加入你的 DSH 配置(web profile)即可启用。

READMEREADME

dsh-eval-harness

DSH 插件/skill 作者的回归评测门禁:写 yaml 用例 → headless 驱动真实 agent 跑 → 解析 session trace 断言 → 对比 baseline 出 PASS/WARN/FAIL 报告与 CI 退出码。

简介

给 DSH 插件/skill 的回归评测流程提供一个可进 CI 的门禁工具:

  1. 用 yaml 写评测用例(prompt + 期望行为断言);
  2. eval_run 逐条 fork dsh --profile headless --patch <overlay> <prompt> 子进程跑真实 agent(overlay 把会话落盘切到隔离目录,每条用例独立 workspace),解析落盘的 session.jsonl / session.jsonl.zstd trace(多帧 zstd 直读),执行断言,写 report.json + report.md;
  3. eval_gate 把本次报告与 baseline 报告对比,输出 OVERALL=PASS|WARN|FAIL|N/A 与退出码,供 CI 拦截回归。

安装

npm 渠道已废弃:registry 上的 dsh-eval-harness 停留在 0.3.1,不再更新,请勿从 npm 安装。分发只走 GitHub。

从 GitHub 源码安装:

dsh plugin --profile headless add github:BiBoyang/dsh-eval-harness

# 验证挂载
dsh --profile headless --dump-config | grep dsh-eval-harness

能力面

Tools

工具 说明
eval_run 跑 cases_dir 下全部用例:headless 驱动真实 agent → 采集 session trace → 断言 → 写 report.json/report.md
eval_gate 对比 baseline 与本次报告,输出门禁判定(OVERALL/EXIT_CODE),strict 模式收紧 WARN 退出码
eval_judge_validate 在人工标注集上校准 LLM judge:报混淆矩阵与 TPR/TNR(分开看,agreement 会骗人),双指标达标才算 calibrated

Skills

Skill 作用
eval 教模型帮用户编写评测用例(用例格式、断言编写要点、解析子集约束)

用例格式(cases/*.yml)

一个文件一条用例:

name: 用例名                    # 唯一,gate 按 name 对比 baseline
prompt: "发给 agent 的内容"      # 多行可用块标量 `|`
require_plugins: [some-plugin]  # 可选,元信息
tags: [fast]                    # 可选,标签;eval_run 的 tags 筛选按任一命中匹配
retries: 1                      # 可选,失败重跑次数(非负整数,缺省用 eval_run 的全局 retries)
trials: 3                       # 可选,可靠性测量的独立 trial 次数(正整数,缺省用 eval_run 的全局 trials,默认 1);
                                # trials > 1 时忽略 retries——测量必须是没有重试干预的原始单次成功率
mock:                           # 可选,mock 模式(见下节;不用真实 API,离线确定性)
  fault: F4                     #   注入故障形态 F0-F5(缺省 F0)
  api: openai-completions       #   协议端点:openai-completions / openai-responses / anthropic-messages(缺省 openai-completions)
  once: true                    #   可选,一次性故障:仅首个 LLM 请求命中,此后回 F0(横评矩阵语义)
  plugins: [dsh-find-plugin@0.4.0]  # 可选,挂载被测插件(见下节;版本必须钉死)
assert:
  turn_end: completed           # turn/end 事件的 reason.kind
  exit_code: 0                  # 可选,dsh 子进程退出码;声明后非零退出进断言层比对(不再直接记 error)
  tools_called: [tool_a]        # tool/call 名称序列须按序包含(保序子序列)
  output_contains: ["关键词"]    # 最终 assistant 文本须包含全部
  max_steps: 8                  # 可选,step/end 数上限
  max_tokens: 50000             # 可选,token 上限(input+output+reasoning;cacheRead/cacheWrite 不计入,防多步膨胀)
  no_tool_errors: true          # 可选,任何 tool/result 硬错误(data.error / isError)即 fail
  tools_exact: [tool_a]         # 可选,工具调用名称序列须完全一致(长度+顺序+内容)
  tools_not_called: [tool_b]    # 可选,列出的工具一次都不能被调用
  output_not_contains: ["抱歉"]  # 可选,最终 assistant 文本不得包含任一子串
  output_matches: ["^okay"]     # 可选,最终 assistant 文本须匹配全部正则(解析期预编译校验)
  tool_args_contains:           # 可选,指定工具至少一次调用的参数 JSON 串包含子串
    - name: tool_a
      contains: '"path"'
  tool_result_contains:         # 可选,指定工具至少一次结果的文本包含子串
    - name: tool_a
      contains: total
  output_judge:                 # 可选,LLM 语义评审(结构断言全过后才调,判 FAIL 记 fail)
    rubric: "回答应解释原因而非只给结论"

Showing the opening section of the README — the full document lives in the repository以上为 README 开头摘要,完整文档在仓库内 · View the full README on GitHub →在 GitHub 查看完整 README →

← 上一个 Prev dsh-plugin-ya-workspace-sidebar 下一个 Next Fairy-DSH-Optimized →