BiBoyang/dsh-eval-harness
DSH插件评测工具:基于YAML用例驱动的真实代理回归评测,并支持与基线对比的PASS/WARN/FAIL门禁检查|DeepSeek Harness插件的回归评测框架
Project Overview项目介绍
dsh-eval-harness is a native plugin built exclusively for DeepSeek Harness (DSH) plugin and skill creators. It adds a full end-to-end regression testing workflow that can be integrated directly into continuous integration pipelines to catch breaking changes before they are merged. The workflow starts with developers writing test cases in simple YAML format, then the tool runs the actual DSH agent in headless mode, parses saved session traces, runs custom assertions, and outputs a structured report for CI gates to process.
Installation is done directly via the DSH plugin command, pulling the source code straight from the GitHub repository. The original npm distribution channel is deprecated and no longer receives updates, so users are strongly advised to avoid installing from npm. The plugin ships three core commands: eval_run to execute all test cases and generate reports, eval_gate to compare results against a baseline and output a pass/fail verdict, and eval_judge_validate to calibrate LLM-based judges. It also includes a helper skill called eval that teaches DSH how to write properly formatted test cases for new plugins.
Each test case runs in an isolated workspace, so concurrent execution of multiple test cases does not cause interference between runs. The tool automatically detects both compressed zstd session traces and plain uncompressed JSONL traces, so no extra configuration is needed to read existing DSH session logs. Test cases support a wide range of assertions, from checking tool call sequences and output content to running LLM-based semantic judgments for open-ended requirements. The entire workflow can be configured to run automatically on GitHub Actions, with scheduled full regression runs and manual baseline updates after review.
这是一款专为 DeepSeek Harness (DSH) 插件和技能开发者打造的原生 DSH 插件,可为 DSH 插件开发提供可集成进 CI 的回归评测门禁工具。用户使用 YAML 格式编写测试用例,工具会以后台无头模式驱动真实 DSH 代理运行,解析会话轨迹,执行断言,再与基准报告对比输出 PASS/FAIL 结果和退出码,供 CI 拦截不合格的回归。
安装方式是通过 DSH 官方的插件命令从 GitHub 源码拉取,原 npm 渠道已停止更新不再推荐使用。工具提供三个核心能力:eval_run 运行所有测试用例并生成报告,eval_gate 对比基准输出门禁判定,eval_judge_validate 用于校准 LLM 评审模型。它还附带一个 eval 技能,可帮助开发者学习编写规范的测试用例。
测试用例支持多种断言类型,包括工具调用序列检查、输出内容匹配、正则匹配、LLM 语义评审等,还支持标签筛选、重试、可靠性测试等配置。每个测试用例运行在独立隔离的工作区,支持并发运行互不干扰,工具可直接集成进 GitHub Actions 等常见 CI 流程,方便开发者做持续回归测试。
请帮我安装这个 DSH 插件。安装前先完成【兼容性检查 + 安全性检查】,检查通过再动手。
插件:dsh-eval-harness(BiBoyang/dsh-eval-harness)
仓库:https://github.com/BiBoyang/dsh-eval-harness
本站详情页:https://www.yhbd.top/plugins/biboyang-dsh-eval-harness/
本站登记:类型 plugin · 归类 原生 DSH 插件 · 许可证未声明 · ⭐ 13 · 最近提交 2026-09-26 · 主语言 TypeScript
按下面顺序执行,每步先把结论告诉我,再进入下一步:
【1 兼容性检查】
① 我这边:DSH 版本、Node 版本、操作系统、当前 profile(web / desktop)。
② 读它的 README、package.json、插件 manifest,列出它要求的 DSH 版本 / Node 版本 / 操作系统 / 外部依赖 / 需要另外先装的运行时。
③ 逐条比对,结论只写「满足 / 不满足 / 未知」三种;不满足的给出可行替代方案。
④ 检查是否和我已装的插件冲突:命令名重复、skill / tool 重名、端口占用、重复注册的 MCP server。
【2 安全性检查】
① 仓库可信度:和上面「本站登记」是否一致;star / fork 数、创建时间、最近提交,是否归档或长期停更。
② 安装脚本:逐行看 package.json 的 preinstall / install / postinstall,以及 install.sh、setup.ps1 之类脚本。出现 curl|bash、下载后直接执行、混淆代码、访问与插件功能无关的域名,立刻停下来告诉我,不要继续装。
③ 依赖:列出新增依赖,标出无人维护、或与知名包拼写近似的可疑包(typosquatting)。
④ 权限与副作用:它会读写哪些目录、访问哪些域名、需要哪些 DSH 权限(filesystem / network / shell / clipboard 等),以及怎么卸载和回滚。
⑤ 如果它要求 sudo / 管理员权限,或权限明显超出功能所需,先停下来问我。
【3 安装】
上面两步没有「不满足」和「高危项」时才执行;用官方推荐方式安装,不要自行提权。
【4 汇报】
用表格输出:检查项 / 结论 / 依据 / 是否需要我决策。拿不准的一律写「未知」并说明要我怎么确认——不要猜,也不要替我决定。
Send this message to DSH in your current session: it verifies compatibility and security first (answering met / not met / unknown item by item) and only installs once everything checks out — it will stop and ask you if it finds a high-risk item. The box scrolls; the copy is the full prompt. CLI install commands may not be accurate across systems, so DSH is the safer route.把上面这条消息直接发给当前会话里的 DSH:它会先核对兼容性与安全性(逐条给「满足 / 不满足 / 未知」),确认没问题再安装,有高危项会停下来问你。框内可滚动,复制到的是完整提示词;安装命令不一定准确,发给 DSH 更稳。
- No license declared - all rights reserved by default; ask the author before commercial use or redistribution未声明开源许可证 —— 默认「保留所有权利」,商用或再分发前先问作者
- 13 stars - an early-stage project星标 13,属于早期项目
DSH walks through these 9 checksDSH 会逐条核对这 9 项
Compatibility兼容性
- DSH, Node, OS and profile requirementsDSH 版本 / Node 版本 / 操作系统 / profile 是否满足要求
- External dependencies and runtimes (Electron / Python / Docker, ...)外部依赖与运行时(Electron / Python / Docker 等)是否齐备
- Conflicts with installed plugins: command names, skill / tool names, ports, duplicate MCP registration与已装插件是否冲突:命令名、skill / tool 重名、端口占用、重复 MCP 注册
Security安全性
- Repo matches the facts registered here; archived or abandoned?仓库是否与页面登记一致,是否归档或长期停更
- Safety of preinstall / install / postinstall and install.sh / setup.ps1preinstall / install / postinstall 与 install.sh、setup.ps1 是否安全
- curl|bash, download-then-execute, obfuscation, unrelated domains → stop immediatelycurl|bash、下载即执行、混淆代码、无关域名 → 立刻停止
- Typosquatting or unmaintained packages among the new dependencies新增依赖里有没有 typosquatting 或无人维护的包
- Requested permissions vs. what the feature actually needs申请了哪些权限、是否超出功能所需(filesystem / network / shell / clipboard)
- Any sudo / admin requirement, plus uninstall and rollback是否要求 sudo / 管理员权限,以及卸载与回滚方式
Anything uncertain must be marked unknown with a note on how to confirm it. This site's signal screen is a static snapshot, not a security audit.拿不准的必须标「未知」并说明要我怎么确认。本站的信号筛查是静态快照,不能替代安全审计。
Or use CLI install (for developers)或使用命令行安装(适合开发者)
CLI Install命令行安装
dsh plugin --profile headless add github:BiBoyang/dsh-eval-harness
把 BiBoyang/dsh-eval-harness 加入你的 DSH 配置(web profile)即可启用。
READMEREADME
dsh-eval-harness
DSH 插件/skill 作者的回归评测门禁:写 yaml 用例 → headless 驱动真实 agent 跑 → 解析 session trace 断言 → 对比 baseline 出 PASS/WARN/FAIL 报告与 CI 退出码。
简介
给 DSH 插件/skill 的回归评测流程提供一个可进 CI 的门禁工具:
- 用 yaml 写评测用例(prompt + 期望行为断言);
eval_run逐条 forkdsh --profile headless --patch <overlay> <prompt>子进程跑真实 agent(overlay 把会话落盘切到隔离目录,每条用例独立 workspace),解析落盘的session.jsonl/session.jsonl.zstdtrace(多帧 zstd 直读),执行断言,写report.json+report.md;eval_gate把本次报告与 baseline 报告对比,输出OVERALL=PASS|WARN|FAIL|N/A与退出码,供 CI 拦截回归。
安装
npm 渠道已废弃:registry 上的
dsh-eval-harness停留在 0.3.1,不再更新,请勿从 npm 安装。分发只走 GitHub。
从 GitHub 源码安装:
dsh plugin --profile headless add github:BiBoyang/dsh-eval-harness
# 验证挂载
dsh --profile headless --dump-config | grep dsh-eval-harness
能力面
Tools
| 工具 | 说明 |
|---|---|
eval_run |
跑 cases_dir 下全部用例:headless 驱动真实 agent → 采集 session trace → 断言 → 写 report.json/report.md |
eval_gate |
对比 baseline 与本次报告,输出门禁判定(OVERALL/EXIT_CODE),strict 模式收紧 WARN 退出码 |
eval_judge_validate |
在人工标注集上校准 LLM judge:报混淆矩阵与 TPR/TNR(分开看,agreement 会骗人),双指标达标才算 calibrated |
Skills
| Skill | 作用 |
|---|---|
eval |
教模型帮用户编写评测用例(用例格式、断言编写要点、解析子集约束) |
用例格式(cases/*.yml)
一个文件一条用例:
name: 用例名 # 唯一,gate 按 name 对比 baseline
prompt: "发给 agent 的内容" # 多行可用块标量 `|`
require_plugins: [some-plugin] # 可选,元信息
tags: [fast] # 可选,标签;eval_run 的 tags 筛选按任一命中匹配
retries: 1 # 可选,失败重跑次数(非负整数,缺省用 eval_run 的全局 retries)
trials: 3 # 可选,可靠性测量的独立 trial 次数(正整数,缺省用 eval_run 的全局 trials,默认 1);
# trials > 1 时忽略 retries——测量必须是没有重试干预的原始单次成功率
mock: # 可选,mock 模式(见下节;不用真实 API,离线确定性)
fault: F4 # 注入故障形态 F0-F5(缺省 F0)
api: openai-completions # 协议端点:openai-completions / openai-responses / anthropic-messages(缺省 openai-completions)
once: true # 可选,一次性故障:仅首个 LLM 请求命中,此后回 F0(横评矩阵语义)
plugins: [dsh-find-plugin@0.4.0] # 可选,挂载被测插件(见下节;版本必须钉死)
assert:
turn_end: completed # turn/end 事件的 reason.kind
exit_code: 0 # 可选,dsh 子进程退出码;声明后非零退出进断言层比对(不再直接记 error)
tools_called: [tool_a] # tool/call 名称序列须按序包含(保序子序列)
output_contains: ["关键词"] # 最终 assistant 文本须包含全部
max_steps: 8 # 可选,step/end 数上限
max_tokens: 50000 # 可选,token 上限(input+output+reasoning;cacheRead/cacheWrite 不计入,防多步膨胀)
no_tool_errors: true # 可选,任何 tool/result 硬错误(data.error / isError)即 fail
tools_exact: [tool_a] # 可选,工具调用名称序列须完全一致(长度+顺序+内容)
tools_not_called: [tool_b] # 可选,列出的工具一次都不能被调用
output_not_contains: ["抱歉"] # 可选,最终 assistant 文本不得包含任一子串
output_matches: ["^okay"] # 可选,最终 assistant 文本须匹配全部正则(解析期预编译校验)
tool_args_contains: # 可选,指定工具至少一次调用的参数 JSON 串包含子串
- name: tool_a
contains: '"path"'
tool_result_contains: # 可选,指定工具至少一次结果的文本包含子串
- name: tool_a
contains: total
output_judge: # 可选,LLM 语义评审(结构断言全过后才调,判 FAIL 记 fail)
rubric: "回答应解释原因而非只给结论"
Showing the opening section of the README — the full document lives in the repository以上为 README 开头摘要,完整文档在仓库内 · View the full README on GitHub →在 GitHub 查看完整 README →
chuspeeism/dashi-taskboard
zhoushoujianwork/easyeda-agent
morluto/rea
gitroomhq/postiz-agent
linhay/harmony-next.skills
yzlnew/infra-skills
liceses/dsh-gitbash-preset
sjh9714/dsh-win32