LeemanCheung/dsh-agent-arena
隔离的多模型编码匹配,附带确定性验证、评分与报告。
Project Overview项目介绍
This is a native plugin built exclusively for DeepSeek Harness (DSH) that enables head-to-head comparison of up to four different configured AI coding models. It creates isolated Git worktrees in the system temporary directory to keep each model’s changes separate, ensuring deterministic, reproducible comparison results. To install the plugin, you run the DSH CLI command dsh plugin --profile web add github:LeemanCheung/dsh-agent-arena, then restart your running DSH Web process and refresh the browser page to activate it. Your selected DSH profile must have support for agents, subprocess handling, storage domains, and at least two active model routes for the plugin to work.
The plugin is designed for developers who want to compare the coding output of different large language models on the same task. Before starting a match, you need to ensure your current Git working tree is clean, then add your project’s custom validation commands (such as test, typecheck, or build) in the plugin’s settings panel. Each model runs the user’s requested task in its own independent session and worktree, then the plugin calculates a score based on how well the model’s output passes the configured validations. After scoring completes, it generates a diff of all changes for you to review before selecting a winner to apply.
This plugin is released under the permissive MIT open source license, so it is free to use and modify. There are several known limitations to be aware of: validation commands cannot accept arguments with whitespace, task cancellation depends on the model provider honoring the abort signal, and patches larger than 1MB are rejected before any changes are applied. You must have the required DSH dependencies enabled on your profile to run this plugin, and it is currently tested and marked compatible with DSH 0.1.2-rc.1 web profile. It has passed automated testing with 15 passing tests on both Windows and Linux.
这是一个专为DeepSeek Harness(DSH)开发的原生插件,核心功能是搭建一个编码对比竞技场,支持同时比较2到4个不同配置模型对同一任务的输出结果。插件会在系统临时目录创建独立的Git工作树,隔离每个模型的执行过程,保证对比结果的确定性。安装方式为通过DSH命令行执行 dsh plugin --profile web add github:LeemanCheung/dsh-agent-arena,之后重启DSH Web进程并刷新页面即可使用。
插件适用于需要对比不同大模型编码效果的开发者,典型工作流程为:先保证当前仓库Git状态干净,再在DSH设置中配置验证命令(如项目测试、类型检查、构建命令),设置竞技场目标后启动对比。每个模型会在独立会话中处理任务,完成后插件会根据验证结果计算得分,最后生成diff供用户手动审查并选择胜者应用。
本插件采用MIT许可证开源,完全免费使用,目前已知存在一些限制:验证命令参数不支持包含空格,终止任务依赖模型提供商响应中止信号,超过1MB的补丁会被直接拒绝。使用前需要确保当前DSH配置文件提供agents、subprocess、存储域等必要能力,且至少配置两个可用的模型路由。
请帮我安装这个 DSH 插件。安装前先完成【兼容性检查 + 安全性检查】,检查通过再动手。
插件:dsh-agent-arena(LeemanCheung/dsh-agent-arena)
仓库:https://github.com/LeemanCheung/dsh-agent-arena
本站详情页:https://www.yhbd.top/plugins/leemancheung-dsh-agent-arena/
本站登记:类型 plugin · 归类 原生 DSH 插件 · 许可证 MIT · ⭐ 2 · 最近提交 2026-09-05 · 主语言 TypeScript
按下面顺序执行,每步先把结论告诉我,再进入下一步:
【1 兼容性检查】
① 我这边:DSH 版本、Node 版本、操作系统、当前 profile(web / desktop)。
② 读它的 README、package.json、插件 manifest,列出它要求的 DSH 版本 / Node 版本 / 操作系统 / 外部依赖 / 需要另外先装的运行时。
③ 逐条比对,结论只写「满足 / 不满足 / 未知」三种;不满足的给出可行替代方案。
④ 检查是否和我已装的插件冲突:命令名重复、skill / tool 重名、端口占用、重复注册的 MCP server。
【2 安全性检查】
① 仓库可信度:和上面「本站登记」是否一致;star / fork 数、创建时间、最近提交,是否归档或长期停更。
② 安装脚本:逐行看 package.json 的 preinstall / install / postinstall,以及 install.sh、setup.ps1 之类脚本。出现 curl|bash、下载后直接执行、混淆代码、访问与插件功能无关的域名,立刻停下来告诉我,不要继续装。
③ 依赖:列出新增依赖,标出无人维护、或与知名包拼写近似的可疑包(typosquatting)。
④ 权限与副作用:它会读写哪些目录、访问哪些域名、需要哪些 DSH 权限(filesystem / network / shell / clipboard 等),以及怎么卸载和回滚。
⑤ 如果它要求 sudo / 管理员权限,或权限明显超出功能所需,先停下来问我。
【3 安装】
上面两步没有「不满足」和「高危项」时才执行;用官方推荐方式安装,不要自行提权。
【4 汇报】
用表格输出:检查项 / 结论 / 依据 / 是否需要我决策。拿不准的一律写「未知」并说明要我怎么确认——不要猜,也不要替我决定。
Send this message to DSH in your current session: it verifies compatibility and security first (answering met / not met / unknown item by item) and only installs once everything checks out — it will stop and ask you if it finds a high-risk item. The box scrolls; the copy is the full prompt. CLI install commands may not be accurate across systems, so DSH is the safer route.把上面这条消息直接发给当前会话里的 DSH:它会先核对兼容性与安全性(逐条给「满足 / 不满足 / 未知」),确认没问题再安装,有高危项会停下来问你。框内可滚动,复制到的是完整提示词;安装命令不一定准确,发给 DSH 更稳。
- Only 2 stars - very few users, little community feedback星标只有 2,几乎没人在用,遇到问题缺少社区反馈
DSH walks through these 9 checksDSH 会逐条核对这 9 项
Compatibility兼容性
- DSH, Node, OS and profile requirementsDSH 版本 / Node 版本 / 操作系统 / profile 是否满足要求
- External dependencies and runtimes (Electron / Python / Docker, ...)外部依赖与运行时(Electron / Python / Docker 等)是否齐备
- Conflicts with installed plugins: command names, skill / tool names, ports, duplicate MCP registration与已装插件是否冲突:命令名、skill / tool 重名、端口占用、重复 MCP 注册
Security安全性
- Repo matches the facts registered here; archived or abandoned?仓库是否与页面登记一致,是否归档或长期停更
- Safety of preinstall / install / postinstall and install.sh / setup.ps1preinstall / install / postinstall 与 install.sh、setup.ps1 是否安全
- curl|bash, download-then-execute, obfuscation, unrelated domains → stop immediatelycurl|bash、下载即执行、混淆代码、无关域名 → 立刻停止
- Typosquatting or unmaintained packages among the new dependencies新增依赖里有没有 typosquatting 或无人维护的包
- Requested permissions vs. what the feature actually needs申请了哪些权限、是否超出功能所需(filesystem / network / shell / clipboard)
- Any sudo / admin requirement, plus uninstall and rollback是否要求 sudo / 管理员权限,以及卸载与回滚方式
Anything uncertain must be marked unknown with a note on how to confirm it. This site's signal screen is a static snapshot, not a security audit.拿不准的必须标「未知」并说明要我怎么确认。本站的信号筛查是静态快照,不能替代安全审计。
Or use CLI install (for developers)或使用命令行安装(适合开发者)
CLI Install命令行安装
dsh plugin --profile web add github:LeemanCheung/dsh-agent-arena
把 LeemanCheung/dsh-agent-arena 加入你的 DSH 配置(web profile)即可启用。
READMEREADME
dsh-agent-arena
English | 中文
A DSH coding arena for comparing 2–4 configured models in isolated Git worktrees, validating their changes deterministically, reviewing every diff, and explicitly applying one winner.
Screenshot

Generated with GPT Image from the implemented Client layout and feature set; runtime appearance follows the active DSH theme and viewport.
Execution and persistence
- Creates detached worktrees under the operating system temporary directory, outside the compared repository.
- Restricts recursive cleanup to the exact two-level match/contestant path under that dedicated Arena directory, including resolved junction and symlink checks.
- Starts each contestant with the public
ctx.agents.create,SessionId,createUserMessage, andAgent.followupAPIs. - Executes Git and validation argv through
ctx.subprocess.spawn; a shell is never used and shell operators are rejected. - Persists match reports through
storageDomain. An interrupted nonterminal match is marked failed after Host restart. - Polls the generated
agentArenaRemote namespace in Settings and supports Start, Cancel, diff review, and Apply winner without browser globals.
Scoring and application
Each validation has a positive weight. A contestant score is the percentage of total validation weight that exits successfully; score ties use contestant id as a stable deterministic tie-breaker. No LLM judge is used. The form starts without a generic validation because git status --porcelain=v1 normally exits successfully regardless of whether it prints changes; add the repository's own test, typecheck, or build command before starting.
A match requires clean git status --porcelain=v1 before worktrees are created. The winner worktree remains available until explicit application. Apply repeats the cleanliness check, verifies that HEAD still equals the recorded base revision, runs git apply --check, then git apply --index --whitespace=error. Arena never auto-applies, commits, pushes, or rewrites history.
Showing the opening section of the README — the full document lives in the repository以上为 README 开头摘要,完整文档在仓库内 · View the full README on GitHub →在 GitHub 查看完整 README →
tt-a1i/archify
loopx-project/loopx
ZSeven-W/openpencil
omdsh-dev/DSH-better-sidebar
NanmiCoder/dsh-agent-teams
LiPu-jpg/Openwrite