hi-fangj/dsh-models-radar
Model capability radar plugin for the DeepSeek Harness Web GUI
Project Overview项目介绍
dsh-models-radar is a DeepSeek Harness plugin that pulls public benchmarks from deng.codexradar.com, adds a Model Radar page under Settings, and surfaces the current session model's DeepSWE capability in the composer toolbar. It delivers a 0–110 capability overview, 24h/7d trend charts, a cost-by-IQ scatter, and community ratings, with harness attribution badges per row. Use it to compare base models and reasoning tiers on one screen before picking a session model. Note: data refreshes on per-dataset windows, not in real time, and unranked base models are hidden by design.
dsh-models-radar 是 DeepSeek Harness 插件,从 deng.codexradar.com 拉取公开基准,在设置页新增“Model Radar”,并于撰写器工具栏显示当前模型的 DeepSWE 能力值。提供 0–110 分的能力总览、24h/7d 趋势、成本×能力散点与社区评分。适用场景:在同一屏比较多个基础模型与推理档位的实际表现,依据能力与成本挑选会话模型。注意:数据为非实时,依赖上游公开榜单刷新;能力读数仅显示榜单收录模型,未收录基础模型按设计隐藏。
请帮我了解并安装插件:【dsh-models-radar】【https://github.com/hi-fangj/dsh-models-radar】
Send this message to DSH in your current session. CLI install commands may not be accurate across systems — DSH will figure it out for you.把上面这条消息直接发给当前会话里的 DSH,让它帮你了解并安装。安装命令不一定准确,发给 DSH 更稳。
Or use CLI install (for developers)或使用命令行安装(适合开发者)
CLI Install命令行安装
dsh plugin --profile web add github:hi-fangj/dsh-models-radar
把 hi-fangj/dsh-models-radar 加入你的 DSH 配置(web profile)即可启用。
READMEREADME
dsh-models-radar
dsh-models-radar · Model Capability Radar
Crowd-benchmarked capability scores from deng.codexradar.com, inside the DeepSeek Harness: overview, trend, and cost on one screen.
中文文档 · Usage · Data & Privacy · Troubleshooting
A model capability radar plugin for the DeepSeek Harness Web GUI. It reads public benchmark data from deng.codexradar.com, adds a Model Radar page to Settings, and shows the selected session model's live DeepSWE score in the composer tool row, left of the model selector.
Screenshots
Settings · capability overview — best-effort-per-base ranking with per-row Harness attribution (Codex / Claude Code / DSH / ZCode / Grok / Kimi Code / Antigravity / CodeBuddy); click any row to switch the charts below to that tier.

Capability popover — opened from the composer readout: cross-base comparison plus the current tier's details, with a live "current" mark following the session model.

Highlights
- Score attribution at a glance. Every base-model row carries a Harness badge (Codex / Claude Code / DSH / ZCode / Grok / Kimi Code / Antigravity / CodeBuddy, site palette), and the tier selector options read
model · effort · harness; unmatchable bases get no badge — never a guess. - Best-effort-per-base ranking. The capability overview groups by base model with a fixed
0–110absolute-scale magnitude bar and a 24h trend signal per row; expand a row for the base's full reasoning-effort ladder. - 24h / 7d dual-window IQ trend. Tab between two time windows, each independently scaled with its own full stats (net change, low, average, high); the curve is colored by capability band.
- Cost × IQ from three angles. Tabs for composite cost (the site's own 2.5×-price-for-1.35×-speed trade-off, normalized per chart), time cost, and price cost; color = base, shape = reasoning effort, same-base tiers joined by ladder lines. Upper-left = more efficient. Hovering surfaces the site's three-line reading: attribution (display name · billing · harness · effort), IQ with its pass/total, and the active metric with sample counts.
- Community ratings. The codexradar.com main-site community's 0–10 experience scores over rolling 7-day / 24-hour windows as a bar chart: grouped by base model with efforts ordered within each group, color = base, the selected tier highlighted (≈ marks an approximate match), and unrated slots kept as explicit placeholders. Each window carries its own freshness window; the data is global and independent of the benchmark channels.
- Live readout beside the model selector. Exact
model@reasoningEffortmatching through DSH's official per-session model directory, updating immediately on model switches; click it to open the capability popover for cross-base comparison. - Lightweight, credential-free, offline-tolerant. The browser never hits upstream directly (same-origin host proxy), freshness windows mean zero upstream requests inside a window, the latest local snapshot serves as fallback, and no credentials are requested or submitted.
Features
- Settings → Model Radar page through the additive
settings.sectionslot - A Model Radar card under Settings → Plugins → Configurable plugins through
settings.plugin.item(the live-readout display switch, persisted in Host settings) - Capability overview grouped by base model with expandable reasoning-effort tiers and per-row Harness attribution badges
- Entry default = the current conversation's model: every visit resolves the session's selected model through the tier-match rule (the deployment default only stands in when no session selection is resolvable), the overview marks the in-use row, and a manual pick lasts for the visit only
- Fixed
0–110IQ scale with consistent capability-band semantics across channels - Trend tabs: last 24 hours / last 7 days, each independently y-scaled with full stats; the choice persists
- Cost × IQ card: composite / time / price tabs on a log x-axis, model filter chips synced across tabs, codex-run DSV4 bases hidden by default (site parity)
- Community ratings card: 7-day / 24-hour tabs (choice remembered), identical slot layout across windows, an independent 15-minute freshness window each; settings page only, never the popover
- Two benchmark channels:
deep-swe: code-repair tasks, binary-majority scoringpompeii-adjacency: visual reconstruction tasks, continuous Adjacency F1
- Semantic task diagnostics:
- DeepSWE: passed / split vote / failed
- Pompeii: low / general / good / excellent F1 bands
- attention-first sorting and local filters
- task titles link to their source repo (GitHub for DeepSWE), with the site's language badge after the title (Py/JS/TS/Go/Rust)
- Efficiency metric badges: IQ, average cost, average duration, cache hit rate, 24-hour run count
- Capability popover: opened from the composer readout; cross-base comparison plus the viewed tier's full details (badges, dual-window trend, task composition)
- Refresh within freshness windows: overview & per-task composition 15 min, channel list / IQ trend / task source info 60 min; only expired datasets are refetched, everything else is served from cache (single-flight, zero upstream hits inside a window)
- Manual refresh button in the footer (skips the windows) next to the last-fetch timestamp
- Offline fallback to the latest persisted snapshot when the upstream API is unavailable
- Chinese and English UI copy
Requirements
- DeepSeek Harness Web GUI
dsh-super-injectorfor runtime or persistent local installation- Node.js 22 or newer
- npm
Install
From Git
The simplest installation fetches the repository directly into the web profile:
dsh plugin --profile web add github:hi-fangj/dsh-models-radar
The equivalent full Git URL:
dsh plugin --profile web add git+https://github.com/hi-fangj/dsh-models-radar.git
The repository includes the built Host and browser bundles required by DSH, so direct Git installation does not run dependency lifecycle scripts and does not require a pnpm build allowlist. Refresh http://127.0.0.1:3080 after installation; restart DSH once if the running process does not hot-load the new package.
To remove the Git-installed package:
dsh plugin --profile web remove dsh-models-radar
From a local clone
To develop or modify the plugin, clone and build manually:
git clone https://github.com/hi-fangj/dsh-models-radar.git
cd dsh-models-radar
npm ci
npm run build
Build artifacts:
lib/index.js: Host ESM bundlelib/client.js: browser CJS bundle with the DSHModuleLoaderhandshake
After building, add the local package to the web profile:
dsh plugin --profile web add /absolute/path/to/dsh-models-radar
dsh plugin delegates dependency installation to the named profile, so the package lands in ~/.dsh/profiles/web. The repository ships build artifacts for Git-based installs; after changing local sources, run npm run build before adding or reloading the local clone.
Refresh http://127.0.0.1:3080 after installation. If the running DSH process does not hot-load the new package, restart DSH once and refresh again.
To remove a CLI-installed dependency:
dsh plugin --profile web remove dsh-models-radar
Restart DSH afterwards so the profile reassembles without the plugin.
Runtime injection and persistent install
Runtime injection suits quick trials. Ask the Agent in a DSH session to call:
dev_inject_plugin({
"dir": "/absolute/path/to/dsh-models-radar"
})
Then refresh http://127.0.0.1:3080 once. Injection lasts until the DSH process restarts or the plugin is unloaded explicitly.
To write the dependency and bundle list into the web profile, ask the Agent to call:
dev_install_package({
"dir": "/absolute/path/to/dsh-models-radar",
"profile": "web"
})
Refresh the Web GUI afterwards. DSH re-assembles the plugin from the profile on restart.
Usage
The Model Radar page
- Open Settings from the bottom-left of the DSH Web GUI.
- Choose Model Radar.
- Pick the
DeepSWEorPompeiichannel at the top. - Select a model tier from the capability overview or the tier selector.
Every page activation refreshes within the freshness windows: cached data costs no upstream requests and only expired datasets are refetched. Clicking any overview row switches the tier used by the efficiency badges, trend, and task diagnostics below.
Capability overview
Each base model shows its currently strongest tier by default; expanding a row reveals the full reasoning-effort ladder. The IQ progress bar uses a fixed 0–110 absolute scale:
| IQ | Capability band |
|---|---|
< 70 |
Developing |
70–84.9 |
General |
85–94.9 |
Steady |
95–99.9 |
Excellent |
≥ 100 |
Leader |
IQ trend
The last-24h and last-7d tabs are time-sliced views of the same hourly series, each with its own y-axis scaling and full stats (net change, low, average, high). The curve is colored by capability band with a matching translucent area fill; endpoint and hover markers use the band color.
Cost × IQ comparison
Three tabs (composite / time / price) plot every tier on a log cost axis × linear IQ axis: color = base model (site palette), shape = reasoning effort (off=× · low=○ · medium=△ · high=□ · xhigh=◇ · max=⬡ · ultra=★), same-base tiers joined by ladder lines in effort order. Upper-left = more efficient. The model chip row multi-selects, synced across the active tab; the codex-run DSV4 bases are hidden by default, matching the site.
Task diagnostics
DeepSWE uses the upstream's real majority-vote verdicts; Pompeii keeps continuous F1 semantics. Filtering, counting, and sorting all happen locally in the browser — switching filters adds no API requests.
Task titles and language badges come from the site's task catalog (the /table endpoint, 60-minute window, proxied through the Host which keeps only the tasks array): a title with a source repo is an outbound link (GitHub for DeepSWE, dataset page for Pompeii), and the badge follows the site's vocabulary — Py/JS/TS/Go/Rust, unknown languages shown verbatim. A failed catalog fetch only drops the badges and links; the channel view is unaffected.
Capability popover
Click the readout in the composer tool row (left of the model selector) to open the popover: a full base overview (for comparison) on top and the viewed tier's details below (efficiency badges, dual-window trend, task composition). The viewed tier follows the session model by default; clicking an overview row or using the trend card's tier selector views another tier temporarily until the session model changes or the popover closes.
The composer capability capsule
The compact capsule reads the model selected for the session's next request, not a guess from the last completed reply. Display:
SWE IQ 90.2
Match order:
- Exact
model@reasoningEffortmatch - The same base model's highest-IQ tier, prefixed with
≈ - Hidden entirely when the base is absent from the DeepSWE leaderboard
Switching models in the composer updates the capsule immediately. The readout polls the host every 15 minutes (the shortest freshness window) — one local request per tick and at most one upstream fetch per channel per window; on failure the last successful value is kept.
Prefer no capsule by the composer? The「Show live capability readout」switch on the Model Radar card under Settings → Plugins → Configurable plugins hides it entirely; while hidden the capsule renders nothing and stops background polling. The preference persists in Host settings (shared across browsers, immune to clearing browser storage; a legacy localStorage choice is migrated once on upgrade).
Update
cd /absolute/path/to/dsh-models-radar
git pull
npm ci
npm run build
For runtime-injected plugins, ask the Agent to hot-reload:
dev_reload_package({
"packageName": "dsh-models-radar"
})
Refresh the page if the client dependency graph changed.
Uninstall
Unload a runtime-injected plugin:
dev_uninject_plugin({
"match": "dsh-models-radar"
})
For persistently installed plugins, remove dsh-models-radar through the profile / plugin manager, then restart DSH. Snapshot history stays in ~/.dsh/plugin-data/dsh-models-radar/; delete that directory separately only if you no longer want the history.
Data and privacy
The plugin reads public, unauthenticated endpoints of https://api.codexradar.com/api/v1:
/benchmarks/intelligence-efficiency/iq-history/leaderboard
The browser never calls the upstream API directly. Because api.codexradar.com allowlists browser origins, the host half proxies requests through the same-origin route /model-radar/api/data. See ADR-0001 for the rationale.
This plugin:
- never requests passwords, tokens, or other credentials
- never submits benchmark results
- never sends session content
- stores only public benchmark snapshots locally
Snapshot directory:
~/.dsh/plugin-data/dsh-models-radar/
├── latest-deep-swe.json
├── latest-pompeii-adjacency.json
└── iq-timeline.jsonl
Architecture
Browser
├── settings.section → Model Radar page
├── settings.plugin.item → plugin-configuration card (live-readout display switch)
├── conversation.composer.dock → session capability capsule + popover
├── GET /model-radar/api/data
└── GET/POST /model-radar/api/pref
│
▼
Host plugin
├── per-dataset freshness windows (efficiency/tasks 15 min, channels/trend 60 min)
├── single-flight upstream requests + channel-global benchmarks cache
├── normalization into RadarView
├── settings namespace dsh-models-radar (live-readout preference, persisted in Host settings)
└── local snapshot persistence (served within its window across restarts)
The capability capsule subscribes to DSH's official modelDirectories per-session store, so model switches propagate without polling.
Development
npm ci
npm run build
GitHub Actions runs a build check on every push/PR; pushing a v* tag automatically builds, packs, and publishes a GitHub Release with the tgz attached. The release flow:
# after bumping the version in package.json and committing:
git tag v0.1.x
git push origin main --tags
Common DSH development operations:
dev_inject_plugin({ "dir": "/absolute/path/to/dsh-models-radar" })
dev_reload_package({ "packageName": "dsh-models-radar" })
dev_uninject_plugin({ "match": "dsh-models-radar" })
Main sources:
src/index.ts: host proxy, refresh throttling, snapshotssrc/client/RadarSection.tsx: settings-page state and compositionsrc/client/Overview.tsx: capability overviewsrc/client/charts.tsx: trend charts and task diagnosticssrc/client/costScatter.tsx: cost × IQ comparison scattersrc/client/harness.ts: harness attribution and tier-selector labelssrc/client/LiveCapability.tsx: composer readout and popoversrc/client/ScrollFrame.tsx: scrollbar for overflowing listssrc/client/scoreMetrics.ts: IQ bands and trend semanticsCONTEXT.md: the project's domain vocabulary
Docs
| Doc | Contents |
|---|---|
| ADR-0001 | Why the browser never hits upstream directly (host proxy) |
| ADR-0002 | Per-dataset freshness-window refresh throttling |
| CONTEXT.md | Domain glossary (model tier, harness, trend, …) |
Troubleshooting
No Model Radar tab in Settings
- Make sure
npm run buildproducedlib/client.js. - Use
dev_plugin_statusto confirm the plugin is active. - Refresh the Web GUI once to load the latest client dependency graph.
No readout in the composer
- Make sure the current model exists on the DeepSWE leaderboard.
- Same-base fallback shows
≈; a fully unknown base is hidden by design. - Make sure the official model-selection UI plugin is enabled.
Upstream refresh failures
With snapshot history present, the settings page shows the last successful data. Check network reachability of api.codexradar.com; the API needs no credentials.
liangmianya/dsh-synapse
alaliqing/claude-paper
omdsh-dev/dsh-annotation
Anionex/dsh-turn-rewind
hanshenmesen/dsh-turn-delete
qkycir-123/dsh-run2skill
Tyan66666/billion-context-dsh
PKUfudawei/dsh-capability-menu