Xidong-AI/dsh-rate-limiter

Plugin插件 Native原生 ⭐ 3 MIT Models & Routing模型与路由

DeepSeek Harness的主动速率限制器插件。

Project Overview项目介绍

dsh-rate-limiter is a proactive rate-limiting plugin for DeepSeek Harness (dsh). It applies a per-provider token bucket before each request is issued; when the limit is exceeded, the request is queued with a delay instead of failing, preventing upstream 429 errors. Use it when calling rate-limited providers, alongside the official dsh-llm-retry which handles post-failure backoff. Unconfigured providers pass through unchanged. The plugin hooks agent/request, never alters request content or routing, and honors the abort signal during queued waits. It uses a hand-written reservation-based token bucket with no third-party rate-limiting dependencies. Note: it only controls when a request is issued, not what it contains.

dsh-rate-limiter 是 DeepSeek Harness (dsh) 的主动限流插件:按 provider 用令牌桶在请求发出前控速,超限时排队延迟而非失败,从而避免上游 429。与官方 dsh-llm-retry(事后退避重试)互补,一个事前预防、一个事后兜底,互不干扰。未配置的 provider 原样放行,零侵入。挂载于 agent/request,等待可被 abort 信号立即中断。手写预留式令牌桶,零三方限流依赖。注意:仅控制请求时机,不改动请求内容与路由。

Or use CLI install (for developers)或使用命令行安装(适合开发者)

CLI Install命令行安装

dsh plugin --profile web add @xidong-ai/dsh-rate-limiter

Xidong-AI/dsh-rate-limiter 加入你的 DSH 配置(web profile)即可启用。

READMEREADME

dsh-rate-limiter

English | 中文

A proactive rate limiter plugin for DeepSeek Harness (dsh): it controls the request rate per provider (token bucket) before model requests are issued, and queues the request with a delay instead of failing when the limit is exceeded — avoiding upstream 429s.

It complements the official dsh-llm-retry (exponential backoff after failure): rate limiting comes first (prevention), backoff comes last (safety net); the two do not interfere with each other.

Features

  • Per-provider token bucket, enforced before the request is sent (proactive prevention)
  • Over-limit requests are queued with a delay instead of rejected (no 429s, no lost requests)
  • Unconfigured providers pass through untouched (zero intrusion)
  • Queued waits honor the abort signal: stopping the user interrupts the wait immediately
  • Hand-written reservation-based token bucket (concurrency-safe), zero third-party rate-limiting dependencies
  • Mounts on agent/request, coexists naturally with dsh-llm-retry

Installation

Install from npm:

dsh plugin --profile web add @xidong-ai/dsh-rate-limiter

npm registry URLs are case-sensitive; use the lowercase package name.

Or install directly from GitHub:

dsh plugin --profile web add github:Xidong-AI/dsh-rate-limiter

For local development, add the checkout directly:

dsh plugin --profile web add .

After installing, dsh --profile web --dump-config should show the plugin entry:

- id: rate-limiter
  name: @xidong-ai/dsh-rate-limiter
  config:
    enabled: true
    providers: {}

Configuration

Configure the token bucket per provider in the profile's cordis.patch.yml (or this plugin's cordis.patch.yml):

- id: rate-limiter
  config:
    enabled: true
    providers:
      nvidia:
        rate: 0.5        # tokens/second (long-term average QPS)
        burst: 1         # bucket capacity (allowed burst requests)
      sensenova:
        rate: 0.02778
        burst: 1
  • rate: refill rate (tokens/second), i.e. the long-term average request rate.
  • burst: bucket capacity, the number of burst requests allowed.
  • Providers not listed are not rate-limited; requests pass through untouched (zero intrusion).
  • enabled: false disables the plugin entirely.

How It Works

The plugin hooks onto the agent/request waterfall: it await next() first to obtain the call config (which carries the provider), then performs a per-provider token bucket check; when tokens are insufficient, it queues the request with a delay (interrupted immediately by the abort signal when the user stops), then returns the config unchanged — it never modifies request content, never changes routing, never swallows errors. It only controls when a request is issued.

The rate-limiting algorithm is a hand-written reservation-based token bucket (concurrency-safe), with zero third-party rate-limiting dependencies.

Relationship with dsh-llm-retry

Plugin Timing Behavior
dsh-rate-limiter Before the request is issued Queue with a delay when over the limit (prevents 429s)
dsh-llm-retry After the request fails Exponential backoff retry (safety net)

They mount at different points (agent/request vs agent/request-error) and coexist naturally.

Uninstall

dsh plugin --profile web remove @xidong-ai/dsh-rate-limiter

Development

npm install
npm run typecheck   # tsc --noEmit
npm run test        # vitest run
npm run build       # esbuild transpiles lib/*.ts → lib/*.js

Acknowledgements

Thanks to the Linux.do community for support.

上一个 Prev dsh-scope 下一个 Next dsh-session-handoff