WayneYu430/dsh-voice-agent

a full duplex voice mode for DSH

Project Overview项目介绍

dsh-voice-agent is a native voice frontend built specifically for the dsh (DeepSeek Harness) Web client, providing full-duplex realtime conversation on top of ByteDance's Duplex speech stack. The plugin captures microphone audio, transcribes it, hands the text to dsh, and streams Duplex-synthesized speech back, letting users talk to a coding agent the way one would talk to ChatGPT's advanced voice mode. Inside the browser the frontend exposes only three orchestration tools — realtime_delegation, send_task_message, and cancel_task — so any real work is delegated to a separate backend agent rather than hallucinated by the voice loop.

The typical workflow pairs the browser voice session with a long-lived background dsh session: when the user asks for a task such as searching code, running a build, or editing files, realtime_delegation spawns an independent session that runs while the user keeps talking, switches tabs, or disconnects; STATUS and COMPLETE events are pushed back into the voice stream and read aloud as facts. It targets dsh users who want hands-free, interruptible interaction and a task result read out instead of guessed. Installation is dsh plugin --profile web add @wayneyu430227/dsh-voice-agent, after which dsh web loads the voice UI on top of the standard dsh client.

Runtime dependencies include a global install of @deepseek-ai/dsh to provide the dsh binary, plus two environment variables, DUPLEX_API_KEY (Volcano Engine / ByteDance access key) and DUPLEX_APP_KEY (matching app key); missing either causes the Duplex provider handshake to fail before any audio starts. The mic UI is built from a copied dsh client tsdown preset and loaded via window.__ModuleLoader__, so it is not a framework-agnostic browser extension. Voice sessions persist voice/* events strictly, and stock dsh builds without this plugin will reject those records on load.

dsh-voice-agent 是一款面向 dsh(DeepSeek Harness)Web 端的原生语音前端,基于字节跳动 Duplex 提供全双工实时对话能力。它把麦克风采集的语音请求转换成文本后交给 dsh,并由 Duplex 把模型答复合成语音播报,使用户可以像 ChatGPT 高级语音那样和 dsh 自然聊天。前端浏览器仅暴露 realtime_delegation、send_task_message、cancel_task 三个编排工具,把"任务"派发到独立后台 Agent 执行。

工作流程由前端与后台 Agent 共同完成:用户开口提出查询、构建、文件修改等需求时,语音前端通过 realtime_delegation 创建一个独立后台会话执行任务,期间用户可以继续对话或切换会话,任务完成后再由语音把 STATUS 与 COMPLETE 结果回灌并播报。它面向希望通过语音替代键盘、长时间离开屏幕却想持续推进任务的 dsh 编码 Agent 用户,安装命令为 dsh plugin --profile web add @wayneyu430227/dsh-voice-agent,随后运行 dsh web 即可加载 voice 界面。

依赖方面,运行需全局安装 @deepseek-ai/dsh 以获得 dsh 命令,并需设置环境变量 DUPLEX_API_KEY(火山引擎 access key)与 DUPLEX_APP_KEY(对应 app key),否则 Duplex provider 握手失败。注意浏览器麦克风与播放界面是 dsh Web UI 的一部分,由复制得到的 dsh client tsdown 预设构建,通过 window.__ModuleLoader__ 契约加载,并非框架无关的浏览器插件;voice session 持久化 voice/* 事件,不带本插件的 dsh 构建加载这些记录会被拒绝。

Pre-install check安装前体检Compatibility · Security兼容性 · 安全性 3 warnings3 项注意
  • No license declared - all rights reserved by default; ask the author before commercial use or redistribution未声明开源许可证 —— 默认「保留所有权利」,商用或再分发前先问作者
  • Only 2 stars - very few users, little community feedback星标只有 2,几乎没人在用,遇到问题缺少社区反馈
  • No DSH plugin manifest detected - it may only carry the dsh-plugin topic, so the install method must be confirmed on the spot未检测到 DSH 插件清单:可能只是打了 dsh-plugin 话题,安装方式要现场确认
DSH walks through these 9 checksDSH 会逐条核对这 9 项

Compatibility兼容性

  • DSH, Node, OS and profile requirementsDSH 版本 / Node 版本 / 操作系统 / profile 是否满足要求
  • External dependencies and runtimes (Electron / Python / Docker, ...)外部依赖与运行时(Electron / Python / Docker 等)是否齐备
  • Conflicts with installed plugins: command names, skill / tool names, ports, duplicate MCP registration与已装插件是否冲突:命令名、skill / tool 重名、端口占用、重复 MCP 注册

Security安全性

  • Repo matches the facts registered here; archived or abandoned?仓库是否与页面登记一致,是否归档或长期停更
  • Safety of preinstall / install / postinstall and install.sh / setup.ps1preinstall / install / postinstall 与 install.sh、setup.ps1 是否安全
  • curl|bash, download-then-execute, obfuscation, unrelated domains → stop immediatelycurl|bash、下载即执行、混淆代码、无关域名 → 立刻停止
  • Typosquatting or unmaintained packages among the new dependencies新增依赖里有没有 typosquatting 或无人维护的包
  • Requested permissions vs. what the feature actually needs申请了哪些权限、是否超出功能所需(filesystem / network / shell / clipboard)
  • Any sudo / admin requirement, plus uninstall and rollback是否要求 sudo / 管理员权限,以及卸载与回滚方式

Anything uncertain must be marked unknown with a note on how to confirm it. This site's signal screen is a static snapshot, not a security audit.拿不准的必须标「未知」并说明要我怎么确认。本站的信号筛查是静态快照,不能替代安全审计。

Or use CLI install (for developers)或使用命令行安装(适合开发者)

CLI Install命令行安装

dsh plugin --profile web add @wayneyu430227/dsh-voice-agent

把 WayneYu430/dsh-voice-agent 加入你的 DSH 配置(web profile)即可启用。

READMEREADME

dsh-voice-agent

English | 中文

像 ChatGPT 高级语音那样,用自然的语音和你的 dsh 编码 Agent 对话——而且不止是聊天,它能真正把活干完。

你开口说一句需求,dsh 立刻用语音回应你。当你要的是一份「任务」(查代码、跑构建、改文件……),语音前端会把它派给一个独立的后台 Agent 去执行;你随时可以继续说话、切走看别的会话,等任务做完,它再用语音把结果念给你听,而不是凭空编一个结论。

交互

  • 全双工实时对话:基于 ByteDance Duplex,像和人说话一样自然,边说边听、随时打断、随时插话纠正。
  • 对话式派活:前端只暴露三个编排工具——realtime_delegation(把「帮我查一下 xxx」变成真正的后台任务)、send_task_message(补充要求 / 纠正方向)、cancel_task(取消)。
  • 异步结果回灌:任务跑在独立 Session 里,进度(STATUS)和最终结果(COMPLETE)会回灌进语音对话,前端按事实播报。
  • 不中断的体验:浏览器切走、断线重连,正在跑的语音会话和后台任务都不会停。

安装

dsh plugin --profile web add @wayneyu430227/dsh-voice-agent

dsh 命令来自 npm install -g @deepseek-ai/dsh。启动 web(voice 界面随之加载):

dsh web

凭据

Duplex provider 从环境读取两个凭据引用:

  • DUPLEX_API_KEY —— ByteDance 火山引擎 access key。
  • DUPLEX_APP_KEY —— 对应的 app key。

开始语音对话前需设置两者;缺少时 provider 会话握手失败。

限制

  • 浏览器麦克风与播放界面面向 dsh Web UI:它由复制而来的 dsh client tsdown 预设构建,并通过 dsh web 运行时的 window.__ModuleLoader__ 契约加载,因此不是框架无关的浏览器插件。
  • Voice Session 记录按「读取时必需」处理的持久 voice/* 事件;不带本插件的 dsh 构建加载它们会被拒绝。
← 上一个 Prev dsh-node-nav 下一个 Next dsh-offpeak-saver →