PandaPolo/dsh-voice-call 预览 preview

PandaPolo/dsh-voice-call

Plugin插件 Native原生 ⭐ 2 MIT Vision & Media视觉与多媒体

给DeepSeek Harness代理一个属于它自己的声音——offer_call呼叫人类(接听/拒接/稍后);通过CrispASR和Qwen3-TTS CustomVoice实现本地TTS,支持9个说话人,2种中文方言。本地优先,支持离线使用。

Project Overview项目介绍

This is a native plugin built exclusively for DeepSeek Harness that enables AI agents running on the DSH platform to initiate voice calls and speak text directly to users. Following a core design principle of agent dialing rights and user answering control, the agent decides when to place a call, while the user retains full control over accepting, rejecting, or deferring the incoming call. No audio will play without explicit user consent, and all speech processing runs locally via CrispASR and Qwen3-TTS, enabling full offline usage with all audio stored locally on the user’s machine.

To install the plugin, users can run the dsh plugin --profile web add dsh-voice-call command via the DSH CLI to pull the published version from npm. After installation, users just need to update the plugin configuration in their profile’s cordis.patch.yml file with the correct paths for their local speech engine and Qwen3-TTS model files, then restart the DSH web client. Once set up, users can tell the agent to use the offer_call tool when it has something important to share, and the agent will initiate an incoming call with a styled notification card.

The plugin is released under the open source MIT license, and requires Node.js 20 or newer, DSH CLI 0.1.2-rc.1, and works across popular Windows 10/11, macOS, and Linux desktop operating systems. Audio recording via the transcribe tool is only supported on macOS, and Linux users need to install ALSA system utilities to enable audio playback through the aplay command. Local engine commands run with a full access policy to access engine binaries, model files, and audio directories across different filesystem roots, so users should evaluate correct trust boundaries before deployment.

这是一个专为 DeepSeek Harness 开发的原生插件,为运行在 DSH 中的 AI 代理提供语音通话和语音朗读能力。它遵循「代理拥有拨号权、人类拥有接听权」的设计理念:代理自主决定何时发起通话,用户掌握接听、拒接或延后的决定权,未经用户同意绝不会播放任何声音。插件全程本地优先,可完全离线运行,基于 CrispASR 和 Qwen3-TTS 本地引擎实现语音合成与识别,所有音频文件都存储在本地 ~/.dsh/voice/ 目录下。

安装完成后,用户可以在会话中提示代理,让代理在有重要信息需要传达时调用 offer_call 工具发起通话。来电会以带有脉冲动画的专属浮层卡片呈现,显示来电者身份和内容预览,无网页客户端连接时会自动回退到弹窗询问。拒接或推迟来电后,插件会将用户的决定返回给代理,帮助代理学习调整沟通方式,适合希望 AI 代理用语音主动沟通的 DSH 用户。

插件采用 MIT 许可证开源,要求 Node.js 版本不低于 20,DSH CLI 版本为 0.1.2-rc.1,支持 Windows 10/11、macOS 和 Linux 平台。录音功能仅支持 macOS,Linux 播放需要额外安装 ALSA 工具,本地引擎命令需要开放完整访问权限,部署前需评估信任边界。目前一共有 87 个单元测试全部通过,v0.2 已经稳定可用,未来计划添加语音信箱功能。

Pre-install check安装前体检Compatibility · Security兼容性 · 安全性 1 warning1 项注意
  • Only 2 stars - very few users, little community feedback星标只有 2,几乎没人在用,遇到问题缺少社区反馈
DSH walks through these 9 checksDSH 会逐条核对这 9 项

Compatibility兼容性

  • DSH, Node, OS and profile requirementsDSH 版本 / Node 版本 / 操作系统 / profile 是否满足要求
  • External dependencies and runtimes (Electron / Python / Docker, ...)外部依赖与运行时(Electron / Python / Docker 等)是否齐备
  • Conflicts with installed plugins: command names, skill / tool names, ports, duplicate MCP registration与已装插件是否冲突:命令名、skill / tool 重名、端口占用、重复 MCP 注册

Security安全性

  • Repo matches the facts registered here; archived or abandoned?仓库是否与页面登记一致,是否归档或长期停更
  • Safety of preinstall / install / postinstall and install.sh / setup.ps1preinstall / install / postinstall 与 install.sh、setup.ps1 是否安全
  • curl|bash, download-then-execute, obfuscation, unrelated domains → stop immediatelycurl|bash、下载即执行、混淆代码、无关域名 → 立刻停止
  • Typosquatting or unmaintained packages among the new dependencies新增依赖里有没有 typosquatting 或无人维护的包
  • Requested permissions vs. what the feature actually needs申请了哪些权限、是否超出功能所需(filesystem / network / shell / clipboard)
  • Any sudo / admin requirement, plus uninstall and rollback是否要求 sudo / 管理员权限,以及卸载与回滚方式

Anything uncertain must be marked unknown with a note on how to confirm it. This site's signal screen is a static snapshot, not a security audit.拿不准的必须标「未知」并说明要我怎么确认。本站的信号筛查是静态快照,不能替代安全审计。

Or use CLI install (for developers)或使用命令行安装(适合开发者)

CLI Install命令行安装

dsh plugin --profile web add github:PandaPolo/dsh-voice-call

把 PandaPolo/dsh-voice-call 加入你的 DSH 配置(web profile)即可启用。

READMEREADME

dsh-voice-call —— agent 拥有的声音,由它主动打给你

dsh-voice-call 标志 —— 一声向外荡开的振铃

CI License: MIT npm version DSH 0.1.7-rc.2 DSH 0.2.0-rc.1 与 0.2.0-rc.2,声明范围 ^0.2.0-rc.1 桌面版(dsh-desktop)运行时 0.2.0-rc.2 已逐项核对 289 个测试全绿

dsh-voice-call 演示 —— 自动放映:旁白、翻页、字幕同步

中文 · English

"这个项目的开始是朴素的——我想知道如果 Agent 知道自己可以发出声音,他会说什么?" —— 人类伙伴,关于这个项目如何开始

给 DeepSeek Harness 的 agent 一个它拥有的声音。 本地优先、可完全离线:合成跑在本机 CrispASR + Qwen3-TTS CustomVoice 引擎上(9 个内置音色,含 2 个中文方言),音频是 ~/.dsh/voice/ 下的普通文件。任何声音都不会自动响起——必须由模型调用工具,或者你接听一次来电。

三条规矩,整个项目围着它们转:

  • agent 拥有拨号权 —— 它在自己觉得值得说的时候调用 offer_call:一个完成的念头、一个里程碑、一句想大声说出来的话。
  • 人类拥有接听权 —— 来电以弹窗或专属来电卡片呈现(接听 / 拒接 / 稍后再说),未经同意绝不播放。
  • 拒接也是教育 —— 被拒接或推迟时,工具把你的决定原样返回给 agent,它学会改用文字写下来,或者只在真正重要时再试一次。

🚀 安装(桌面版)

  1. 打开 DeepSeek Harness 桌面版 → 「插件」页面(设置里的插件管理)。
  2. 在插件市场搜索 dsh-voice-call 并安装;装好后在同一页让它生效(本插件不热加载,可能需要重载插件或重启桌面版)。
  3. 首次使用在插件自己的设置页里配置。

旧版的命令行安装(npm install -g @deepseek-ai/dsh,然后 dsh plugin --profile web add dsh-voice-call)随全局 CLI 一起退役了——本机已卸载全局 DSH,插件管理请走桌面版界面。

版本要求:需要 DSH 0.1.7-rc.2 / 0.2.0-rc.1 / 0.2.0-rc.2(桌面版当前带的就是 0.2.0-rc.2)。

装完到 DSH 的插件设置页,「运行环境」区点一下就把引擎和模型装齐。然后开个会话,对 agent 说:「你有 offer_call 工具——有什么值得说的就打电话给我。」

从零到能听见声音的完整步骤(手动下载路径、全字段配置、故障排查、信任边界)在 docs/deployment.md。

Showing the opening section of the README — the full document lives in the repository以上为 README 开头摘要,完整文档在仓库内 · View the full README on GitHub →在 GitHub 查看完整 README →

← 上一个 Prev dsh-plugin-dev-kb 下一个 Next dsh-siyuan →