beiyege-01/dsh-voice-ai-girlfriend-plugin

Plugin插件 Native原生 ⭐ 12 Vision & Media视觉与多媒体

Project Overview项目介绍

This is a native plugin built exclusively for DeepSeek Harness (DSH), that adds full voice interaction and AI girlfriend digital human functionality to the DSH platform. It can be installed directly with a single dsh plugin add github:beiyege-01/dsh-voice-ai-girlfriend-plugin command via DSH's built-in plugin manager. It currently supports both DSH version 0.1.3 (profile web-v013) and the older rc.8 version, following DSH's official plugin manifest and slot protocol. It enables microphone voice input recognition powered by local FunASR, streaming TTS reply reading via local OmniVoice, a dockable AI animation window, and two-way QQ conversation forwarding.

Before you can use the plugin, you first need to clone the companion main repository beiyege-01/dsh-voice-ai-girlfriend, which contains the Python voice bridge service, one-click startup scripts, and full setup documentation. You will need to deploy the voice bridge, which runs FunASR for speech recognition and OmniVoice for TTS, on your local machine. If you want to use the two-way QQ conversation feature, you will also need to deploy a separate NapCatQQ service. This plugin is designed for users who want a fully local interactive voice and digital AI girlfriend experience directly within DSH.

The digital human feature requires pulling a specified Docker image, and the cold first run takes between 1 and 3 minutes to complete. You can turn off the digital human service when not in use to free up GPU video memory. The plugin author is currently preparing for a wedding, so adaptation work for future DSH versions will resume after the National Day holiday, but current supported versions are fully functional and ready to use. If you install with pnpm version 10 or newer, you will need to manually add esbuild to the allowBuilds list in your pnpm-workspace.yaml to complete the build.

这是一款专为 DeepSeek Harness 开发的原生插件,提供语音交互和 AI 女友数字人完整能力。它支持通过 FunASR 实现麦克风语音输入识别,通过 OmniVoice 实现流式 TTS 回复朗读,还带可停靠的 AI 女友动画窗口,以及 QQ 双向对话转发功能。用户可直接通过 dsh plugin add 命令安装,目前同时适配 DSH 0.1.3 和 rc.8 两个版本。

使用本插件前,需要先克隆配套的主仓库,部署包含 FunASR 语音识别和 OmniVoice TTS 的 Python 语音桥接服务,若需要 QQ 双向对话功能,还需额外部署 NapCatQQ。用户可通过自行添加对应资源文件,自定义 TTS 音色、数字人形象、待机动画,全程无需修改代码。它适合想要基于 DSH 实现带语音交互和数字人的 AI 对话体验的用户使用。

数字人模块需要拉取指定的 Docker 镜像,冷启动需要 1 到 3 分钟,不使用时可关闭数字人服务释放显存。作者目前正在筹备婚礼,新版本适配工作将顺延至国庆后,当前列出的适配版本功能完整,均可正常使用。若使用 pnpm ≥10 安装,需要手动将 esbuild 添加到 pnpm-workspace.yaml 的 allowBuilds 列表中才能完成构建。

Pre-install check安装前体检Compatibility · Security兼容性 · 安全性 1 warning1 项注意
  • No license declared - all rights reserved by default; ask the author before commercial use or redistribution未声明开源许可证 —— 默认「保留所有权利」,商用或再分发前先问作者
  • 12 stars - an early-stage project星标 12,属于早期项目
DSH walks through these 9 checksDSH 会逐条核对这 9 项

Compatibility兼容性

  • DSH, Node, OS and profile requirementsDSH 版本 / Node 版本 / 操作系统 / profile 是否满足要求
  • External dependencies and runtimes (Electron / Python / Docker, ...)外部依赖与运行时(Electron / Python / Docker 等)是否齐备
  • Conflicts with installed plugins: command names, skill / tool names, ports, duplicate MCP registration与已装插件是否冲突:命令名、skill / tool 重名、端口占用、重复 MCP 注册

Security安全性

  • Repo matches the facts registered here; archived or abandoned?仓库是否与页面登记一致,是否归档或长期停更
  • Safety of preinstall / install / postinstall and install.sh / setup.ps1preinstall / install / postinstall 与 install.sh、setup.ps1 是否安全
  • curl|bash, download-then-execute, obfuscation, unrelated domains → stop immediatelycurl|bash、下载即执行、混淆代码、无关域名 → 立刻停止
  • Typosquatting or unmaintained packages among the new dependencies新增依赖里有没有 typosquatting 或无人维护的包
  • Requested permissions vs. what the feature actually needs申请了哪些权限、是否超出功能所需(filesystem / network / shell / clipboard)
  • Any sudo / admin requirement, plus uninstall and rollback是否要求 sudo / 管理员权限,以及卸载与回滚方式

Anything uncertain must be marked unknown with a note on how to confirm it. This site's signal screen is a static snapshot, not a security audit.拿不准的必须标「未知」并说明要我怎么确认。本站的信号筛查是静态快照,不能替代安全审计。

Or use CLI install (for developers)或使用命令行安装(适合开发者)

CLI Install命令行安装

dsh plugin --profile web add github:beiyege-01/dsh-voice-ai-girlfriend-plugin

把 beiyege-01/dsh-voice-ai-girlfriend-plugin 加入你的 DSH 配置(web profile)即可启用。

READMEREADME

DSH 语音 AI 女友插件(voice-ai-girlfriend)

DeepSeek Harness 的浏览器语音插件:麦克风语音输入、回复朗读(TTS)、AI 女友动画窗、QQ 双向对话。

✅ 已适配 DSH rc.8 与 dsh 0.1.3(最近实测 2026-09-14,v0.3.0):dsh.bundle manifest + conversation.input.dock/left 槽位按 rc.8 协议实现;2026-09-05 已在 dsh v0.1.3-alpha.1(profile web-v013)实测通过 —— 聊天节点读取用 0.1.3 的 useChat 视图(assistant-step 节点)、桥接 /api/tts(语音朗读)与 /api/dh/speak(数字人)均正常触发;build.mjs 同时产出服务端入口 lib/index.js 与内联 <style> 的浏览器包,修复 dsh plugin add 安装与样式注入。含数字人(DUIX)控制、DeepSeek 余额 badge(stats 行)、五态麦克风。dsh plugin add 即可安装。

🎥 成品展示(抖音):

⚠️ 本插件是完整方案的一部分:语音识别/合成依赖配套的 voice bridge(Python 服务,FunASR + OmniVoice TTS(WSL2 + FlashInfer)),QQ 对话依赖 NapCatQQ。请配合完整仓库使用:beiyege-01/dsh-voice-ai-girlfriend(含桥接代码、模型准备、一键启动脚本、安装文档)。

适配版本

组件 适配版本 状态
DSH(主用) dsh 0.1.3-alpha.1 / profile web-v013(端口 3080) ✅ 2026-09-14 实测通过
DSH(兼容) rc.8(E:\DSH\deepseek-harness master 141eb6fef)/ profile web ✅ 仍可加载
本插件 @beiyege-01/dsh-voice-ai-girlfriend v0.3.0 构建:node build.mjs → lib/client.js
语音桥接 voice_bridge.py(FastAPI/uvicorn :8765,Python 3.14 venv) 配套主仓库
TTS 引擎 OmniVoice(WSL2 + FlashInfer :9877) 音色全部为 OmniVoice 克隆
数字人 DUIX guiji2025/duix.avatar-5090:trt10.9(宿主 :9000) 冷启动 1-3 分钟

v0.3.0 新增:数字人开关下沉到桥接(关掉=桥接停止预热/提交,并可联动 docker stop 释放显存;打开=docker start + 就绪后预热)、生成中按钮周期性绿光、卡片左下角分段进度(第 N/M 段)与语音总开关提示。

维护节奏(2026-09)

作者本月正在筹备婚礼,DSH 新版本的适配与兼容性跟进顺延到国庆之后(10 月上旬)。这段时间欢迎继续提 issue / PR,但响应可能会慢一些;上表所列版本功能完整、可正常使用。

安装

dsh plugin --profile web add github:beiyege-01/dsh-voice-ai-girlfriend-plugin

pnpm ≥10 对 git 依赖的构建脚本有 allowBuilds 限制:按安装时 pnpm 打印的提示,在 profile 的 pnpm-workspace.yaml 的 allowBuilds 里加一行允许 esbuild 后重跑。

功能

  • 🎙️ 语音输入:麦克风 → 桥接 FunASR 中文识别 → 注入对话(连续聆听、自动端点)
  • 🔊 语音回复:回复流式 TTS 朗读(OmniVoice 声音克隆,跑在 WSL2 + FlashInfer 加速,音色由参考音频决定,支持 600+ 语言)
  • ⚡ 插话/排队:亮=说话打断回复;灭=回复读完句子排队接上
  • 👧 数字人窗口:右侧动画窗(空闲/说话视频,素材自备),自动贴右侧插件边缘不被遮挡
  • 💬 QQ 双向:QQ 消息注入对话 + 回复文本/语音/图片推送(NapCat)
  • 🎭 三类预设一键切换:工具行三个按钮分别循环切换 TTS 音色 / 数字人形象 / 待机动画,选择自动记忆(localStorage)。音色(voices/<名字>/)、待机(assets/bg-images/<名字>/)、形象(共享卷 temp 放 mp4)均可自行添加,无需改代码

使用前提

  1. 克隆主仓库并按其 README 装好 voice bridge(start-all.cmd 一键)
  2. 本插件的桥接地址默认 http://127.0.0.1:8765(localStorage s2s.voice.bridge 可改)

构建

npm install
node build.mjs   # 产物 lib/client.js(browser bundle)
← 上一个 Prev dsh-plugin-skills 下一个 Next dsh-multi-tenant →