baisama-cloud/dsh-stt-input
DeepSeek Harness(DSH)网页图形界面语音转文字输入插件:在编辑器中点击麦克风图标,即可将语音转换为输入框中的文字。采用浏览器网页语音API及兼容OpenAI的Whisper(OpenAI/Groq),支持可选择的模型。DSH语音输入插件
Project Overview项目介绍
This is a native speech-to-text voice input plugin built exclusively for the DeepSeek Harness (DSH) web GUI. After installation, it adds a clickable microphone button to the left side of the conversation input box. Tapping the button once starts recording, and tapping it again stops the recording process, then automatically inserts the recognized text into the input box ready for sending. The plugin supports two different recognition engines to choose from, giving users flexibility based on their needs and the browser they are using. The first engine is browser-based local recognition, and the second is cloud API-based recognition.
The browser local recognition engine uses the Web Speech API built into Chrome and Edge, so it requires no extra configuration or API keys. It outputs intermediate recognition results as you speak, updating the input box in real time. The API recognition engine records audio via MediaRecorder, then sends it to any OpenAI-compatible /v1/audio/transcriptions endpoint, including OpenAI, Groq, or self-hosted endpoints. Users can select from common Whisper models or enter a custom model name for their provider. Groq currently offers free access to the whisper-large-v3 model, making it a popular low-cost option for users.
To install the plugin, you first pack it with pnpm pack, then add the resulting tarball as a dependency to your DSH web profile in ~/.dsh/profiles/web, then restart the dsh web service. Privacy is handled carefully: your API key is only stored in page memory, never saved to disk or logged. Non-secret configuration is stored in localStorage so it persists across page refreshes. The plugin is released under the permissive MIT license, so it is free to use and modify. Note that Firefox does not support the local recognition engine, so Firefox users must use the API recognition option.
这是专为 DeepSeek Harness (DSH) Web 图形界面开发的语音输入原生插件。它会在对话输入框旁添加一个麦克风按钮,点击开始录音,再次点击停止,识别完成后的文字会自动填入输入框。插件支持两种识别引擎:基于浏览器 Web Speech API 的本地零配置识别,以及基于 OpenAI 兼容接口的 API 识别,可自定义选择模型与服务提供商。
用户安装插件后,需要先进入设置页的「语音输入」板块配置识别引擎。使用浏览器本地识别仅需在 Chrome 或 Edge 浏览器中开启,无需额外配置 API 密钥,即可边识别边输出中间结果。如果使用 API 识别,可以选择 OpenAI、Groq 预设或者自定义服务,填入 API Base URL 和密钥即可使用。
插件遵循 MIT 开源协议,可免费使用。隐私方面,API 密钥仅保存在页面内存中,不落盘也不记录日志,非密钥配置会通过 localStorage 保留,刷新页面后无需重新配置。安装时需要打包后添加到 DSH 的 web 配置依赖,重启 DSH 服务后生效,仅 Chrome、Edge 支持本地识别,Firefox 需使用 API 识别。
请帮我安装这个 DSH 插件。安装前先完成【兼容性检查 + 安全性检查】,检查通过再动手。
插件:dsh-stt-input(baisama-cloud/dsh-stt-input)
仓库:https://github.com/baisama-cloud/dsh-stt-input
本站详情页:https://www.yhbd.top/plugins/baisama-cloud-dsh-stt-input/
本站登记:类型 plugin · 归类 原生 DSH 插件 · 许可证 MIT · ⭐ 3 · 最近提交 2026-08-31 · 主语言 JavaScript
按下面顺序执行,每步先把结论告诉我,再进入下一步:
【1 兼容性检查】
① 我这边:DSH 版本、Node 版本、操作系统、当前 profile(web / desktop)。
② 读它的 README、package.json、插件 manifest,列出它要求的 DSH 版本 / Node 版本 / 操作系统 / 外部依赖 / 需要另外先装的运行时。
③ 逐条比对,结论只写「满足 / 不满足 / 未知」三种;不满足的给出可行替代方案。
④ 检查是否和我已装的插件冲突:命令名重复、skill / tool 重名、端口占用、重复注册的 MCP server。
【2 安全性检查】
① 仓库可信度:和上面「本站登记」是否一致;star / fork 数、创建时间、最近提交,是否归档或长期停更。
② 安装脚本:逐行看 package.json 的 preinstall / install / postinstall,以及 install.sh、setup.ps1 之类脚本。出现 curl|bash、下载后直接执行、混淆代码、访问与插件功能无关的域名,立刻停下来告诉我,不要继续装。
③ 依赖:列出新增依赖,标出无人维护、或与知名包拼写近似的可疑包(typosquatting)。
④ 权限与副作用:它会读写哪些目录、访问哪些域名、需要哪些 DSH 权限(filesystem / network / shell / clipboard 等),以及怎么卸载和回滚。
⑤ 如果它要求 sudo / 管理员权限,或权限明显超出功能所需,先停下来问我。
【3 安装】
上面两步没有「不满足」和「高危项」时才执行;用官方推荐方式安装,不要自行提权。
【4 汇报】
用表格输出:检查项 / 结论 / 依据 / 是否需要我决策。拿不准的一律写「未知」并说明要我怎么确认——不要猜,也不要替我决定。
Send this message to DSH in your current session: it verifies compatibility and security first (answering met / not met / unknown item by item) and only installs once everything checks out — it will stop and ask you if it finds a high-risk item. The box scrolls; the copy is the full prompt. CLI install commands may not be accurate across systems, so DSH is the safer route.把上面这条消息直接发给当前会话里的 DSH:它会先核对兼容性与安全性(逐条给「满足 / 不满足 / 未知」),确认没问题再安装,有高危项会停下来问你。框内可滚动,复制到的是完整提示词;安装命令不一定准确,发给 DSH 更稳。
- Only 3 stars - very few users, little community feedback星标只有 3,几乎没人在用,遇到问题缺少社区反馈
DSH walks through these 9 checksDSH 会逐条核对这 9 项
Compatibility兼容性
- DSH, Node, OS and profile requirementsDSH 版本 / Node 版本 / 操作系统 / profile 是否满足要求
- External dependencies and runtimes (Electron / Python / Docker, ...)外部依赖与运行时(Electron / Python / Docker 等)是否齐备
- Conflicts with installed plugins: command names, skill / tool names, ports, duplicate MCP registration与已装插件是否冲突:命令名、skill / tool 重名、端口占用、重复 MCP 注册
Security安全性
- Repo matches the facts registered here; archived or abandoned?仓库是否与页面登记一致,是否归档或长期停更
- Safety of preinstall / install / postinstall and install.sh / setup.ps1preinstall / install / postinstall 与 install.sh、setup.ps1 是否安全
- curl|bash, download-then-execute, obfuscation, unrelated domains → stop immediatelycurl|bash、下载即执行、混淆代码、无关域名 → 立刻停止
- Typosquatting or unmaintained packages among the new dependencies新增依赖里有没有 typosquatting 或无人维护的包
- Requested permissions vs. what the feature actually needs申请了哪些权限、是否超出功能所需(filesystem / network / shell / clipboard)
- Any sudo / admin requirement, plus uninstall and rollback是否要求 sudo / 管理员权限,以及卸载与回滚方式
Anything uncertain must be marked unknown with a note on how to confirm it. This site's signal screen is a static snapshot, not a security audit.拿不准的必须标「未知」并说明要我怎么确认。本站的信号筛查是静态快照,不能替代安全审计。
Or use CLI install (for developers)或使用命令行安装(适合开发者)
CLI Install命令行安装
dsh plugin --profile web add github:baisama-cloud/dsh-stt-input
把 baisama-cloud/dsh-stt-input 加入你的 DSH 配置(web profile)即可启用。
READMEREADME
中文 · English
dsh-stt-input
DeepSeek Harness (DSH) Web GUI 的语音输入插件。
点击输入框旁的 🎤 麦克风按钮开始说话,再点一次停止,识别文字自动填入输入框。 识别引擎与模型可在 设置 → 语音输入 中选择。
功能
- 两种识别引擎
- 浏览器本地识别 — 使用浏览器自带的 Web Speech API(
SpeechRecognition, Chrome/Edge)。零配置、无需 API Key,边说边把中间结果写进输入框。 - API 识别 — 用
MediaRecorder录音,通过任意 OpenAI 兼容/v1/audio/transcriptions接口(OpenAI、Groq、自定义)转写。
- 浏览器本地识别 — 使用浏览器自带的 Web Speech API(
- 模型可选 —
whisper-1(OpenAI)、whisper-large-v3、whisper-large-v3-turbo、distil-whisper-large-v3-en(Groq),或自定义模型名。 - 可配置 — 服务预设(OpenAI / Groq / 自定义)、API Base URL、API Key、 识别语言、写入方式(追加到输入框 / 替换输入框内容)。
- 实时状态条 — 输入框下方显示录音计时、识别中状态与错误信息。
- 隐私 — API Key 仅保存在页面内存中,不落盘、不打日志;非密钥配置通过
localStorage在刷新后保留。
安装
打包 tarball 后安装到你的 DSH web profile(与其他 dsh-* 插件一致):
pnpm pack
# 把 dsh-stt-input-*.tgz 复制到 web profile 并添加依赖,
# 例如在 ~/.dsh/profiles/web 下:pnpm add ../path/to/dsh-stt-input-0.1.0.tgz
# 然后重启 `dsh web`。
插件注册了三个界面位:
- 输入框工具行的麦克风按钮(
conversation.input.left) - 输入框下方的状态条(
conversation.composer.dock) - 设置页(
settings.section→ 语音输入)
使用
- 打开 设置 → 语音输入 选择引擎。
- 浏览器本地识别:无需其他配置(Chrome/Edge)。
- API 识别:选择预设(OpenAI 或 Groq)、模型,并粘贴 API Key。
Groq 的
whisper-large-v3目前免费。
- 点击输入框旁的 🎤 开始录音,说话,再点一次停止。识别文字进入输入框,回车发送。
浏览器本地引擎依赖 Chrome/Edge 的 Web Speech API;Firefox 请使用 API 引擎。
工作原理
┌──────────┐ 点击🎤 ┌───────────────┐
│ 客户端 │ ─────────────▶ │ MediaRecorder │ (api 引擎)
│ (浏览器) │ │ SpeechRecog. │ (browser 引擎)
└──────────┘ └──────┬────────┘
▲ ▼
│ setDraft(text) base64 音频 (JSON)
│ POST /stt-input/transcribe
│ │
┌─────┴──────┐ ┌───────▼────────┐
│ 输入框 │ ◀──────────│ Host (Node) │
└────────────┘ {ok,text} │ fetch → /v1/ │
│ audio/transcr.│
└────────────────┘
客户端(lib/client.js)录音后把 base64 JSON POST 到宿主路由
/stt-input/transcribe;宿主(lib/index.js)解码音频并以
multipart/form-data 上传到 ${baseUrl}/v1/audio/transcriptions
(使用 Node ≥ 18 的全局 fetch / FormData / Blob)。
License
MIT
omdsh-dev/dsh-browser
yinnho/aginxbrowser
BrambleXu/dsh-annotate
welsione/dsh-mmx-bridge
linhut/dsh-stock-terminal