baisama-cloud/dsh-stt-input

DeepSeek Harness(DSH)网页图形界面语音转文字输入插件:在编辑器中点击麦克风图标,即可将语音转换为输入框中的文字。采用浏览器网页语音API及兼容OpenAI的Whisper(OpenAI/Groq),支持可选择的模型。DSH语音输入插件

Project Overview项目介绍

This is a native speech-to-text voice input plugin built exclusively for the DeepSeek Harness (DSH) web GUI. After installation, it adds a clickable microphone button to the left side of the conversation input box. Tapping the button once starts recording, and tapping it again stops the recording process, then automatically inserts the recognized text into the input box ready for sending. The plugin supports two different recognition engines to choose from, giving users flexibility based on their needs and the browser they are using. The first engine is browser-based local recognition, and the second is cloud API-based recognition.

The browser local recognition engine uses the Web Speech API built into Chrome and Edge, so it requires no extra configuration or API keys. It outputs intermediate recognition results as you speak, updating the input box in real time. The API recognition engine records audio via MediaRecorder, then sends it to any OpenAI-compatible /v1/audio/transcriptions endpoint, including OpenAI, Groq, or self-hosted endpoints. Users can select from common Whisper models or enter a custom model name for their provider. Groq currently offers free access to the whisper-large-v3 model, making it a popular low-cost option for users.

To install the plugin, you first pack it with pnpm pack, then add the resulting tarball as a dependency to your DSH web profile in ~/.dsh/profiles/web, then restart the dsh web service. Privacy is handled carefully: your API key is only stored in page memory, never saved to disk or logged. Non-secret configuration is stored in localStorage so it persists across page refreshes. The plugin is released under the permissive MIT license, so it is free to use and modify. Note that Firefox does not support the local recognition engine, so Firefox users must use the API recognition option.

这是专为 DeepSeek Harness (DSH) Web 图形界面开发的语音输入原生插件。它会在对话输入框旁添加一个麦克风按钮,点击开始录音,再次点击停止,识别完成后的文字会自动填入输入框。插件支持两种识别引擎:基于浏览器 Web Speech API 的本地零配置识别,以及基于 OpenAI 兼容接口的 API 识别,可自定义选择模型与服务提供商。

用户安装插件后,需要先进入设置页的「语音输入」板块配置识别引擎。使用浏览器本地识别仅需在 Chrome 或 Edge 浏览器中开启,无需额外配置 API 密钥,即可边识别边输出中间结果。如果使用 API 识别,可以选择 OpenAI、Groq 预设或者自定义服务,填入 API Base URL 和密钥即可使用。

插件遵循 MIT 开源协议,可免费使用。隐私方面,API 密钥仅保存在页面内存中,不落盘也不记录日志,非密钥配置会通过 localStorage 保留,刷新页面后无需重新配置。安装时需要打包后添加到 DSH 的 web 配置依赖,重启 DSH 服务后生效,仅 Chrome、Edge 支持本地识别,Firefox 需使用 API 识别。

Pre-install check安装前体检Compatibility · Security兼容性 · 安全性 1 warning1 项注意
  • Only 3 stars - very few users, little community feedback星标只有 3,几乎没人在用,遇到问题缺少社区反馈
DSH walks through these 9 checksDSH 会逐条核对这 9 项

Compatibility兼容性

  • DSH, Node, OS and profile requirementsDSH 版本 / Node 版本 / 操作系统 / profile 是否满足要求
  • External dependencies and runtimes (Electron / Python / Docker, ...)外部依赖与运行时(Electron / Python / Docker 等)是否齐备
  • Conflicts with installed plugins: command names, skill / tool names, ports, duplicate MCP registration与已装插件是否冲突:命令名、skill / tool 重名、端口占用、重复 MCP 注册

Security安全性

  • Repo matches the facts registered here; archived or abandoned?仓库是否与页面登记一致,是否归档或长期停更
  • Safety of preinstall / install / postinstall and install.sh / setup.ps1preinstall / install / postinstall 与 install.sh、setup.ps1 是否安全
  • curl|bash, download-then-execute, obfuscation, unrelated domains → stop immediatelycurl|bash、下载即执行、混淆代码、无关域名 → 立刻停止
  • Typosquatting or unmaintained packages among the new dependencies新增依赖里有没有 typosquatting 或无人维护的包
  • Requested permissions vs. what the feature actually needs申请了哪些权限、是否超出功能所需(filesystem / network / shell / clipboard)
  • Any sudo / admin requirement, plus uninstall and rollback是否要求 sudo / 管理员权限,以及卸载与回滚方式

Anything uncertain must be marked unknown with a note on how to confirm it. This site's signal screen is a static snapshot, not a security audit.拿不准的必须标「未知」并说明要我怎么确认。本站的信号筛查是静态快照,不能替代安全审计。

Or use CLI install (for developers)或使用命令行安装(适合开发者)

CLI Install命令行安装

dsh plugin --profile web add github:baisama-cloud/dsh-stt-input

把 baisama-cloud/dsh-stt-input 加入你的 DSH 配置(web profile)即可启用。

READMEREADME

中文 · English

dsh-stt-input

DeepSeek Harness (DSH) Web GUI 的语音输入插件。

点击输入框旁的 🎤 麦克风按钮开始说话,再点一次停止,识别文字自动填入输入框。 识别引擎与模型可在 设置 → 语音输入 中选择。

功能

  • 两种识别引擎
    • 浏览器本地识别 — 使用浏览器自带的 Web Speech API(SpeechRecognition, Chrome/Edge)。零配置、无需 API Key,边说边把中间结果写进输入框。
    • API 识别 — 用 MediaRecorder 录音,通过任意 OpenAI 兼容 /v1/audio/transcriptions 接口(OpenAI、Groq、自定义)转写。
  • 模型可选 — whisper-1(OpenAI)、whisper-large-v3、 whisper-large-v3-turbo、distil-whisper-large-v3-en(Groq),或自定义模型名。
  • 可配置 — 服务预设(OpenAI / Groq / 自定义)、API Base URL、API Key、 识别语言、写入方式(追加到输入框 / 替换输入框内容)。
  • 实时状态条 — 输入框下方显示录音计时、识别中状态与错误信息。
  • 隐私 — API Key 仅保存在页面内存中,不落盘、不打日志;非密钥配置通过 localStorage 在刷新后保留。

安装

打包 tarball 后安装到你的 DSH web profile(与其他 dsh-* 插件一致):

pnpm pack
# 把 dsh-stt-input-*.tgz 复制到 web profile 并添加依赖,
# 例如在 ~/.dsh/profiles/web 下:pnpm add ../path/to/dsh-stt-input-0.1.0.tgz
# 然后重启 `dsh web`。

插件注册了三个界面位:

  • 输入框工具行的麦克风按钮(conversation.input.left)
  • 输入框下方的状态条(conversation.composer.dock)
  • 设置页(settings.section → 语音输入)

使用

  1. 打开 设置 → 语音输入 选择引擎。
    • 浏览器本地识别:无需其他配置(Chrome/Edge)。
    • API 识别:选择预设(OpenAI 或 Groq)、模型,并粘贴 API Key。 Groq 的 whisper-large-v3 目前免费。
  2. 点击输入框旁的 🎤 开始录音,说话,再点一次停止。识别文字进入输入框,回车发送。

浏览器本地引擎依赖 Chrome/Edge 的 Web Speech API;Firefox 请使用 API 引擎。

工作原理

┌──────────┐  点击🎤        ┌───────────────┐
│  客户端  │ ─────────────▶ │ MediaRecorder │  (api 引擎)
│ (浏览器) │                │ SpeechRecog.  │  (browser 引擎)
└──────────┘                └──────┬────────┘
      ▲                           ▼
      │ setDraft(text)      base64 音频 (JSON)
      │              POST /stt-input/transcribe
      │                           │
┌─────┴──────┐           ┌───────▼────────┐
│  输入框    │ ◀──────────│  Host (Node)   │
└────────────┘  {ok,text} │  fetch → /v1/  │
                          │  audio/transcr.│
                          └────────────────┘

客户端(lib/client.js)录音后把 base64 JSON POST 到宿主路由 /stt-input/transcribe;宿主(lib/index.js)解码音频并以 multipart/form-data 上传到 ${baseUrl}/v1/audio/transcriptions (使用 Node ≥ 18 的全局 fetch / FormData / Blob)。

License

MIT

← 上一个 Prev dsh-moyan 下一个 Next dsh-unread-dot →