ch1bug/dsh-mimo-agent-tools
小米MiMo搜索+多模态工具,用于DeepSeek Harness代理:mimo_search/vision/audio/video/asr/tts
Project Overview项目介绍
DSH plugin that exposes the Xiaomi MiMo API as agent model tools. Capabilities include web search with cited sources (mimo_search), deep thinking with full reasoning chains (mimo_think), structured JSON output (mimo_json), image/audio/video understanding (mimo_vision/audio/video), speech-to-text (mimo_asr), and text-to-speech with voice cloning (mimo_tts/voiceclone). An audio-tools skill guides selection among audio tools. Requires Python3, an XIAOMI_API_KEY resolved through the DSH credentials service, and the connected-search plugin enabled in the MiMo console for web search. Multimodal payloads run through a local Python driver to bypass shell stdout truncation, so the driver must be installed at the default path or pointed to via MIMO_DRIVER. Video base64 input is capped at 50MB per MiMo documentation.
DSH 插件,将小米 MiMo API 封装为代理模型工具。提供网页搜索(mimo_search,含来源引用)、深度思考(mimo_think,返回推理链)、JSON 结构化输出(mimo_json)、图像/音频/视频理解(mimo_vision/audio/video)、语音转文字(mimo_asr)、文字转语音与声音克隆(mimo_tts/voiceclone)等能力;附 audio-tools 技能引导音频工具选用。需 Python3 与 DSH 凭证服务提供的 XIAOMI_API_KEY,网页搜索功能需在 MiMo 控制台启用 connected-search。多模态负载经本地 Python 驱动处理以规避 shell 输出截断;驱动需手动拷贝至默认路径或通过 MIMO_DRIVER 指定。注意:mimo_vision 受文档声明的视频 base64 50MB 上限约束。
请帮我了解并安装插件:【dsh-mimo-agent-tools】【https://github.com/ch1bug/dsh-mimo-agent-tools】
Send this message to DSH in your current session. CLI install commands may not be accurate across systems — DSH will figure it out for you.把上面这条消息直接发给当前会话里的 DSH,让它帮你了解并安装。安装命令不一定准确,发给 DSH 更稳。
Or use CLI install (for developers)或使用命令行安装(适合开发者)
CLI Install命令行安装
dsh plugin --profile web add github:ch1bug/dsh-mimo-agent-tools
把 ch1bug/dsh-mimo-agent-tools 加入你的 DSH 配置(web profile)即可启用。
READMEREADME
mimo-agent-tools
DSH (DeepSeek Harness) Cordis plugin that turns the Xiaomi MiMo API into model tools for an agent: web search, image/audio/video understanding, speech-to-text, and text-to-speech.
Backed by the OpenAI-compatible endpoint https://api.xiaomimimo.com/v1.
Tools
| Tool | Model | Purpose |
|---|---|---|
mimo_search |
mimo-v2.5-pro | Web search via the native web_search tool, returns answer + cited sources (incl. site_name/logo_url); optional force_search, user_location |
mimo_think |
mimo-v2.5-pro / mimo-v2.5 | Deep thinking — full reasoning chain (reasoning_content) + final answer |
mimo_json |
mimo-v2.5-pro / mimo-v2.5 | Structured JSON output via response_format: json_object |
mimo_vision |
mimo-v2.5 | Image understanding — local files (auto base64) or public URLs, multi-image |
mimo_audio |
mimo-v2.5 | Audio understanding / transcription (wav/mp3/flac/ogg/m4a) — local file or public URL |
mimo_video |
mimo-v2.5 | Video understanding (mp4/webm/mov) — local file or public URL; optional fps / media_resolution |
mimo_asr |
mimo-v2.5-asr | Speech-to-text with optional language hint — local file or public URL |
mimo_tts |
mimo-v2.5-tts / -voicedesign | Text-to-speech to a .wav/.mp3 file — preset voices or free-form voice design; optional style (speaking tone) and format (wav/mp3) |
mimo_voiceclone |
mimo-v2.5-tts-voiceclone | Voice cloning — reference clip + text → speech in that voice; optional format (wav/mp3) |
audio-tools skill
The plugin registers an audio-tools skill that guides when to use the four
audio tools (mimo_asr, mimo_tts, mimo_voiceclone, mimo_audio). The
tools themselves are always registered — they are lightweight pure-API calls —
so the skill only teaches usage, it does not gate the tools.
Requirements
- A MiMo API key — the
web_searchtool requires the connected-search plugin enabled in the MiMo console. python3on the host (used bydriver/mimo_driver.py).- DSH host with the
shellandsandboxPolicyservices.
Configuration
The API key is resolved at tool-call time from the DSH credentials
service first (key name XIAOMI_API_KEY — the web Models page writes keys
there), falling back to the environment. Nothing is hardcoded:
| Source | Key | Priority |
|---|---|---|
DSH credentials service (~/.dsh/.credentials.yaml, web Models page) |
XIAOMI_API_KEY |
1 |
| Environment | XIAOMI_API_KEY or MIMO_API_KEY |
2 |
Other options (environment, read at apply time):
| Env var | Purpose | Default |
|---|---|---|
MIMO_DRIVER |
path to driver/mimo_driver.py |
~/.local/lib/mimo-agent-tools/driver/mimo_driver.py |
MIMO_TMP |
temp dir for spec/response files | /tmp |
Install the driver at the default path (or point MIMO_DRIVER at it):
mkdir -p ~/.local/lib/mimo-agent-tools/driver
cp driver/mimo_driver.py ~/.local/lib/mimo-agent-tools/driver/
Why a python driver?
Multimodal payloads are multi-megabyte base64 strings. Agent shells commonly cap captured stdout (DSH's bash seam caps at 64KB), which silently truncates large payloads. The driver therefore does ALL file reading, body assembly and HTTP POSTing inside one python3 process, using spec/response files on disk — nothing large ever crosses the shell.
Install (DSH bundle)
Standard DSH bundle — install with the official plugin command (auto-inits
the profile, pnpm-installs, and appends the bundle layer per
dsh.bundle.patch):
# From a local checkout, or via git/npm:
dsh plugin --profile web add /path/to/dsh-mimo-agent-tools
# or: dsh plugin --profile web add github:you/dsh-mimo-agent-tools
# Install the python driver to the default path (MIMO_DRIVER points at it):
mkdir -p ~/.local/lib/mimo-agent-tools/driver
cp driver/mimo_driver.py ~/.local/lib/mimo-agent-tools/driver/
# Restart dsh web; the tools mount automatically.
Dependencies are declared as peerDependencies (ecosystem convention —
@deepseek-ai/dsh-tools is already loaded in the DSH process, so nothing is
duplicated). dsh plugin add installs the bundle into the profile's
node_modules where peer deps resolve against the running harness.
Notes on the MiMo API (from the official docs)
- TTS target text goes in the assistant message; the voice description
(voicedesign model) goes in the user message; a
styleinstruction rides the user message for preset voices and becomes an inline(风格)tag prefix for voicedesign voices. mimo-v2.5-tts-voicedesigndoes not accept anaudio.voicefield — it usesoptimize_text_previewinstead.- ASR (
mimo-v2.5-asr) must not receive athinkingfield. input_audio.data/video_url.urlaccept either a public URL or adata:<mime>;base64,...data URL (video base64 capped at 50MB per the docs).- Web search costs per keyword round (
max_keyword, default 3) — see MiMo pricing.force_search(default true) trades freshness against cost;user_locationbiases results, e.g.{"type":"approximate","country":"China","region":"Hubei","city":"Wuhan"}. - Deep thinking (
mimo_think) returnsreasoning_content+content; in multi-turn agent conversations with tool calls,reasoning_contentfrom earlier turns must be echoed back or the API returns 400. - Structured output (
mimo_json) needs an explicit JSON shape description in the prompt (fields, types, nesting); keepmax_completion_tokensgenerous so the JSON is not truncated mid-document.
Tests
python3 tests/test_driver.py # driver request-body assembly (12 cases)
node --test tests/tools.test.mjs # tool registration surface (9 cases)
License
MIT
nexu-io/open-design
Devin-AXIS/iPolloWork
liustack/modlens
ysr666/dsh-vision-router
EthanYoQ/AI-Novel-Writer
Anionex/dsh-vision-toolkit
Lum1104/dsh-browser
fufankeji/deepseek-harness-studio