xiaohou521/fusion-moa

Plugin插件 ⭐ 2 Apache-2.0 Models & Routing模型与路由

Fusion MoA: model- and GPU-independent Mixture-of-Agents runtime for coding agents

Project Overview项目介绍

Fusion MoA is an open-source, model- and GPU-independent Mixture-of-Agents runtime for coding agents. It lets users connect to any model endpoints they already use, select an orchestration policy, and expose a single OpenAI- or Anthropic-compatible API endpoint that works with multiple coding agents. Supported coding agents include Claude Code, Codex, OpenCode, and DeepSeek Harness, so it does not tie users to a single agent platform. Its core architecture keeps the main model as the only public writer, while independent expert models only provide private corrections to the main model’s output. The recommended expert-constrained policy requires a valid independent review from experts before the main model generates its final output.

This tool is built for developers who want to combine multiple LLM capabilities to improve performance on coding tasks. To get started, after installing the runtime, you copy a pre-built YAML recipe that matches your desired orchestration policy. You then edit the YAML file to add your main model and expert model endpoint URLs and credential references, validate the configuration with a built-in check command, and start the local runtime service. Finally, you configure your coding agent to connect to the Fusion MoA endpoint as a custom OpenAI- or Anthropic-compatible model, and you can start using the blended model capability right away.

Fusion MoA requires Python 3.11 or newer to run, and you need at least one working model endpoint to use it. Any local service like vLLM, SGLang, or llama.cpp, or any cloud endpoint with an OpenAI-compatible Chat Completions API works for a first run. Since Fusion MoA is still early-stage software, expert orchestration can cost more or perform worse than direct inference on some workloads. The project is released under the permissive Apache 2.0 open-source license, and it follows the security practice of storing all credentials in environment variables, never in configuration files or code. You should always validate the complete expert path on your own coding tasks before using it in production.

Fusion MoA 是一个开源、与具体模型和GPU无关的混合智能体(Mixture-of-Agents)运行时。它支持接入用户已有的多个模型端点,选择编排规则后,对外暴露一个兼容OpenAI或Anthropic格式的统一模型API,可供给Claude Code、Codex、OpenCode、DeepSeek Harness等各类编码智能体调用。它的核心设计是主模型负责公开输出,独立专家模型仅提供私有修正,默认推荐的专家约束策略要求最终生成前必须经过独立专家评审。

适合需要整合多个模型能力提升编码任务质量的开发者使用。典型工作流是安装后复制预设的YAML配方,配置主模型和专家模型的端点地址与凭证,验证配置后启动运行时服务,最后在编码智能体中配置自定义的兼容OpenAI或Anthropic的模型来源,即可使用融合后的模型能力。用户可根据需求选择不同编排策略,包括直接输出、推理预留、评论家评审、自适应自审等多种选项。

需要Python 3.11及以上版本,至少配置一个可用的模型端点,本地vLLM、SGLang、llama.cpp或任何兼容OpenAI接口的服务都可使用。项目目前处于早期开发阶段,专家编排在部分任务上可能成本更高或效果不如直接推理,建议生产使用前先在自有任务上验证完整路径。项目采用Apache 2.0开源许可证,凭证一律存放在环境变量,不写入配置文件或代码。

Pre-install check安装前体检Compatibility · Security兼容性 · 安全性 2 warnings2 项注意
  • Only 2 stars - very few users, little community feedback星标只有 2,几乎没人在用,遇到问题缺少社区反馈
  • No DSH plugin manifest detected - it may only carry the dsh-plugin topic, so the install method must be confirmed on the spot未检测到 DSH 插件清单:可能只是打了 dsh-plugin 话题,安装方式要现场确认
  • Not DSH-native: a multi-platform tool that may require Node / Electron or another runtime first非 DSH 原生,是多平台兼容工具:可能要先装 Node / Electron 等运行时
DSH walks through these 9 checksDSH 会逐条核对这 9 项

Compatibility兼容性

  • DSH, Node, OS and profile requirementsDSH 版本 / Node 版本 / 操作系统 / profile 是否满足要求
  • External dependencies and runtimes (Electron / Python / Docker, ...)外部依赖与运行时(Electron / Python / Docker 等)是否齐备
  • Conflicts with installed plugins: command names, skill / tool names, ports, duplicate MCP registration与已装插件是否冲突:命令名、skill / tool 重名、端口占用、重复 MCP 注册

Security安全性

  • Repo matches the facts registered here; archived or abandoned?仓库是否与页面登记一致,是否归档或长期停更
  • Safety of preinstall / install / postinstall and install.sh / setup.ps1preinstall / install / postinstall 与 install.sh、setup.ps1 是否安全
  • curl|bash, download-then-execute, obfuscation, unrelated domains → stop immediatelycurl|bash、下载即执行、混淆代码、无关域名 → 立刻停止
  • Typosquatting or unmaintained packages among the new dependencies新增依赖里有没有 typosquatting 或无人维护的包
  • Requested permissions vs. what the feature actually needs申请了哪些权限、是否超出功能所需(filesystem / network / shell / clipboard)
  • Any sudo / admin requirement, plus uninstall and rollback是否要求 sudo / 管理员权限,以及卸载与回滚方式

Anything uncertain must be marked unknown with a note on how to confirm it. This site's signal screen is a static snapshot, not a security audit.拿不准的必须标「未知」并说明要我怎么确认。本站的信号筛查是静态快照,不能替代安全审计。

Or use CLI install (for developers)或使用命令行安装(适合开发者)

CLI Install命令行安装

dsh plugin --profile web add github:xiaohou521/fusion-moa

把 xiaohou521/fusion-moa 加入你的 DSH 配置(web profile)即可启用。

READMEREADME

Fusion MoA

简体中文

One model API for coding agents, backed by your main model and independent expert models.

Fusion MoA is an open, model- and GPU-independent Mixture-of-Agents runtime. You connect the model endpoints you already use, choose an orchestration recipe, and expose one OpenAI- or Anthropic-compatible model to Claude Code, Codex, OpenCode, DeepSeek Harness, or another coding agent.

Coding agent
    │  OpenAI Chat / Responses / Anthropic Messages
    ▼
Fusion MoA  ── policy, budgets, fallback, accounting
    │
    ├── authoritative main model ──► native final stream
    └── independent read-only experts ─► private corrections only

The main model is always the only public writer. Experts receive no coding tools, their output is bounded and treated as untrusted, and the recommended expert-constrained policy requires a valid independent review before final generation. Fusion MoA does not require a particular model family, inference server, cloud, GPU, or coding-agent harness.

What works today

The current community release provides:

  • one fusion/v1 YAML recipe for providers, models, pools, policies, completion rules, and serving;
  • OpenAI-compatible and Anthropic-compatible upstream providers;
  • OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages client endpoints;
  • native streaming from the authoritative final model, including incremental tool-call arguments;
  • portable function/tool-call round trips and /v1/models discovery;
  • direct, reasoning-reserve, critic, review-board, adaptive self-review, and mandatory expert-constrained policies;
  • explicit capability checks for thinking controls, tools, and schema-constrained output;
  • bounded fallback, usage aggregation, and completeness flags across all model calls;
  • Python entry points for third-party provider and policy plugins;
  • a version-pinned DeepSeek Harness integration.

Fusion MoA is early-stage software. Expert orchestration can cost more or perform worse than direct inference on some workloads. Keep direct as an experimental control, and validate the complete expert path on your own coding tasks before using it in production.

Showing the opening section of the README — the full document lives in the repository以上为 README 开头摘要,完整文档在仓库内 · View the full README on GitHub →在 GitHub 查看完整 README →

← 上一个 Prev dsh-mcp-workspace-scope 下一个 Next dsh-chatgpt-codex →