论文

BAR将数学、代码等能力分别训练为可更新的MoE专家

Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts

模型架构模型训练MoE监督微调与指令调优偏好优化强化学习Agentic RL

摘要

BAR将数学、代码、工具使用与安全能力分别训练,再组合成同一个MoE模型。各领域可以采用不同的训练步骤;训练中按阶段释放必要的共享参数,组合时合并共享参数并训练路由器。 在Olmo 2 7B基座实验中,更换代码训练数据或为数学专家增加强化学习,能够把对应能力改善带入组合模型,同时保留其他领域表现。其模块化针对领域能力升级;更换共同基座仍需要重新训练领域专家。模型及方法于4月20日由Ai2发布,论文首次提交的北京时间为4月21日,两个事件分别留档。

BAR独立训练领域专家、合并共享参数与训练路由器的流程。
Figure 1: Overview of BAR. The initial model M M (a dense transformer). For each target domain, a two-expert MoE is created: the anchor expert preserves M M ’s capabilities while the domain expert is trained on new data. Each domain follows its applicable pipeline—math and code use the full pipeline (mid-training → \rightarrow SFT → \rightarrow RLVR), while tool use and safety use SFT only. Shared parameters are progressively unfrozen across stages to minimize divergence between experts. All experts are merged into a single MoE, and a lightweight router is trained on a small sample of SFT data. New experts can later be added (to add a new capability) or swapped in (to upgrade a capability) without retraining previous experts.