论文

EMO通过预训练形成可按任务保留的专家组

EMO: Pretraining Mixture of Experts for Emergent Modularity

摘要

普通MoE虽然每个词元只使用少数专家,但通常仍需保存完整专家权重;直接删除大量专家会破坏模型能力。EMO在预训练中让同一文档的词元共享一个候选专家池,同时允许不同文档使用不同专家池,使专家组按领域形成分工。 研究训练约1B激活参数、14B总参数的模型,并对数学、代码等领域选择不同大小的专家子集。保留25%和12.5%的专家时,对应评测绝对分数下降约1和3个百分点。这里的比例指部署保留的专家池,而不是每个词元激活参数比例;效果针对这一专门训练的模型。

EMO按任务选择专家子集及其部署评测流程。
Figure 1: (Left) Emo is an MoE trained with modularity as a first-class objective. For a given domain (e.g., math, code, biomedical), users can select a small subset of experts of any size and retain near full-model performance. This turns a single model into a composable architecture, enabling flexible deployment with improved memory-accuracy tradeoffs for large, sparse MoEs. (Right) Averaged performance over 16 MMLU categories across different memory budgets. Emo (purple) and Reg. MoE (green) are single models evaluated with expert subsets of different sizes. Emo expert subsets push the Pareto frontier in memory-accuracy trade-off, outperforming standard MoEs and even fixed-budget models trained from scratch.