Figure 1: (Left) Emo is an MoE trained with modularity as a first-class objective.
For a given domain (e.g., math, code, biomedical), users can select a small subset of experts of any size and retain near full-model performance. This turns a single model into a composable architecture, enabling flexible deployment with improved memory-accuracy tradeoffs for large, sparse MoEs. (Right) Averaged performance over 16 MMLU categories across different memory budgets. Emo (purple) and Reg. MoE (green) are single models evaluated with expert subsets of different sizes. Emo expert subsets push the Pareto frontier in memory-accuracy trade-off, outperforming standard MoEs and even fixed-budget models trained from scratch.