Figure 1: Overview of BAR. The initial model M M (a dense transformer). For each target domain, a two-expert MoE is created: the anchor expert preserves M M ’s capabilities while the domain expert is trained on new data. Each domain follows its applicable pipeline—math and code use the full pipeline (mid-training → \rightarrow SFT → \rightarrow RLVR), while tool use and safety use SFT only. Shared parameters are progressively unfrozen across stages to minimize divergence between experts. All experts are merged into a single MoE, and a lightweight router is trained on a small sample of SFT data. New experts can later be added (to add a new capability) or swapped in (to upgrade a capability) without retraining previous experts.