论文

CoX-MoE以CPU与GPU协同执行缓解专家卸载开销

CoX-MoE: Coalesced Expert Execution for High-Throughput MoE Inference with AMX-Enabled CPU-GPU Co-Execution

摘要

MoE完整权重占用较大,专家卸载会受到PCIe传输与碎片化小批次计算的影响。CoX-MoE采用CPU与GPU协同执行,利用CPU的AMX能力,并以普通批次合并专家计算,提高硬件利用率。 系统依据专家使用频率预先划分GPU与CPU执行位置,并按阶段选择性卸载注意力计算。论文与FlexGen、MoE-Lightning比较端到端吞吐,收益来自具体CPU、GPU、批次与模型配置。其核心是执行安排与数据搬运优化,不是用缓存目标替换原模型的专家选择。

MoE推理与不同CPU、GPU协同执行策略的结构示意。
Figure 1. (a) Architecture of Mixture-of-Experts along with the inference flow. (b), (c) Inference flow for each strategy. A block diagram illustrating the Mixture-of-Experts (MoE) architecture, showing tokens being routed by a router to sparse experts, along with the overall inference data flow.