论文
使用交错堆叠的快速语音基础模型 蒸馏
Fast Speech Foundation Model Distillation Using Interleaved Stacking
摘要
将大型语音基础模型(SFM)提炼为高效的学生模型已成功应用于资源匮乏的环境。虽然 蒸馏 减少了推理延迟,但它需要额外的学生 模型训练。然而,SFM 蒸馏 的训练效率仍有待探索。在这项工作中,我们探索 SFM 蒸馏 的训练加速,以加快模型部署。 We examine the potential of stacking, in which the model depth is progressively increased through training until the target model depth is reached.虽然现有的堆叠方法提高了训练速度,但它们的性能却下降了。 To handle this limitation, we propose interleaved stacking, a novel stacking method that consistently preserves layer position throughout the stacking process.这一属性在 SFM 中尤其重要,其中每一层都编码不同的特定于层的知识。我们在 SUPERB 上验证了所提出方法的有效性。