Hugging Face Blog·· 2026-08-10AI 评分51
Multiverse Computing 提出低成本 LLM 知识蒸馏方法:离线 top-K logits 与融合分块 KL 损失
Making Knowledge Distillation Cheap Enough to Run at Scale
AI 导读
Multiverse Computing 发布论文《Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss》,通过缓存教师模型每位置的 top-100 logits 做离线蒸馏,并把输出投影融合进分块 KL 损失计算,使峰值显存只随序列长度线性增长。
来源:Hugging Face Blog · huggingface.co