Lightning Weave: Improving the Accuracy-Efficiency Frontier of Reasoning Models through Capability Composition

A core goal of efficient reasoning is to improve the accuracy-efficiency frontier. However, jointly improving reasoning accuracy and inference efficiency can be challenging, as the two objectives can favor different reasoning behaviors. Independently post-trained models already offer distinct strengths in accuracy and efficiency. We introduce Lightning Weave, a post-training framework that extracts and composes these independently learned capabilities in a single student through on-policy distillation. Each acquired capability is represented by the policy shift from the model before post-train

阅读原文(HuggingFace Daily Papers)↗ ← 返回资讯列表