Weak-to-Strong Reasoning
Yuqing Yang, Yan Ma, Pengfei Liu
Proposes a progressive weak-to-strong reasoning framework where a strong model refines its own training data without human or advanced model input, significantly improving reasoning on GSM8K and MATH.