What Happened in LLMs Layers when Trained for Fast vs. Slow Thinking: A Gradient Perspective - View it on GitHub
Star
0
Rank
14124007