A very simple GRPO implement for reproducing r1-like LLM thinking. - View it on GitHub
Star
1702
Rank
26053