A reference implementation for training math reasoning models with RL on Modal — parallel vLLM inference, LEAN proof verification, and serverless GPU orchestration. - View it on GitHub
Star
1
Rank
6192692