Minimalistic large language model 3D-parallelism training - View it on GitHub
Star
2457
Rank
15923