Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine — pure C, zero deps, experts streamed from disk. Radical LLM cost reduction: no cloud API, no GPU farm. Tiny engine, immense model. 🐦 - View it on GitHub
Star
2
Rank
4336756