Extreme weight + KV cache compression for LLMs on Apple Silicon (MLX implementation of Google's TurboQuant) - View it on GitHub
Star
0
Rank
14358822