Tiered shared KV cache storage for LLM inference. Extends GPU HBM to host memory, NVMe, and JBOF pools — enabling cross-instance prefix reuse for vLLM, Dynamo, and NIXL. - View it on GitHub
Star
6
Rank
2258653