[ICLR 2025] SWIFT: On-the-Fly Self-Speculative Decoding for LLM Inference Acceleration - View it on GitHub
Star
0
Rank
14120501