Preprint2026
Fastkv: Decoupling of context reduction and kv cache compression for prefill-decoding acceleration
Unknown
FastKV decouples context reduction from KV cache compression to accelerate both prefill and decoding phases in LLM inference.
0Jan 1, 2026