Products & tools · first seen 9 Sep, updated 9 Sep
RedKnot-MLA: Multi-Head Offline-Online Reuse for DeepSeek-V4 Long-Context Serving
arXiv:2609.07008v1 Announce Type: new Abstract: Multi-head latent attention (MLA) exposes many logical query heads through one packed latent KV stream. This representation is memory efficient, but it removes the physical per-head cache boun…
Summary from arXiv cs.AI.