Tensorwire
Products & tools · first seen 9 Sep, updated 9 Sep

RedKnot-MLA: Multi-Head Offline-Online Reuse for DeepSeek-V4 Long-Context Serving

1 outlet DeepSeek

arXiv:2609.07008v1 Announce Type: new Abstract: Multi-head latent attention (MLA) exposes many logical query heads through one packed latent KV stream. This representation is memory efficient, but it removes the physical per-head cache boun…

Summary from arXiv cs.AI.

Coverage 1 article · 1 outlet

  1. arXiv cs.AI
    RedKnot-MLA: Multi-Head Offline-Online Reuse for DeepSeek-V4 Long-Context Serving