RedKnot: Efficient Long-Context LLM Serving with Head-Aware KV Reuse and SegPagedAttention
arXiv:2606.06256v2 Announce Type: replace Abstract: As the input length of large language model (LLM) serving continues to grow, the KV cache has become a dominant bottleneck in AI infrastructure. It limits GPU memory capacity, serving concurrency, cache reuse, and distributed scalability. Multiple important problems, including position-independent KV cache, prefix KV cache compression, hot/cold KV cache separation, and distributed KV cache management, all depend on how the KV cache is represent
![Hiding messages in the least significant mantissa bits of fine-tuned ONNX model weights [P]](https://external-preview.redd.it/xL20TWLoDXtutsGuMHS1qdEyNEn6zkliHGGNaYV1H4A.png?width=640&crop=smart&auto=webp&s=2b035893689551e28412391b858ab5c0323b052d)


![A debugger for RL reward functions that detects reward hacking during training [P]](https://preview.redd.it/r5m95bf5cn9h1.gif?width=640&crop=smart&s=f9e1900b5e007ea3a72c74d4089c56fdeed22f49)
