search
KV-cache
Trends
- 1New Self-Pruning Transformer Promises Extreme KV-Cache Compression▼A Self-Pruning Transformer: Extreme KV-Cache Compression w/Universal Attention
A new paper on arXiv introduces a self-pruning transformer architecture claiming extreme KV-cache compression using a universal attention mechanism. KV-cache memory is a major bottleneck for running large language models efficiently, so techniques that shrink it could cut inference costs and enable longer contexts on ordinary hardware. Early reactions are focused on the technical details and whether the compression claims hold up in practice.
- 2Critical unpatched RCE flaw reported in LMCache used by vLLM●🤖 CVE-2026-105192 (CVSS 9.8): unpatched RCE in LMCache, the KV-cache server used by vLLM. Multiprocess mode exposes an u
Security researchers are flagging CVE-2026-105192, a CVSS 9.8 remote code execution vulnerability in LMCache, the KV-cache server used with the vLLM inference framework. In multiprocess mode, an unauthenticated ZeroMQ socket unpickles attacker-controlled data, allowing code execution as root on official images. Versions 0.3.9 through 0.5.5 are affected, and no fix is currently available, raising concern among teams running AI infrastructure.
- 3Unpatched critical flaw allows remote code execution in LMCache●🤖 Unpatched critical RCE in LMCache, the LLM KV-cache server used with vLLM: in multiprocess mode an unauthenticated att
Security researchers have disclosed an unpatched critical remote code execution vulnerability in LMCache, the KV-cache server used alongside the vLLM inference framework. In multiprocess mode, an unauthenticated attacker can execute arbitrary code over ZeroMQ. No fixed version is available yet, leaving deployments exposed until a patch or mitigation is released.