Executive Summary
A high-severity vulnerability identified as CVE-2026-61120 has been uncovered in the widely deployed open-source large language model (LLM) serving engine vLLM. The flaw resides within the distributed memory management protocol responsible for transferring PagedAttention Key-Value (KV) cache pages across multi-GPU compute nodes, allowing unauthenticated network actors within the inference cluster VPC to trigger arbitrary remote code execution on worker pods and extract proprietary model parameters directly from GPU memory.
Vulnerability Mechanics & Protocol Flow
vLLM achieves high-throughput inference performance by utilizing dynamic PagedAttention algorithms that split contiguous KV-cache states into non-contiguous virtual memory blocks. In distributed deployments leveraging pipeline parallelism (PP) or tensor parallelism (TP) across multiple worker hosts, cache pages are synchronized using an internal remote procedure call (RPC) endpoint.
During the deserialization of cross-node KV-cache metadata packets in versions prior to 0.6.4, the tensor transport receiver relied on an unsafe Python object serialization parser. Because authentication was disabled by default on internal worker ports, an adversary who gained initial access to the internal cluster subnet or an adjacent container could dispatch malicious serialized byte arrays directly to the worker service port.
# Vulnerable vLLM inter-worker RPC listener binding (default unauthenticated)
python3 -m vllm.entrypoints.openai.api_server --model meta-llama/Llama-3-70B --tensor-parallel-size 4 --distributed-executor-backend ray --worker-port-range 29500-29510
Upon receipt, the runtime invokes deserialization before verifying cryptographic checksums, executing arbitrary system instructions within the security context of the inference runner (often running as root with direct access to NVIDIA host drivers and IPC shared memory).
Threat Modeling & Impact Analysis
The impact of CVE-2026-61120 is particularly severe for enterprise AI providers hosting proprietary fine-tuned weights or multi-tenant inference APIs. Once remote execution is achieved inside a worker container, the attacker can:
- Dump Proprietary Weights: Read host VRAM buffers via CUDA device memory mappings, extracting proprietary model checkpoints prior to disk encryption.
- Eavesdrop on Inference Prompts: Intercept plain-text user prompts, system instructions, and generated tokens residing unencrypted in KV-cache pages.
- Subvert Model Outputs: Inject subtle mathematical noise or poisoned tensor weights into running batches to manipulate generative responses dynamically.
Defensive Playbook & Mitigation Guidelines
| Component | Vulnerable Version | Remediated Version | Mitigation Action |
|---|---|---|---|
| vLLM Core Inference Engine | < 0.6.4 | 0.6.4 or later | Upgrade pip package; enforce safe tensor serializers |
| Ray Cluster Backend | All unsegmented | mTLS enforced | Enable Ray TLS authentication on cluster communications |
| Kubernetes NetworkPolicy | Permissive ingress | Strict Pod Ingress | Block worker port access (29500-29510) from non-head pods |
Security engineering teams deploying vLLM on Kubernetes or bare-metal GPU clusters must immediately upgrade to vLLM 0.6.4 or apply strict network isolation policies preventing unauthorized connections to distributed worker sockets.



