Executive Summary
The maintainers of vLLM, the widely adopted open-source high-throughput inference engine for large language models (LLMs), have issued a critical security advisory addressing CVE-2026-72110. The flaw, assigned a CVSS v3.1 score of 9.4 (Critical), exists within the core PagedAttention memory management subsystem and allows unauthenticated API clients in multi-tenant inference deployments to trigger out-of-bounds GPU memory reads and cross-tenant key-value (KV) cache extraction.
vLLM powers generative AI inference infrastructure across major cloud providers, private enterprise clusters, and frontier AI research labs. In continuous batching setups, where hundreds of parallel user requests are dynamically scheduled into shared GPU memory pages, CVE-2026-72110 breaks the fundamental privacy boundary between concurrent sessions.
Root Cause: Integer Truncation in PagedAttention Block Mapping
PagedAttention partitions the continuous KV cache of an LLM into non-contiguous physical memory blocks on the GPU, mirroring virtual memory paging in operating systems. When an incoming request arrives with dynamic prefix caching enabled, the scheduler maps existing token prefixes to shared blocks.
Security researchers discovered that the CUDA kernel responsible for updating the block mapping table contained an integer truncation vulnerability when processing requests with sequence lengths approaching the maximum context window (e.g., 128k or 256k tokens). A specially constructed sequence length caused a 32-bit offset calculation to wrap around, resulting in a miscalculated GPU block index:
// Vulnerable snippet in PagedAttention CUDA kernel offset calculation
__global__ void paged_attention_kernel(
const int* __restrict__ block_tables,
const int max_num_blocks_per_seq,
const int context_len,
...) {
// VULNERABLE: 16-bit intermediate cast causes overflow for context lengths > 65535
short block_idx = (short)(context_len / BLOCK_SIZE);
// Reads or writes into an unintended physical GPU block table address
int physical_block_number = block_tables[seq_idx * max_num_blocks_per_seq + block_idx];
}
Exploitation Mechanics & Cross-Tenant Data Leakage
In a shared SaaS deployment where enterprise customers query an OpenAI-compatible endpoint served by vLLM, an attacker can exploit the vulnerability through the following sequence:
- Long-Context Sequence Spray: The attacker transmits a crafted prompt padded with repetitive whitespace and specific token prefixes to align context length near the 64k boundary.
- Block Index Wrap-Around: The integer wrap-around forces the CUDA kernel to read from an unallocated or foreign physical block index belonging to another active user.
- Attention Sampling Bleed: During the autoregressive decoding phase, the model samples attention scores over the victim's KV cache tokens, outputting coherent fragments of the victim's conversational history, internal code, or system instructions into the attacker's completion response.
- Worker Denial of Service: If the wrapped block index points to unmapped GPU memory space, the NVIDIA CUDA driver triggers an illegal memory access exception, terminating the entire inference worker pod.
Remediation Guidance for AI Infrastructure Teams
Enterprise platform engineers and MLOps teams hosting vLLM or derivative inference frameworks (e.g., Ray Serve, vLLM-on-Triton) must enact immediate mitigations:
- Upgrade vLLM: Update the
vllmPython package to version0.6.4or higher, which enforces 64-bit block offset indexing across all PagedAttention CUDA and ROCm kernels. - Disable Dynamic Prefix Caching in Multi-Tenant Environments: As a temporary operational safeguard, run vLLM with
--enable-prefix-caching falseuntil clusters are fully patched. - Enforce Worker-Level Isolation: Partition user workloads across dedicated inference pods mapped to specific tenant groups or classification domains rather than sharing single GPU pods across multi-tenant boundaries.
- Implement Request Token Guards: Configure API gateway rate limiting and token boundaries to reject requests that attempt anomalous padding around context boundary thresholds.



