NVIDIA Product Security has issued an emergency bulletin addressing a critical memory corruption vulnerability (CVE-2026-34902, CVSS 9.0) in NVIDIA Triton Inference Server. The defect, categorized under CWE-120: Buffer Copy without Checking Size of Input, allows remote attackers to trigger arbitrary code execution within host inference environments, posing acute risks to multi-tenant cloud AI clusters and enterprise generative model hosting pipelines.
Vulnerability Dynamics: Shared Memory Tensor Parsing (CWE-120)
NVIDIA Triton Inference Server serves as the standard multi-model execution backbone across enterprise cloud environments, serving deep learning models on GPUs via HTTP/REST, gRPC, and CUDA shared memory. To minimize latency, Triton supports System Shared Memory (shm) and CUDA IPC shared memory for exchanging high-volume tensor arrays between client processes and server workers.
Researchers identified an integer calculation flaw within the gRPC shared memory registration endpoint (RegisterSystemSharedMemory). When a remote caller submits an inference request declaring tensor dimensions with negative strides or arithmetic overflows in byte size calculations, the shared memory allocator miscomputes the required buffer span:
// Vulnerable tensor size calculation in Triton backend handler
size_t calculate_tensor_byte_size(const Dims& dims, DataType dtype) {
int64_t element_count = 1;
for (int i = 0; i < dims.nbDims; ++i) {
element_count *= dims.d[i]; // Defect: 32-bit signed overflow allows undersized buffer allocation
}
return element_count * GetDataTypeByteSize(dtype);
}Exploitation Blast Radius in Production AI Environments
By exploiting this flaw, unauthenticated callers with network access to Triton's inference port (TCP 8001 for gRPC or 8000 for HTTP) can overwrite adjacent memory pages on the host machine. In Kubernetes clusters utilizing NVIDIA GPU Operator, this enables container escapes, extraction of neighboring model weights, and exfiltration of confidential inference prompt tokens from other tenants.
Enterprise AI Infrastructure Remediation
- Upgrade Triton Inference Server: Immediately deploy NVIDIA Triton container release 24.10 (Triton core version 2.50.0) or later.
- Restrict Shared Memory Endpoints: Disable CUDA and System Shared Memory endpoints on internet-exposed Triton instances by setting
--allow-shm=false. - Enforce Kubernetes Network Policies: Isolate Triton inference worker pods within private backend subnets accessible only via authenticated API gateway proxies.



