A critical memory corruption vulnerability cataloged as CVE-2026-103764 (CVSS 9.8, GHSA-mp8g-r7h2-5x36) has been discovered in the Mooncake transfer engine, the open-source distributed Key-Value Cache (KVCache) transfer architecture engineered to power disaggregated large language model (LLM) inference clusters. An untrusted pointer dereference flaw (CWE-822) in the engine's TCP transport layer allows unauthenticated remote attackers to execute arbitrary memory read and write operations across GPU cluster host memory, enabling LLM prompt cache theft and arbitrary code execution on inference server nodes.

Disaggregated KVCache in Frontier LLM Serving

In modern generative AI inference architectures (such as those serving LLaMA-3, DeepSeek, and Mistral models), processing is split into two distinct operational phases: the prefill phase (computation-heavy prompt ingestion) and the decoding phase (memory-bandwidth-bound token generation).

To prevent decoding GPU nodes from stalling while waiting for prefill calculations, distributed serving engines employ Mooncake. Mooncake coordinates a unified, high-speed memory pool spanning host DRAM, local NVMe SSDs, and remote RDMA/TCP memory pools, transferring gigabytes of intermediate transformer attention matrices (KV caches) between prefill and decode instances in sub-millisecond latencies.

Vulnerability Mechanics: SessionHeader Pointer Dereference

When Mooncake nodes communicate over the TCP transport layer, incoming request packets are parsed by the ServerSession::readHeader function in mooncake-transfer-engine/src/transport/tcp_transport.cpp.

The protocol format expects a binary SessionHeader containing operation metadata, buffer lengths, and memory segment addresses. However, prior to version 0.3.13, the parser directly trusted 64-bit client-supplied memory address pointers without verifying that the target address fell within registered Mooncake segment allocation descriptors:

// Vulnerable pointer extraction in ServerSession::readHeader
void ServerSession::readHeader(const char* buffer, size_t len) {
    auto header = reinterpret_cast<const SessionHeader*>(buffer);
    
    // Attacker supplies arbitrary 64-bit pointer and length
    uintptr_t target_addr = header->remote_addr;
    size_t transfer_size = header->payload_size;
    
    if (header->opcode == OpCode::READ) {
        // Direct memory dereference without bounds verification!
        void* src_ptr = reinterpret_cast<void*>(target_addr);
        sendResponse(src_ptr, transfer_size);
    } else if (header->opcode == OpCode::WRITE) {
        void* dst_ptr = reinterpret_cast<void*>(target_addr);
        readPayload(dst_ptr, transfer_size);
    }
}

Because the Mooncake server daemon runs with unconstrained access to host memory, an attacker connecting to the raw TCP port (default port 13801) can dispatch crafted SessionHeader packets to:

  1. Dump Entire KVCache Tensors: Exfiltrate private user conversational prompts, confidential corporate documents ingested via RAG (Retrieval-Augmented Generation), and system instructions from co-located tenant sessions.
  2. Snoop Host Secrets: Read application memory containing cloud IAM credentials, model weights, and API authentication tokens.
  3. Corrupt Function Pointers: Overwrite C++ virtual method tables (vtables) or stack return addresses using the WRITE opcode, hijacking the execution thread to execute arbitrary shellcode with the privileges of the GPU orchestrator daemon.
Operation Code Attacker-Controlled Parameters Exploitation Capability
OpCode::READ remote_addr (64-bit), payload_size Arbitrary kernel/userland memory read (Information Disclosure)
OpCode::WRITE remote_addr (64-bit), payload bytes Arbitrary memory write, vtable overwrite (Remote Code Execution)

Defensive Remediation Guidelines for AI Infrastructure Engineers

  • Upgrade Mooncake Immediately: Deploy Mooncake version 0.3.13 or higher. The patch introduces strict pointer canonicalization against verified memory segment tables in SegmentManager::validate_buffer_range().
  • Isolate KVCache Data Ports: Ensure Mooncake TCP and RDMA transport ports (ports 13800–13850) are strictly confined to isolated cluster backend networks (e.g. dedicated Kubernetes overlay networks or AWS VPC private subnets) and never exposed to the public internet or untrusted tenant networks.
  • Implement Mutual TLS on Transfer Transports: Configure Mooncake clusters with mTLS authentication, preventing rogue containers on the same host node from establishing unauthenticated TCP sessions to the memory daemon.