Google Cloud has resolved a critical input sanitization and classifier evasion flaw in the safety inspection pipeline of Vertex AI Model Armor. The defect allowed attackers to bypass automated prompt injection defenses using zero-width Unicode permutations and token-splitting structures, enabling indirect prompt injection attacks against enterprise Retrieval-Augmented Generation (RAG) knowledge stores.
Vulnerability Dynamics: Unicode Token Obfuscation (CWE-20)
Vertex AI Prompt Armor operates as a pre-inference defense gateway designed to detect adversarial jailbreaks, jailbreak heuristics, and indirect prompt injections before user queries reach underlying Gemini models. The detector uses a dedicated fine-tuned neural classifier alongside heuristic pattern matchers.
Researchers discovered that when user input documents embedded Cyrillic homoglyphs interspersed with directional formatting control characters (such as Right-to-Left Override U+202E), the pre-processing tokenizer stripped the non-printable bytes during classification but reconstructed the full UTF-8 payload when piping data into the model inference context:
// Tokenizer Discrepancy Flow in Vertex AI Gateway
[Raw Ingress Query] -> [Pre-Tokenizer: Strips Unicode Control Chars] -> [Classifier: Clean / Benign]
|
+--> [Model Inference Gateway: Preserves Raw Bytes] -> [LLM Executes Injected System Override]Exploitation Blast Radius in Autonomous Agent Pipelines
In enterprise customer service and internal document synthesis deployments, poisoned external PDFs or web scrapes could coerce the autonomous agent into executing unprompted tool actions. Telemetry demonstrated that attackers could execute unauthorized document summaries, query enterprise BigQuery data sets, and exfiltrate customer account records through outgoing webhook parameters.
Hardening Recommendations for Cloud Architects
- Implement Dual-Layer Character Normalization: Enforce strict Unicode NFKC normalization at the API gateway layer prior to transmitting payloads to any cloud safety filter.
- Enforce Strict RAG Read Boundaries: Restrict vector database service accounts to read-only scopes limited strictly to the authenticated user's authorization tier.
- Deploy VPC Service Controls: Isolate Vertex AI Model Armor endpoints within protected perimeter zones blocking unauthorized outbound internet access from inference containers.



