Adversa AI’s Cryptographic Context Injection Turns Encryption Into a Guardrail Blind Spot
Adversa AI researchers found that AES-encrypted prompts can slip malicious instructions past Grok and Gemini's safety guardrails entirely, and xAI still has not fixed the flaw more than two months...
AI safety guardrails work by reading the text a model is about to act on and deciding whether it looks dangerous. Security researchers at Adversa AI found a way to make that entire premise irrelevant: encrypt the dangerous part first. Their new technique, which they call Cryptographic Context Injection, hides malicious instructions inside AES-encrypted ciphertext that a guardrail’s text classifier cannot read, then gets the AI model itself to decrypt those instructions inside its own code execution sandbox, at which point the model treats them as trusted output rather than untrusted input. Adversa demonstrated the technique against both xAI’s Grok and Google’s Gemini, with a zero-click version against Grok capable of quietly exfiltrating a user’s chat history.
Table Of Content
How Cryptographic Context Injection Slips Past Guardrails
Static safety guardrails classify text; they do not execute it. That distinction is the whole vulnerability. An attacker places a block of AES-256-GCM ciphertext on a web page or inside a prompt, along with the key material and an instruction telling the model to decrypt it using its own Python runtime. To a guardrail’s input filter, the ciphertext looks like random noise, so it passes the scan without incident. The model then decrypts the payload itself, and the resulting plaintext instructions arrive inside the model’s trusted execution context rather than as flagged, external input.
Earlier cipher-based jailbreak attempts, using something reversible like base64 encoding, mostly failed because models can often decode those schemes natively from patterns in their own training data, which gives a guardrail a fighting chance of catching the decoded result. Strong encryption removes that shortcut. Recovering AES-256-GCM ciphertext under a PBKDF2-derived key requires actually running the decryption algorithm, which forces the attack through the one part of the system, the code execution sandbox, that guardrails were never built to inspect. Adversa’s own description of the mechanism reaches for a familiar analogy: “The closest classical analogy is SQL injection: a system that fails to distinguish its own trusted query from attacker-supplied data flowing through the same channel.”
A Zero-Click Path Into Grok’s Chat History
The more dangerous of the two demonstrations targets Grok’s web chat agent and its agentic browsing framework at grok.com. The attacker hosts an ordinary-looking web page containing an encrypted JSON object and simple, Python-implementable instructions for decrypting it. A user does nothing unusual: they ask Grok to summarize or analyze the page, and Grok fetches it as part of normal browsing. Once decrypted, the hidden instructions direct the agent to resolve the user’s private session context, their name, coarse location, subscription tier, and the full set of prompts from the ongoing conversation, and embed that data into a URL the agent is told to open in order to “fetch additional context.” Grok’s own privileged, internet-connected navigation tool then loads that URL, transmitting the stolen data to an attacker-controlled server through the request’s query parameters. No click, no warning, no visible sign anything happened.
Part of the attack’s design is misdirection. Rather than asking the model outright to leak data, the payload tells it to generate what looks like a second “decryption key.” That value is not real key material at all: it is a template string built from the user’s own private session data, which the agent then treats as a legitimate parameter rather than as the outbound data leak it actually is. Adversa says it could still reproduce the attack against Grok.com as of August 19, 2026.
Getting Gemini to Break Its Own Rules
Gemini’s public chat interface at gemini.google.com resists the exfiltration half of the attack for a mundane reason: its Python sandbox does not give the model network access to external sites, so there is no privileged tool for a decrypted payload to hijack. Rony Utevsky, Adversa’s lead researcher, told The Register that “the Grok scenario doesn’t work because Gemini doesn’t provide Python with access to external websites,” adding that against Gemini the technique “is useful only to sneak bad questions and answers past guardrails.”
That narrower use is still notable. As SecurityWeek reported, a single prompt instructing Gemini to decrypt a supplied ciphertext was enough, through a series of tricks the researchers describe, to get the model to produce multi-paragraph content its safety filters normally suppress, including instructions related to building an incendiary weapon. The researchers say the technique can also work the bypass in reverse on the way out: telling the model to frame its restricted answer as something it needs to encrypt “for safety,” then hand that still-encrypted material back to the user, using the same cryptographic cover to smuggle output past guardrails that inspect what the model is about to say. Adversa reports that the attack’s success rate against Gemini has fallen sharply since June 2026, though it is not fully closed, and the researchers say they are not sure whether the decline traces to filter updates, model version changes, or both.
A Disclosure Timeline With No Fix Yet
Adversa reported the Grok vulnerability to xAI on June 3, 2026, both directly and through xAI’s HackerOne bug bounty program. According to Utevsky, xAI acknowledged the report but never provided a mitigation timeline. Adversa followed up on August 4 and August 10. As of August 19, the technique still worked against Grok.com. There is no CVE identifier attached to the finding and no user-facing workaround. SpaceX, which acquired xAI earlier this year, did not respond to The Register’s request for comment on the story.
Google was never notified of the Gemini bypass at all. Adversa says it considers jailbreaks, meaning attacks that bypass guardrails to make a model emit content it would otherwise refuse, to be out of scope for Google’s vulnerability disclosure program, so there was no formal channel to report it through.
The Bigger Lesson for Agentic AI Security
Asked whether the technique resembles return-oriented programming, the decades-old binary exploitation trick of chaining together small, already-present snippets of code into a malicious sequence, Utevsky called the comparison close but not exact: ROP attackers can’t inject new code at all and are stuck reusing gadgets already sitting in memory, out of necessity rather than choice. The shape of the failure, he argued, is the same either way. “A static guardrail reads text one artifact at a time. If no single artifact is harmful, they all pass, and the malicious meaning appears only once the runtime assembles them. And guardrails can’t see into the runtime.”
Adversa’s guidance for defenders follows from that same observation: stop trusting artifacts individually and start watching the chain. The company recommends gating any tool call whose arguments derive from content the model itself fetched or decrypted, tagging the provenance of that data as it moves through the system, and alerting on the assembled sequence of actions rather than on any single payload that looked clean in isolation. That is a harder engineering problem than adding another input filter, and it echoes a pattern sxz.io has covered elsewhere this month. Docker’s own analysis of the OpenAI and Hugging Face breach made a similar case that agent security has to be built as a system rather than a stronger version of human review. Tracebit’s earlier research went the other direction and showed the same underlying leverage can cut both ways: its “context bomb” technique plants refusal-triggering text in decoy secrets to turn a model’s own safety training into a kill switch against an attacking agent. Cryptographic Context Injection is the offensive mirror image of that idea, using the same trust a model places in its own guardrail-cleared output, just aimed at getting past the guardrail instead of triggering it. Either way, the lesson is the same: as models gain code execution and privileged tool access, the boundary a guardrail actually protects keeps getting narrower than the one enterprises assume it covers.








No Comment! Be the first one.