Check Point’s Black Hat Research Turns AI Agent Frameworks Into the Real Attack Surface
Check Point researchers spent a year breaking LangChain, LangGraph, CrewAI, AutoGen, Microsoft Agent Framework, and Google ADK, and found the real vulnerability sits in the framework code, not the...
Check Point researchers spent a year trying to break the software that enterprises use to build AI agents, and the frameworks lost. In a Black Hat USA 2026 briefing on Wednesday, August 5, researchers Yarden Porat and Shahar Tal disclosed 11 vulnerabilities across six of the most widely used agent frameworks: LangChain, LangGraph, CrewAI, AutoGen, Microsoft Agent Framework, and Google’s Agent Development Kit (ADK). Their conclusion cuts against the industry’s current obsession with prompt injection defenses: the prompt is just the delivery mechanism. The actual bug lives in the framework code underneath it.
Table Of Content
“Our research shows a deeper failure: in many agentic frameworks, prompt-controlled content can cross the boundary into trusted framework logic itself,” Porat and Tal wrote in the research write-up accompanying their talk, titled “No Tools Required: Post-Injection Exploitation Across AI Agent Frameworks,” which other outlets covering Black Hat’s opening day also flagged as a reframing of agent security around the framework layer rather than the tools an agent can call.
Old Bug Classes, New Blast Radius
What makes the finding notable isn’t novelty. It’s the opposite. “Almost none of it was a completely new bug class,” Tal told The Register, which first reported the disclosure. “That’s insecure deserialization, server-side request forgeries, path traversals, use-after-free. These are bugs that we learned to fix 20 years ago, and they’re sitting underneath agents that now read your inbox, or update your database.”
The pattern the pair kept finding: frameworks fail to keep attacker-controlled content confined to the data plane, letting it leak into orchestration logic, memory, state handling, routing, and system instructions that are supposed to be trusted. An attacker doesn’t need the agent to have dangerous tool access at all. “The agent needs no dangerous tools to be turned against you: reading the wrong document is enough,” Tal said. “We’re building this layer faster than we know how to defend it.”
How a Chat Message Became a Shell on Microsoft’s Agent Framework
The most severe individual finding was a checkpoint deserialization bug in Microsoft Agent Framework that led to remote code execution. Agent checkpoints save an agent’s state, including conversation history, to persistent storage so a session can resume or roll back after an error. Check Point found that the framework would deserialize that stored data without verifying it hadn’t been tampered with, and that prompt injection could plant tampered data in the first place.
Tal described the resulting attack chain in concrete terms: “One person’s message plants the payload, and then a different person rewinds their own session, which triggers the payload, and now the attacker has a shell on that server.”
Microsoft credited the researchers and paid a $10,000 bug bounty. Because Agent Framework wasn’t yet a generally available product when Check Point reported the bug, Microsoft didn’t assign it a CVE. A Microsoft spokesperson told The Register: “We have released protections to harden the Agent Framework and prevent the concrete exploitation path demonstrated in the proof of concept. In addition, we updated the specific checkpoint file with additional language to define the security boundary.”
Google’s ADK Took a Different Path
Check Point also found flaws in Google’s ADK, and says Google’s response was less complete: no full fix, and no CVE. Porat described a built-in development assistant that ships with ADK, one capable of writing files to disk, that “stays reachable over the HTTP API even though it is hidden from the app listing.” Because that API carries no authentication by default, an attacker who opens a session can ask the assistant to write a malicious Python file, then ask the server to run it, crossing from prompt-controlled chat into arbitrary code execution on import.
Check Point’s Second Look at the Same Layer
This wasn’t Check Point’s first pass at framework internals this year. In a June 2026 disclosure covering LangGraph specifically, Check Point’s Porat reported three vulnerabilities built on the same checkpoint mechanism. Two were chainable into remote code execution on self-hosted deployments: CVE-2025-67644, a SQL injection flaw in LangGraph’s SQLite checkpoint implementation (CVSS 7.3, fixed in langgraph-checkpoint-sqlite 3.0.1), combined with CVE-2026-28277, an unsafe msgpack deserialization bug that could trigger object reconstruction from a tampered checkpoint (CVSS 6.8, fixed in LangGraph 1.0.10), by exploiting LangGraph’s get_state_history() endpoint to retrieve and then poison stored checkpoints. A third, separate flaw, CVE-2026-27022, was a RediSearch query-injection issue in the Redis checkpointer (CVSS 6.5, fixed in version 1.0.1). The Hacker News, which covered that disclosure, put LangGraph’s adoption at roughly 46.5 million monthly downloads, a reminder of how much production software sits on top of a single checkpointer implementation.
Taken together, the June LangGraph chain and the six-framework Black Hat findings point at the same architectural seam: agent frameworks persist state (conversation history, task progress, memory) so sessions can pause and resume, and that persistence layer keeps trusting data it should be verifying.
What This Changes for Teams Building on These Frameworks
Porat and Tal’s framing is a deliberate correction, not just a disclosure. Prompt injection defenses filter what a model is told; they don’t touch what a framework does with data once an agent has already been influenced. Check Point’s advice, echoed across both disclosures, is to treat prompt injection as a given rather than a solved problem, and to audit the framework layer itself: how checkpoints are deserialized, whether internal development or debugging APIs are reachable without authentication, and whether memory and state stores validate their inputs the way any other persistence layer handling untrusted data would have to. “There’s a lot of research going into prompt injection and defenses, which are important,” Tal said, “but that’s just the beginning.”








No Comment! Be the first one.