TRENDING
Rows of identical brass-colored apartment mailboxes with small locks and name labels along an orange corridor wall
October 9, 2026
How to Prevent Broken Object Level Authorization (IDOR) in a FastAPI App
Street-level upward view of the Monetary Authority of Singapore building and neighbouring office towers under a pale sky
October 9, 2026
Singapore’s AI Guidelines Turn Independent Review Into a Question of Who Sets the Risk Rating
Cast-iron late Qing dynasty coin minting press with a large flywheel, displayed in a museum case
October 9, 2026
Attackers Hijacked the .gh, .sl and .as Country Domains and Minted HTTPS Certificates for Google
Rows of closed oak library card catalog drawers, each with a brass pull and a blank label holder
October 9, 2026
How to Encrypt PII in Python and Keep It Searchable With Blind Indexes
Close-up of a vintage Western Electric manual telephone switchboard with orange lamps, red patch cords plugged into jacks, a rotary dial and a black handset
October 9, 2026
Microsoft’s Agent Lightning v1.0 Turns Agent Training Into a Sample-Accounting Problem
09 Oct 2026
SXZ.io SXZ.io
  • Home
Search the Site
Popular Searches:
Technology Amazon AI
Recent Posts
Two orange safety relief valves on grey pressure vessels in an industrial plant
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
October 8, 2026
Yellow diamond-shaped merging traffic warning sign showing a side road joining a main road
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
October 8, 2026
A lugworm lying on wet sand and mud at low tide
A Compromised Admin Account Put the Shai-Hulud Worm Into AI Sandbox Maker Tensorlake’s npm SDK
October 8, 2026
SXZ.io SXZ.io
  • Home

Categories

Articles 232 Posts
News 234 Posts
Learning Hub 204 Posts
Home/Articles/Docker’s Latest Horror Story Turns a Patched Cursor Bug Into a Sandboxing Argument
Articles

Docker’s Latest Horror Story Turns a Patched Cursor Bug Into a Sandboxing Argument

A patched Cursor vulnerability shows how a shell built-in can quietly rewrite what an approved command does, and Docker is using the case to argue coding agents need execution-layer isolation, not...

August 19, 2026 4 Min Read
42

Docker published the fifth installment of its “AI Coding Agent Horror Stories” series on August 18, and this one is about the safety net most teams lean on when they let an AI coding agent run shell commands with minimal supervision: the command allowlist. The post, written by Docker’s Ajeet Singh Raina, revisits a patched Cursor vulnerability to make the case that an allowlist only ever checks the name of a command, not what that command actually does once it runs, and that gap was enough to defeat it completely.

Table Of Content

  • An approval prompt that didn’t tell the whole story
  • A “systemic issue” with a slow fix
  • Docker’s pitch: contain the execution, not just the approval
  • Why this outlasts one Cursor patch

An approval prompt that didn’t tell the whole story

The vulnerability is CVE-2026-22708, disclosed publicly by Pillar Security researcher Dan Lisichkin on January 14, 2026. It affects Cursor’s Agent when running in Auto-Run Mode with Allowlist mode enabled, a configuration meant to let routine commands through automatically while still prompting a developer for anything unfamiliar. According to GitHub’s security advisory, “certain shell built-ins can still be executed without appearing in the allowlist and without requiring user approval.” Cursor rated the issue High severity and fixed it in version 2.3.

Docker’s post walks through the mechanics with a two-line example. Programs commonly read configuration from environment variables at startup: Git checks one called PAGER to decide how to display its output. The first line quietly sets that variable:

# This one runs silently. You are never asked.
export PAGER="open -a Calculator"

# This one you are asked about, and you say yes, because obviously.
git branch

The export command never shows up on the allowlist and never triggers a prompt, because Cursor’s checker was built to recognize programs on disk, and shell built-ins such as export, typeset, and declare are not files sitting on disk to check. By the time the developer sees and approves the completely ordinary-looking git branch, the environment has already been rewritten. Git looks up PAGER, finds the attacker’s payload sitting in it, and runs that instead. There is no memory corruption and no privilege escalation involved, just a command whose meaning was changed a minute before it was shown to a human for a decision.

Pillar’s research found the technique worked even against a completely empty allowlist, the most restrictive setting Cursor offers. An allowlist can only ever check whether a command’s name appears on a list; it has no way to know what that command will actually do once environment state has already been altered by something the check never saw in the first place.

A “systemic issue” with a slow fix

Pillar’s own account of the disclosure describes reporting the bug to Cursor in August 2025. By September, Pillar’s timeline says, Cursor had acknowledged the report reflected a “systemic issue” with “two major initiatives underway” to address it, though Pillar does not name who at Cursor made that statement. Public disclosure followed on January 14, 2026, and Cursor’s eventual fix requires explicit user approval for any command its server-side parser cannot classify, rather than silently letting an unclassified shell built-in run without one. Pillar’s researchers were direct about what they think the durable fix looks like: “The right way to handle this issue is to allow full command execution to all agents in an isolated or sandbox environment and deprecate the use of allowlists completely.”

Docker’s pitch: contain the execution, not just the approval

Docker is using the case to argue for its own answer to that problem. Docker Sandboxes runs a coding agent’s shell inside an isolated microVM rather than directly on a developer’s machine, so the operative question stops being “was this specific command approved” and becomes “what can this sandboxed process reach at all,” regardless of which built-in quietly changed which variable. Docker’s sbx command-line tool exposes that boundary directly: sbx policy ls shows what a given sandbox can currently reach, and a pair of commands like sbx policy deny network "**" followed by sbx policy allow network "github.com,registry.npmjs.org" narrows that down to only what a specific task needs.

That local control has a ceiling by design. When an organization has governance turned on, its policy replaces whatever a developer sets locally, visible as a “Governance: Managed by <org>” line in the sbx policy ls output, and a local allow rule can never override an organization-level deny. Docker also describes “kits” that let a team define network and credential rules once and hand the same artifact to every developer instead of maintaining a personal allowlist per laptop, plus audit events, carrying the user, timestamp, and the rule that fired, that stream into whatever SIEM a security team already runs whenever a sandboxed process reaches for something outside its policy.

Among the practices Docker’s post recommends: treat export and other environment-changing built-ins as commands in their own right, since they can silently change what a later, already-approved command does; do not mistake an allowlist for a security boundary, since, in Docker’s characterization, Cursor’s own documentation now describes it as best-effort rather than a guarantee; and isolate untrusted code, meaning anything a developer did not write and has not read, before the first command runs rather than after something looks wrong.

Why this outlasts one Cursor patch

This is the fifth entry in Docker’s horror-stories series, following posts on an agent running rm -rf ~/, the same failure mode reaching into a production cloud environment, and, in Part 4, credentials harvested through the Nx npm supply chain attack. The throughline across all five is that a coding agent runs with the developer’s own filesystem permissions and credentials, and an approval prompt only protects against threats the prompt itself can see.

CVE-2026-22708 is specific to Cursor and already patched. But the assumption it broke, that a command’s name at approval time reliably describes what will execute, is a property of approval-based guardrails in general, not of any single vendor’s implementation. Any agentic coding tool that lets a model run shell commands past a human review step inherits the same open question: what else, besides the command a developer actually sees, gets a vote in deciding what that command does?

Tags:

AI AgentsApplication SecurityCursorDockerPrompt Injection

Share

Close-up photograph of an archaic stone Gorgon relief sculpture from a Greek temple pediment, flanked by carved lion heads
Previous Post

CISA, FBI, and HHS Warn Medusa Ransomware Has Hit Over 500 Critical Infrastructure Orgs

Close-up of a refreshable Braille display connected beneath a computer keyboard, showing raised dot patterns used by blind screen reader users
Next Post

How to Find and Fix WCAG 2.2 Accessibility Bugs With axe-core and Playwright

No Comment! Be the first one.

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Latest
08 Oct
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
08 Oct
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
Trending
October 8, 2026
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
October 8, 2026
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
October 8, 2026
A Compromised Admin Account Put the Shai-Hulud Worm Into AI Sandbox Maker Tensorlake’s npm SDK
October 8, 2026
How to Prevent Broken Object Level Authorization (IDOR) in a FastAPI App
October 8, 2026
Singapore’s AI Guidelines Turn Independent Review Into a Question of Who Sets the Risk Rating
October 8, 2026
Attackers Hijacked the .gh, .sl and .as Country Domains and Minted HTTPS Certificates for Google

Related Posts

Blue-lit server racks in a modern data center, illustrating the compute infrastructure behind the AI boom.
Articles

The AI Boom Is Spending Real Money Before Proving Real Returns

June 7, 2026
Technician working with a laptop beside server racks, representing enterprise AI retrieval infrastructure
Articles

Google’s Agentic RAG Push Makes Enterprise AI Less of a One-Shot Guess

June 7, 2026
A person with a laptop and smartphone, representing digital attention and AI-assisted work
Articles

AI Chatbots Are Making Attention a Design Problem

June 7, 2026
A customer-support representative wearing a headset against a dark studio background.
Articles

The Meta AI Support Hack Was a Plain Old Authorization Failure

June 7, 2026
SXZ.io SXZ.io
  • [email protected]

Categories

Articles
Learning Hub
News

All Rights Reserved by SXZ.io ©2026