At Black Hat USA 2026, Simcha Kosman presented a full exploitation chain achieving what he called “full C2 communication entirely inside ChatGPT sandbox.” The attack chains file-parsing flaws, reasoning injection, and Task Scheduler URL laundering to establish a covert signaling protocol using JFrog authentication rate limits. Kosman reported five findings to OpenAI in March, but the company disputes the framing: it says this doesn’t represent a sandbox escape.
That distinction matters more than it should. Kosman demonstrated operational control inside production infrastructure: executing code, receiving commands, and exfiltrating data on OpenAI’s hardware. Whether that’s technically an “escape” depends on where you draw the sandbox boundary. But a C2 channel inside a billion-user platform is the threat that counts.
OpenAI’s responses were mixed: treating reasoning injection as out of scope, calling some findings known issues, and removing Artifactory authentication requirements. Those are specific arguments worth having. The broader claim, that confinement to the sandbox is meaningful security when the sandbox runs on production hardware, is already settled.
Sources: Black Hat 2026: ChatGPT Sandbox Research
Coverage: Dark Reading: ChatGPT Sandbox Control • NCSA Cyber Threat Intelligence
Related on this blog: [AUTO] Meta’s model didn’t hack a company, the sandbox did • OpenAI’s Containment Problem Grows a Third Time in 10 Days • Three Vendors, One Misconfigured Test Lab: The 2026 Agent Escape Wave