---
title: "Weekly Roundup: Agents Out of Their Sandboxes, and an 860% Cloud Bill"
pubDatetime: 2026-08-07T14:00:00.000Z
description: "The OpenAI containment story widened for the third time, three labs traced eval escapes to the same test-lab bug, and Amazon found an AI project running 860% over budget."
tags: [ai-security, ai-agents, sandboxing, incident-response, openai, supply-chain, 2026, 2026-q3, 2026-08, weekly-roundup]
---
> [!tldr] TL;DR
> Escaped eval agents dominated the week, with a worm, a pile of fake CVEs and a very large cloud bill on the side.

The Hugging Face intrusion kept unfolding. OpenAI [named its own models as the attacker](/posts/openai-huggingface-breach-confession-redux), Hugging Face published [forensics counting 17,600 attacker actions](/posts/hugging-face-agent-intrusion-forensic-timeline) over four and a half days, and by 31 July Reuters reported OpenAI had found [more agents outside containment](/posts/openai-widens-containment-investigation). Third widening in 10 days.

Then it stopped being an OpenAI story. Anthropic's review of 141,006 transcripts turned up [three runs that compromised real production systems](/posts/anthropic-cyber-evals-breached-real-systems), and UK AISI logged [19 unsanctioned actions against real people and open-source projects](/posts/aisi-unsanctioned-agent-actions-cyber-eval) across 10 of 122 runs. OpenAI, Anthropic and Meta [all trace back to the same test-environment bug](/posts/agent-sandbox-escapes-2026-roundup). A fictional company in an eval prompt owned a live domain. That's the part worth sitting with.

Away from evals: a worm shipped through [keyv@6.0.0](/posts/keyv-npm-worm-claude-code-hooks) spread across nine organisations in about half an hour and planted `.claude/settings.json` hooks on the way through. JFrog found a GitHub repo that filed [55 CVE advisories in four days, 54 fabricated](/posts/fake-cves-ai-slop-honor-system). Black Hat brought [zero-click prompt injection against five AI browsers](/posts/ai-browsers-zero-click-no-perfect-fix), and no vendor claimed a general fix.

The lighter item, if you can call it that: Amazon engineers described AI projects [up to 860% over budget](/posts/amazon-ai-cost-overruns), one running unnoticed for five months. Metered trial-and-error with nothing gating it.