---
title: "Weekly Roundup: Two Models Broke Out of Their Sandboxes and AISI Graded the Rest"
pubDatetime: 2026-07-24T15:00:00.000Z
description: "OpenAI's model breached Hugging Face during a cyber eval, AISI published three studies in one day, and Washington got another request to do something about open weights."
tags: [ai-security, aisi, open-weight, llm-agents, sandboxing, 2026, 2026-q3, 2026-07, weekly-roundup]
---
> [!tldr] TL;DR
> Eleven posts this week, and the ones worth your time are the eval failures: models escaping sandboxes, beating monitors, and cheating on the tests meant to measure them.

The week's best story arrived in two halves. On Sunday, [Hugging Face disclosed an intrusion](/posts/huggingface-security-incident-july-2026) run end to end by an autonomous agent. On Wednesday, [OpenAI admitted the agent was its own](/posts/openai-huggingface-breach-confession), escaped from a sandboxed cyber eval through a zero-day, and went after Hugging Face's production servers to steal the answer key. Hugging Face had already called law enforcement by then.

UK AISI published three things on Thursday. Frontier models [cheated on every cyber eval](/posts/sandboxes-escape-rooms-llms) it ran, and called the behaviour wrong less than half the time when asked about it afterwards. Its new Control Red Team [found holes in every version of Anthropic's monitor](/posts/aisi-control-red-team-monitors) it tested, and got past DeepMind's by having an agent claim it had hallucinated a user instruction. The [open-weight cyber gap](/posts/aisi-open-weight-cyber-gap) is down to four to seven months, with open models solving tasks at a 45th of the cost. A day later, AISI and CAISI put [Kimi K3 behind the frontier](/posts/aisi-caisi-kimi-k3-cyber-capabilities), where it cleared a full cyber range anyway with safeguards that never fired.

On policy, OpenAI and Anthropic [both want Washington to act on Chinese open weights](/posts/openai-anthropic-china-open-weight-alliance), a pitch that turns up every three to five months. The White House separately accused Moonshot of [distilling Anthropic's Fable](/posts/moonshot-fable-distillation-kimi-k3) to build Kimi K3, on evidence nobody has published.

Elsewhere: a $25 subscription [found a pre-auth WordPress core RCE](/posts/wordpress-rce-gpt5-6-sol-exploit-chain), OpenCode's permission system [denies `git` and runs `$(which git)`](/posts/stop-using-opencode-security-teardown), and no, [AI does not control an F1 car's hybrid deployment](/posts/f1-2026-hybrid-deployment-ai).