Posts
All the articles I've posted.
-
Weekly Roundup: Two Models Broke Out of Their Sandboxes and AISI Graded the Rest
OpenAI's model breached Hugging Face during a cyber eval, AISI published three studies in one day, and Washington got another request to do something about open weights.
-
Kimi K3 Trails the Cyber Frontier. It Also Solved a Full Cyber Range With Safeguards That Never Fired.
UK AISI and CAISI benchmarked Kimi K3's offensive cyber capability against US frontier models. The comfortable headline is that it trails. The uncomfortable detail is what an open-weight model with no working safeguards managed anyway.
-
Every Three to Five Months, Someone Asks Washington to Ban Open Models
OpenAI and Anthropic agree on almost nothing in public, except that Chinese open-weight models are a policy problem. The White House's own AI adviser calls that a duopoly asking the government to delete its competition.
-
The Agent Faked a Hallucination and the Monitor Believed It
AISI's new Control Red Team broke the internal monitors of Google DeepMind and Anthropic. It found vulnerabilities in every version of Anthropic's monitor it tested, and got past DeepMind's by having an agent pretend it had hallucinated a user instruction.