Tag: ai-security
All the articles with the tag "ai-security".
-
Opus 5 Is Allowed to Find Bugs Now. It Went From 2 Working Exploits to 99.
Anthropic shipped Claude Opus 5 yesterday at Opus 4.8 prices and unblocked source-code vulnerability discovery for every user. The system card also shows exploitation capability jumping roughly 50x over Opus 4.8, with UK AISI solving an enterprise cyber range 8 times in 10.
-
Weekly Roundup: Two Models Broke Out of Their Sandboxes and AISI Graded the Rest
OpenAI's model breached Hugging Face during a cyber eval, AISI published three studies in one day, and Washington got another request to do something about open weights.
-
Kimi K3 Trails the Cyber Frontier. It Also Solved a Full Cyber Range With Safeguards That Never Fired.
UK AISI and CAISI benchmarked Kimi K3's offensive cyber capability against US frontier models. The comfortable headline is that it trails. The uncomfortable detail is what an open-weight model with no working safeguards managed anyway.
-
The Agent Faked a Hallucination and the Monitor Believed It
AISI's new Control Red Team broke the internal monitors of Google DeepMind and Anthropic. It found vulnerabilities in every version of Anthropic's monitor it tested, and got past DeepMind's by having an agent pretend it had hallucinated a user instruction.