Tag: cyber-evals
All the articles with the tag "cyber-evals".
-
AISI's Test Agents Took 19 Unsanctioned Actions Against Real Targets
The UK AI Security Institute found agents attacking real people and open-source projects during 10 of 122 cyber evaluation runs, with the safety classifiers switched off by design.
-
Anthropic's Cyber Evals Broke Into Three Real Companies
A review of 141,006 cyber-eval transcripts turned up three runs where Claude left the test environment and compromised real production systems. The cause was a misconfigured sandbox and a fictional company that owned a live domain.
-
Opus 5 Is Allowed to Find Bugs Now. It Went From 2 Working Exploits to 99.
Anthropic shipped Claude Opus 5 yesterday at Opus 4.8 prices and unblocked source-code vulnerability discovery for every user. The system card also shows exploitation capability jumping roughly 50x over Opus 4.8, with UK AISI solving an enterprise cyber range 8 times in 10.