Tag: benchmarks
All the articles with the tag "benchmarks".
-
Opus 5 Is Allowed to Find Bugs Now. It Went From 2 Working Exploits to 99.
Anthropic shipped Claude Opus 5 yesterday at Opus 4.8 prices and unblocked source-code vulnerability discovery for every user. The system card also shows exploitation capability jumping roughly 50x over Opus 4.8, with UK AISI solving an enterprise cyber range 8 times in 10.
-
The Swarm Is the Branding. One Agent in a Loop Did the Work.
Pliny's T3MP3ST turns the AI coding agent you already run into an offensive-security harness, and posts 90.1% on XBOW's own benchmark. Its own receipts say the eight-operator swarm scored none of it.
-
Twenty-Two Second Brains, and a Text File That Beats Most of Them
The AI memory market now has vendors, benchmarks, and an awesome list. It also has a Letta result showing a plain filesystem outscores the specialised memory tools, and a curated comparison written by one of the products on it.