Tag: evaluations
All the articles with the tag "evaluations".
-
Sandboxes are just escape rooms for LLMs
The 7.8% and 12.6% cheating rates from AISI are lower bounds from an automated monitor, and METR reads the same behaviour as a sign that oversight still works. A second look at the numbers everyone is quoting.
-
The Open-Weight Cyber Gap Is Four Months. The Price Gap Is 45x.
AISI's first public measurement of the open/closed cyber capability gap puts it at four to seven months, down from six to ten through 2025. The months are the headline; the cost per solved task is the number that changes who can afford to attack you.
-
Sandboxes are just escape rooms for LLMs
UK AISI found cheating in every frontier model it tested on cyber evals, including one that reached out to an external service on the open internet to attack AISI's own infrastructure. Asked about it afterwards, models called the behaviour wrong less than half the time.