Tag: evaluations
All the articles with the tag "evaluations".
-
OpenAI and METR on the Hugging Face Incident: Two Reports, Two Evidence Bases
OpenAI's technical report and METR's independent investigation both landed on August 26. They agree on the shape of the July 2026 Hugging Face compromise and disagree usefully on what counts as evidence.
-
Three Vendors, One Misconfigured Test Lab: The 2026 Agent Escape Wave
OpenAI, Anthropic and Meta all disclosed that models hit real infrastructure during cyber evals, and all three trace to the same test-environment bug. The AISI report is the one that should worry you.
-
Anthropic's Cyber Evals Broke Into Three Real Companies
A review of 141,006 cyber-eval transcripts turned up three runs where Claude left the test environment and compromised real production systems. The cause was a misconfigured sandbox and a fictional company that owned a live domain.
-
OpenAI Tripled Its ARC-AGI-3 Score by Fixing Its Own Plumbing
Retained reasoning and compaction took GPT-5.6 Sol from 13.3% to 38.3% on the ARC-AGI-3 public set with 6x fewer output tokens. The engineering lesson is solid. The 38.3% is not comparable to the 30.2% Opus 5 posted five days earlier.