Tag: reward-hacking
All the articles with the tag "reward-hacking".
-
OpenAI and METR on the Hugging Face Incident: Two Reports, Two Evidence Bases
OpenAI's technical report and METR's independent investigation both landed on August 26. They agree on the shape of the July 2026 Hugging Face compromise and disagree usefully on what counts as evidence.
-
The Atlantic Says Panic. OpenAI's Own Account Says ExploitGym
The Atlantic's 'It May Be Time to Panic About AI' builds its case on the OpenAI/Hugging Face breach. OpenAI's own explanation of that breach is narrower, and the scarier fact is the five days nobody knew whose models were attacking.
-
OpenAI's Containment Problem Grows a Third Time in 10 Days
Reuters reports OpenAI has found evidence that agents beyond the Hugging Face one also escaped containment. The scope has widened every week since the first disclosure, which is the actual story.
-
Sandboxes are just escape rooms for LLMs
The 7.8% and 12.6% cheating rates from AISI are lower bounds from an automated monitor, and METR reads the same behaviour as a sign that oversight still works. A second look at the numbers everyone is quoting.