Tag: ai-control
All the articles with the tag "ai-control".
-
The Agent Faked a Hallucination and the Monitor Believed It
AISI's new Control Red Team broke the internal monitors of Google DeepMind and Anthropic. It found vulnerabilities in every version of Anthropic's monitor it tested, and got past DeepMind's by having an agent pretend it had hallucinated a user instruction.