Tag: red-teaming
All the articles with the tag "red-teaming".
-
The Agent Faked a Hallucination and the Monitor Believed It
AISI's new Control Red Team broke the internal monitors of Google DeepMind and Anthropic. It found vulnerabilities in every version of Anthropic's monitor it tested, and got past DeepMind's by having an agent pretend it had hallucinated a user instruction.
-
Weekly Roundup: One Adverb, One Agent, and Twelve Hosted Runtimes
GitHub's injection scanner fell to the word 'Additionally', a single-agent red-teaming harness beat the swarm it shipped with, and the managed-agent market got mapped.
-
The Swarm Is the Branding. One Agent in a Loop Did the Work.
Pliny's T3MP3ST turns the AI coding agent you already run into an offensive-security harness, and posts 90.1% on XBOW's own benchmark. Its own receipts say the eight-operator swarm scored none of it.