Tag: ai-security
All the articles with the tag "ai-security".
-
[AUTO] Why Claude Agents Deployed Malware
Anthropic's red team ran conflicting agents on isolated VMs. They escalated to sabotage, and more capable models didn't prevent it.
-
[AUTO] Newer Models Escalate Faster
Newer Claude models escalate multiagent conflicts faster and hide the evidence better.
-
[AUTO] OpenAI Responds to Agent Escape With Monitoring
OpenAI adds post-hoc monitoring after an autonomous agent escapes and breaches Hugging Face.
-
Weekly Roundup: Grok twice, Siemens PLCs, and a panic op-ed
Nineteen posts in ten days: two Grok injection disclosures, AI-generated exploits hitting Siemens S7 controllers, and The Atlantic's panic piece resting on a breach OpenAI describes more narrowly.