Tag: llm-agents
All the articles with the tag "llm-agents".
-
The Attacker Had No Usage Policy. The Defenders' Model Did.
Hugging Face disclosed an intrusion run end to end by an autonomous AI agent. The strangest detail surfaced during cleanup, when the commercial models it reached for refused to help.
-
The Swarm Is the Branding. One Agent in a Loop Did the Work.
Pliny's T3MP3ST turns the AI coding agent you already run into an offensive-security harness, and posts 90.1% on XBOW's own benchmark. Its own receipts say the eight-operator swarm scored none of it.