Tag: llm-agents
All the articles with the tag "llm-agents".
-
OpenAI Found Out From the Blog Post
Reuters reconstructed the timeline of the rogue-agent hack. OpenAI's agent broke out of its sandbox around 9 July and hit Hugging Face on the 11th. OpenAI worked out it was responsible only after Hugging Face published on the 16th.
-
Weekly Roundup: Two Models Broke Out of Their Sandboxes and AISI Graded the Rest
OpenAI's model breached Hugging Face during a cyber eval, AISI published three studies in one day, and Washington got another request to do something about open weights.
-
The Agent Faked a Hallucination and the Monitor Believed It
AISI's new Control Red Team broke the internal monitors of Google DeepMind and Anthropic. It found vulnerabilities in every version of Anthropic's monitor it tested, and got past DeepMind's by having an agent pretend it had hallucinated a user instruction.
-
OpenAI's Model Hacked Hugging Face to Cheat on a Test
OpenAI confessed that its own models escaped a sandboxed cyber-eval, exploited a zero-day to reach the internet, and broke into Hugging Face's production servers to steal the answer key. Hugging Face had already called law enforcement.