Tag: sandboxing
All the articles with the tag "sandboxing".
-
Two Cowork Security Reports, Two Acknowledgements, No Fix
SharedRoot walks out of Claude Cowork's Mac sandbox using a public Linux kernel bug and a writable host mount. Anthropic closed it as Informative, which is the second Cowork report in seven months to be acknowledged and left alone.
-
Sandboxes are just escape rooms for LLMs
The 7.8% and 12.6% cheating rates from AISI are lower bounds from an automated monitor, and METR reads the same behaviour as a sign that oversight still works. A second look at the numbers everyone is quoting.
-
Cowork's Sandbox Escape and the Fix That Came First
Accomplish AI chained a kernel bug to a writable host mount and walked out of Claude Cowork's VM with read-write access to the whole Mac. Anthropic closed the report as Informative, and the cloud default everyone is calling the fix shipped two weeks before the disclosure.
-
Weekly Roundup: Two Models Broke Out of Their Sandboxes and AISI Graded the Rest
OpenAI's model breached Hugging Face during a cyber eval, AISI published three studies in one day, and Washington got another request to do something about open weights.