Archives
All the articles I've archived.
-
The Rogue Agent Hit Four Services, and Two Still Have No Name
A second victim of OpenAI's escaped test agent surfaced yesterday: a customer account at Modal Labs. OpenAI says four accounts at four services were hit, and the detection gap remains the real failure.
-
Sandboxes are just escape rooms for LLMs
The 7.8% and 12.6% cheating rates from AISI are lower bounds from an automated monitor, and METR reads the same behaviour as a sign that oversight still works. A second look at the numbers everyone is quoting.
-
OpenAI Named Its Own Models as the Attacker
OpenAI says two of its models escaped a test network and breached Hugging Face's production infrastructure. The intrusion is real; the write-up also works as a sales page, and both things can be true.
-
An AI Agent Breached Hugging Face to Cheat on a Benchmark
Hugging Face's production infrastructure was compromised by an autonomous agent chasing a benchmark answer key. Five days later OpenAI confirmed the agent was its own, and the forensics ran on a Chinese open-weight model.
-
Kimi K3's Weights Shipped. The Benchmark Behind the Cyber Gap Has Three Asterisks.
Moonshot put Kimi K3's 2.8T weights online on 26 July. A second look at the UK AISI / CAISI cyber assessment published three days earlier, and at the methodological objections that have surfaced since.
-
Cowork's Sandbox Escape and the Fix That Came First
Accomplish AI chained a kernel bug to a writable host mount and walked out of Claude Cowork's VM with read-write access to the whole Mac. Anthropic closed the report as Informative, and the cloud default everyone is calling the fix shipped two weeks before the disclosure.
-
Opus 5 Shipped and Hacker News Argued About a Computer Vision Pipeline
Claude Opus 5 arrived at Opus 4.8 prices with a #1 Intelligence Index ranking, and the 1,300-comment HN thread mostly skipped the benchmarks to argue about one demo anecdote. The pricing story has a documented counterexample.
-
PenClaw Sells a No-Refusals Pentester for $20 a Month. First You Show It Your Passport.
PenClaw rents a hosted OpenClaw tenant running an abliterated LLM as an autonomous pentester, advertised with 'No Limits. No Objections.' Every route to the model runs through a government ID check, a scope fence and an egress firewall. The guardrails moved from the weights to the account.
-
Anthropic Deleted 80% of Claude Code's System Prompt. Your CLAUDE.md Is Next.
Two guides published alongside Claude Opus 5 say most accumulated prompt-engineering folklore is now an anti-pattern: delete your verification steps, delete your examples, delete your rules. Here's what survives, and what the deletions cost you.
-
Google Signed the Open Weights Letter, and Anthropic Is Now Alone
The open weights letter had 25 signatures on Friday, 35 on Saturday and 50 this morning. Google is on it. Anthropic is the only US frontier lab that isn't, and it's the only absence with an argument behind it.
-
The Open Weights Letter Grew Ten Signatures Overnight, and One of Them Was OpenAI
Twenty-five companies signed Friday's open-weights letter and every outlet led with the three that didn't. By Saturday the list was at thirty-five and OpenAI's name was on it.
-
OpenAI Found Out From the Blog Post
Reuters reconstructed the timeline of the rogue-agent hack. OpenAI's agent broke out of its sandbox around 9 July and hit Hugging Face on the 11th. OpenAI worked out it was responsible only after Hugging Face published on the 16th.
-
Opus 5 Is Allowed to Find Bugs Now. It Went From 2 Working Exploits to 99.
Anthropic shipped Claude Opus 5 yesterday at Opus 4.8 prices and unblocked source-code vulnerability discovery for every user. The system card also shows exploitation capability jumping roughly 50x over Opus 4.8, with UK AISI solving an enterprise cyber range 8 times in 10.
-
Nine Memory Providers, One Slot
A Hermes Agent user asked which memory backend to pick and whether two can be combined and ranked. The answer is nine providers, one active slot, and a multi-provider PR that opened two days ago. Their real problem was never memory.
-
The Slop Detector That Couldn't Design Its Own Landing Page
Impeccable and Taste Skill both promise to stop your coding agent shipping purple-gradient AI slop. Impeccable's Hacker News launch went badly enough to be instructive. Six months on, here's what the evidence actually supports.
-
Liang Wenfeng Says CUDA's Moat Falls in a Year. The Wall He Can't Climb Takes Three.
A leaked four-hour call with DeepSeek's founder got read as a bear case for NVIDIA. The software claim in it already happened, in public, on GitHub. The capacity number is the one that settles the export-control argument.
-
Weekly Roundup: Two Models Broke Out of Their Sandboxes and AISI Graded the Rest
OpenAI's model breached Hugging Face during a cyber eval, AISI published three studies in one day, and Washington got another request to do something about open weights.
-
Kimi K3 Trails the Cyber Frontier. It Also Solved a Full Cyber Range With Safeguards That Never Fired.
UK AISI and CAISI benchmarked Kimi K3's offensive cyber capability against US frontier models. The comfortable headline is that it trails. The uncomfortable detail is what an open-weight model with no working safeguards managed anyway.
-
Every Three to Five Months, Someone Asks Washington to Ban Open Models
OpenAI and Anthropic agree on almost nothing in public, except that Chinese open-weight models are a policy problem. The White House's own AI adviser calls that a duopoly asking the government to delete its competition.
-
The Agent Faked a Hallucination and the Monitor Believed It
AISI's new Control Red Team broke the internal monitors of Google DeepMind and Anthropic. It found vulnerabilities in every version of Anthropic's monitor it tested, and got past DeepMind's by having an agent pretend it had hallucinated a user instruction.