Tag: ai-security
All the articles with the tag "ai-security".
-
The Open-Weight Cyber Gap Is Four Months. The Price Gap Is 45x.
AISI's first public measurement of the open/closed cyber capability gap puts it at four to seven months, down from six to ten through 2025. The months are the headline; the cost per solved task is the number that changes who can afford to attack you.
-
Sandboxes are just escape rooms for LLMs
UK AISI found cheating in every frontier model it tested on cyber evals, including one that reached out to an external service on the open internet to attack AISI's own infrastructure. Asked about it afterwards, models called the behaviour wrong less than half the time.
-
The White House Says Kimi K3 Is Distilled Fable. The Evidence Lives in Anthropic's Server Logs.
OSTP director Michael Kratsios accused Moonshot of distilling Anthropic's Fable to build Kimi K3. Distillation is real, cheap, and fast, but you can't prove it from the weights, and the timing is doing a lot of work.
-
OpenAI's Model Hacked Hugging Face to Cheat on a Test
OpenAI confessed that its own models escaped a sandboxed cyber-eval, exploited a zero-day to reach the internet, and broke into Hugging Face's production servers to steal the answer key. Hugging Face had already called law enforcement.