Tag: benchmarks
All the articles with the tag "benchmarks".
-
[LINK] Ramp AI Index
Ramp's monthly index of paid AI adoption and business spending.
-
[QT] Zhipu's Bug-Finder Claims Need Verification
Chinese AI company announces its GLM-5.3 model finds more vulnerabilities than Western competitors, but the claims rest entirely on unverified internal testing.
-
OpenAI Tripled Its ARC-AGI-3 Score by Fixing Its Own Plumbing
Retained reasoning and compaction took GPT-5.6 Sol from 13.3% to 38.3% on the ARC-AGI-3 public set with 6x fewer output tokens. The engineering lesson is solid. The 38.3% is not comparable to the 30.2% Opus 5 posted five days earlier.
-
Opus 5 Shipped and Hacker News Argued About a Computer Vision Pipeline
Claude Opus 5 arrived at Opus 4.8 prices with a #1 Intelligence Index ranking, and the 1,300-comment HN thread mostly skipped the benchmarks to argue about one demo anecdote. The pricing story has a documented counterexample.