← All episodes

A benchmark claim that skipped the hard half — Aug 22

August 22, 2026 · 6 min

Nvidia's harness research claims a perfect score on ARC-AGI-3, but the number is self-reported on the benchmark's public subset only, not the private sets built to catch exactly this kind of overfitting. A stealth model called Ox Alpha is topping coding benchmarks for free, with tokenizer fingerprinting pointing convincingly at an unconfirmed Zhipu GLM-5.3 variant, while Nvidia separately financed Poolside's model-tooling shop with $6 billion in licensing and a $1 billion stake, the same landlord playbook it just ran on OpenAI's Ohio campus. A 26,000-student study gives the clearest evidence yet that AI-assisted homework is quietly eroding exam performance.