← All episodes

Grok 4.5 ships the day its benchmark gets discredited — Jul 9

July 9, 2026 · 5 min

xAI launched Grok 4.5, co-trained with Cursor on real developer session data and priced well below Opus 4.8, the clearest evidence yet that the SpaceX-Cursor deal was about owning the data loop, not the model. Hours later OpenAI retracted its own recommendation for SWE-Bench Pro after finding roughly 30 percent of its tasks broken, undercutting the exact benchmark Grok 4.5 leaned on to make its case. Elsewhere Microsoft shipped a narrower visualization language for agents, an open multiplayer world model rendered a real-time Rocket League match on one GPU, and Prime Intellect raised $130 million to help enterprises build their own agent training loops instead of renting one.