Are AI Benchmarks Reliable? How the Scores Get Gamed
AI benchmark scores get gamed by contamination, saturation, and cheating. The 2026 receipts, a trust scorecard, and how to read past any leaderboard.
Chronological archive · page 11 of 13
Independent analysis of the systems, prices and markets behind artificial intelligence.
Archive
AI benchmark scores get gamed by contamination, saturation, and cheating. The 2026 receipts, a trust scorecard, and how to read past any leaderboard.
Seven Chinese firms now ship AI accelerators, the best near NVIDIA H100 class. A fact-checked 2026 map of who makes China's GPUs and what is real.
Cohere North Mini Code is free on the API and open-weight. Here is what it really costs per task once you self-host it on a single H100.
SpaceX is buying Cursor for $60B. What changes for your bill, whether your code now trains xAI models, and the real cost per task of switching away.
GPT-5.6 launched June 26 as Sol at $5/$30, half Claude Fable 5 at $10/$50. Which flagship costs less per task, now that Fable 5 is back from suspension.
Is decentralized GPU compute (Akash, io.net, Render) cheaper than AWS? A grounded 2026 cost-per-hour comparison, plus the caveats the hype skips.
What a self-hosted LLM token really costs in 2026: cost per token across owned hardware, why memory bandwidth sets speed, and where buying beats the API.
Google ended free Gemini CLI access on June 18, 2026. The best alternatives (Claude Code, Aider, OpenCode, Antigravity) ranked by real cost per task.
Claude Fable 5 costs exactly 2x Opus 4.8 per token: $10/$50 vs $5/$25. Whether it is cheaper per task depends on loop count, not the sticker rate.