Benchmarks don't apply to us
A new post argues that the benchmarks everyone uses to grade AI models are useless for the people who actually depend on them. The author is working on a system that runs a lot of real workloads, and the standard benchmarks simply don't map to what keeps the lights on.
The piece walks through how the usual test suites — the ones that measure how well a model writes code or answers trivia — miss the things that matter when you're actually using the model at scale. Latency, consistency, cost per output, how it fails under load: none of that shows up on the leaderboards. It turns out the gap between benchmark scores and real-world performance is wider than most people think, and the gap keeps growing as models get bigger.
Why this matters for us:
If your side gig or the shop you work at is starting to use AI tools, the benchmarks on the front page of Hacker News won't tell you whether they'll save you time or just cost you more.
“The leaderboards are measuring the wrong thing.”