ai_scamsAugust 17, 2026Issue #86

Your benchmarks don't apply to us

Marble is an open-source eval framework for measuring AI product quality, and its core argument is simple: the metrics people use to judge models don't work for apps built on top of them.

The post walks through why the usual benchmarks — accuracy on a test set, latency, cost per call — miss the stuff that actually determines whether a product feels good. A model might score high on a standard eval while still giving a user the wrong answer in context. Conversely, a weaker model can feel better if it's tuned to a specific task and handles edge cases cleanly.

Marble's approach is to evaluate the product, not the model. That means measuring things like whether the output actually solves the problem, how consistent it is across similar prompts, and whether it breaks in predictable ways. The framework is open source.

This is the right call. Everyone's shipping AI features and measuring them with the wrong ruler. The companies that figure out how to actually evaluate their products — not just their models — will know who wins before the benchmarks do.

Why this matters for us: the same thing is true for our apps — we need to measure whether they work for the people using them, not just whether the model scores well on a test set.

The benchmarks everyone uses measure the model, not the product. That's the wrong ruler.

marble.onl

Read the originalOpen in new tab
#ai-evals#product-quality#marble

Daily issue · no spam

Get the daily on your stoop

One short email a day — AI, tech, and what it means for our communities. Plain language, cultural lens, no Silicon Valley jargon.