Skip to content

Benchmark Marketing: How to Read a Launch Post

4 min read · updated August 3, 2026

Launch posts rarely contain false numbers. Fabricating a benchmark score is easy to catch and expensive to be caught at, so almost nobody does it. The distance between a launch chart and the experience of using the model is created by choices that are all individually defensible — which is exactly why they are hard to see.

Start by assuming the numbers are true

This is not generosity; it is the only way to read these documents usefully. If you approach a launch post looking for a lie you will not find one, conclude the post is fine, and take away a picture that is still wrong. The productive question is never “is this true” but “what would also be true if this were misleading”.

The three places the gap opens are selection, conditions and presentation. They are worth separating because the defence against each is different.

Selection: which benchmarks appear

A model is evaluated internally on far more benchmarks than appear in the post. Choosing which ones to show is legitimate — nobody publishes forty charts — and it is also the single largest lever on the impression created.

  • The set changes between releases. If a benchmark that featured prominently last time is absent this time, the interesting question is what it would have shown. The check: pull up the previous launch post from the same source and diff the list.
  • Comparators are chosen too. A chart comparing against three named models is a chart that excluded the others. Ask which current model is missing from the axis, and whether it is the one that would have won.
  • Aggregate scores hide their composition. A suite averaging many subtests can move because of a large gain on one subtest, which may be the one least like your work. What the standard benchmarks actually test is worth knowing before reading any of them as a summary.

Conditions: how each number was produced

Two runs of the same benchmark on the same model can differ by more than the gap being celebrated, and the differences come from the harness rather than the model. The footnotes are where this lives, and the footnotes are usually accurate — they are just not read.

ConditionDescription
promptingZero-shot, few-shot, chain-of-thought, or a bespoke prompt tuned for this benchmark. A model evaluated with a tuned prompt against a comparator evaluated with a default one is not a like-for-like comparison.
samplingHow many attempts, and how the score aggregates them. A best-of-many score and a single-attempt score are different quantities with the same name.
own run vs reportedScores for competitors are often quoted from those competitors' own posts, produced under their conditions. Comparing a self-run number to a quoted one compares two harnesses.
compute spentFor reasoning models, the score depends on how many tokens the model was allowed to spend. Equal-accuracy at unequal cost is a different result from equal-accuracy.
which checkpointThe evaluated model and the served model may differ in quantisation or serving configuration, which is a documented source of quality difference.

The contamination question sits here too and is the most consequential of all: any public test set old enough to appear in a crawl may be in the training data, in which case the score measures recall rather than capability. Posts that address this directly — by reporting on a held-out or newly constructed set — are making a much stronger claim than posts that do not mention it.

Presentation: the chart itself

These are ordinary chart-reading skills, applied to a genre where they are unusually load-bearing because the underlying differences are often small.

  • Truncated axes. A bar chart starting at 70 rather than 0 turns a three-point difference into a visually decisive one. Read the axis before the bars.
  • No uncertainty. Benchmark scores are estimates from finite samples and have sampling error. A one-point gap on a few-hundred-item set is frequently within noise, and the arithmetic for how much data a difference needs is not complicated. A chart with no error bars is not claiming the differences are significant, but it is drawn to look as if it were.
  • Composite indices with unstated weights. If a headline number combines several measures, whoever chose the weights chose the ranking.
  • Price on a second axis, or absent. A quality comparison without cost is half a comparison, and the half that is missing is the one that decides most real choices.

A reading procedure

Four questions, in order. They take about two minutes and they survive any specific post.

  • What is missing? Which benchmark and which comparator would you have expected to see, and is it here?
  • Under what conditions? Find the footnote. If there is no footnote, the numbers are not comparable to anything.
  • How big, relative to noise? Convert the visual gap back into points, and ask whether that many points on that many items means anything.
  • Does any of it resemble my work? This is the one that actually matters, and it is almost always answered no. A benchmark is a proxy chosen by somebody who does not know what you are building.

The constructive alternative is unglamorous and it is the only thing that has ever worked: a small held-out set of your own examples, scored the way you care about. Fifty items from your own traffic tell you more about a model choice than every chart in the post.

Benchmark Marketing: How to Read a Launch Post · Multigrid