Explainer / From the October 9 audit

A lower price is not a smarter model

You can make a benchmark cheaper without making its answers any better. That distinction should survive the headline.

Change one thing

The cache-read audit takes recorded Sonnet usage and substitutes $0.10 for $0.20 per million cache-read tokens. Every other recorded cost component stays in place.

Across the 115 runs, the estimated total falls from $394.7669577 to $286.4297226. That is a meaningful difference in the price assigned to the recorded work.

No task is rerun. No failed task becomes a pass. No answer improves. This is a pricing scenario applied to an existing workload.

Scenario cost = recorded cost
              − cache-read tokens × rate difference / 1,000,000

What the result can tell you

It answers a narrow question: what would this recorded usage cost under the alternative cache-read rate, retaining the rest of the estimate?

It also shows which comparisons are sensitive to that assumption. At Medium effort, Sonnet crosses below Opus on cost per attempt. Within Sonnet, the cost ordering of the five effort levels stays the same.

These are useful findings because the changed input is explicit. Someone else can use the same data, change the assumption, and check the result.

What does not come along for free

The result does not establish a billing error. The benchmark author had already disclosed conflicting official prices and deliberately retained the CLI estimates.

It does not establish what an earlier invoice should have charged. The price check is dated October 9, 2026, and the records are API-equivalent estimates for runs made with a subscription.

It also does not isolate the quality of the underlying models. The runs used different dates, harness versions, and judging-protocol versions. Repricing leaves all of those differences in place.

Keep the assumption attached

When sharing a result like this, carry three things with it: the date of the price snapshot, the input that changed, and the inputs that stayed fixed.

“27.44% cheaper under this cache-read assumption” is a claim someone can reproduce. “27.44% more efficient” quietly asks the calculation to prove much more than it did.

Evidence & scope

This explainer uses the frozen results in the BOZOS VulcanBench cache-read audit. The benchmark and numerical exports are Morgan Linton’s work. No new models were run.

Download the source data and reproducible calculation. Figures here describe that snapshot, not current provider prices or future model reliability.