Explainer / From the October 9 audit

The denominator does the damage

A cheap attempt is attractive. A cheap completed task is usually the more useful thing to buy.

Two prices for the same work

In the audit’s repriced Medium-effort scenario, Sonnet costs about $2.07 per attempt. Opus costs about $2.84. Stop there and Sonnet looks comfortably cheaper.

Now include the outcomes. Sonnet passed 17 of 23 tasks; Opus passed 23 of 23. The spending on Sonnet’s six failed runs still belongs in the total.

Medium effort, frozen October 9 audit. USD, rounded.
MeasureSonnet, repricedOpus, recorded
Total spending$47.6359$65.3706
Attempts2323
Observed successes1723
Cost per attempt$2.0711$2.8422
Cost per success$2.8021$2.8422

Sonnet’s advantage narrows from roughly 27.1% per attempt to 1.4% per observed success. The arithmetic has not become hostile. The denominator has become more relevant.

Keep the failed runs on the receipt

Cost per observed success = total spending / successes

This is different from averaging only the successful runs. Throwing away the cost of failures would make the experiment look cheaper than it was.

It also isn’t a forecast of what repeated attempts will cost. These records contain one attempt per task at each effort level. They don’t tell us whether a failed task would pass on attempt two, whether a person would need to intervene, or how much that intervention would cost.

The question to take into a comparison

Ask what the denominator represents: requests sent, tasks attempted, tasks passed, or outcomes someone actually accepted. Then ask which costs were counted.

For this particular audit, the spending is an API-equivalent estimate for subscription runs, judging costs are excluded, and the models used different harness and judging-protocol versions. A 1.4% difference here is not evidence of a dependable advantage in your workload.

The useful habit is simple: put price, completion, and the definition of success in the same view. A cheap request can be an expensive way to leave a task unfinished.

Evidence & scope

This explainer uses the frozen results in the BOZOS VulcanBench cache-read audit. The benchmark and numerical exports are Morgan Linton’s work. No new models were run.

Download the source data and reproducible calculation. Figures here describe that snapshot, not current provider prices or future model reliability.