Explainer / From the October 9 audit
The denominator does the damage
A cheap attempt is attractive. A cheap completed task is usually the more useful thing to buy.
Two prices for the same work
In the audit’s repriced Medium-effort scenario, Sonnet costs about $2.07 per attempt. Opus costs about $2.84. Stop there and Sonnet looks comfortably cheaper.
Now include the outcomes. Sonnet passed 17 of 23 tasks; Opus passed 23 of 23. The spending on Sonnet’s six failed runs still belongs in the total.
| Measure | Sonnet, repriced | Opus, recorded |
|---|---|---|
| Total spending | $47.6359 | $65.3706 |
| Attempts | 23 | 23 |
| Observed successes | 17 | 23 |
| Cost per attempt | $2.0711 | $2.8422 |
| Cost per success | $2.8021 | $2.8422 |
Sonnet’s advantage narrows from roughly 27.1% per attempt to 1.4% per observed success. The arithmetic has not become hostile. The denominator has become more relevant.
Keep the failed runs on the receipt
Cost per observed success = total spending / successesThis is different from averaging only the successful runs. Throwing away the cost of failures would make the experiment look cheaper than it was.
It also isn’t a forecast of what repeated attempts will cost. These records contain one attempt per task at each effort level. They don’t tell us whether a failed task would pass on attempt two, whether a person would need to intervene, or how much that intervention would cost.
The question to take into a comparison
Ask what the denominator represents: requests sent, tasks attempted, tasks passed, or outcomes someone actually accepted. Then ask which costs were counted.
For this particular audit, the spending is an API-equivalent estimate for subscription runs, judging costs are excluded, and the models used different harness and judging-protocol versions. A 1.4% difference here is not evidence of a dependable advantage in your workload.
The useful habit is simple: put price, completion, and the definition of success in the same view. A cheap request can be an expensive way to leave a task unfinished.
Evidence & scope
This explainer uses the frozen results in the BOZOS VulcanBench cache-read audit. The benchmark and numerical exports are Morgan Linton’s work. No new models were run.
Download the source data and reproducible calculation. Figures here describe that snapshot, not current provider prices or future model reliability.