VULCANBENCH CACHE-READ PRICE COUNTERFACTUAL
Snapshot date: 2026-10-09 UTC

RESULT
Using the current documented Sonnet 5.5 cache-read price of $0.10 per million
instead of $0.20, while retaining every other recorded cost component, lowers
the 115-run Sonnet sweep from $394.7669577 to $286.4297226: $108.3372351 less,
or 27.4433%. These are API-equivalent estimates for subscription runs, not bills.

The useful comparison change is Medium. Sonnet falls from $3.0022 to $2.0711
per task, below Opus 5.5 Medium's $2.8422. Opus's original export already uses
the current $0.20/M cache-read price, so no second discount is applied to it.
Low remains more expensive than Opus; High through Max remain cheaper.

Within Sonnet, no cost ordering changes: High, Low, Medium, Extra-high, Max.
High remains cheapest at $1.7309/task. Scores and observed successes do not
change. Nothing in this audit changes or corrects the author's published record.
The author explicitly retained CLI costs when official pricing text conflicted.

SONNET 5.5 (23 runs per effort; dollars rounded for display)
Effort      Passes  Recorded/run  Current/run  Current/success  Current total
Low         15/23       2.8762       1.9917          3.0540      45.8092676
Medium      17/23       3.0022       2.0711          2.8021      47.6358673
High        20/23       2.3904       1.7309          1.9905      39.8095800
Extra-high  22/23       3.3754       2.5071          2.6211      57.6644127
Max         23/23       5.5196       4.1526          4.1526      95.5105950
Total       97/115      3.4328       2.4907          2.9529     286.4297226

OPUS 5.5 WITH PUBLISHED FALLBACK/AUXILIARY USAGE (no price change needed)
Effort      Passes  Current/run  Current/success  Current total
Low         17/23       1.6953          2.2937      38.9920962
Medium      23/23       2.8422          2.8422      65.3706232
High        22/23       3.2694          3.4181      75.1972071
Extra-high  22/23       4.1116          4.2984      94.5658126
Max         21/23       8.7480          9.5812     201.2043215
Total      105/115      4.1333          4.5270     475.3300606

Per success = all run spending / count of runs with functional == 1. It includes
failed-run spending; it is not the mean cost of successful runs or a prediction
of retry-until-success spending. At Medium, Sonnet's 27.1% per-run advantage is
only 1.4% per observed success because it passed 17/23 versus Opus's 23/23.
Sonnet Max at $4.15 (23/23) is still dearer than Opus Medium at $2.84 (23/23).
Equal success counts do not establish equal code quality or general performance.

METHOD
For every Sonnet run:
  counterfactual USD = recorded estimated_usd
                     - cache_read_tokens * (0.20 - 0.10) / 1,000,000
For Opus and its fallback/auxiliary usage the documented old and current read
prices are identical, so the cache-read-only adjustment is zero.

Current standard USD per million tokens:
Serving model       Input  Output  Cache read  Write 5m  Write 1h
Sonnet 5.5             2      10       0.10        2.5        4
Opus 5.5               4      20       0.20        5          8
Opus 4.8 fallback      5      25       0.50        6.25      10
Opus 5 auxiliary       5      25       0.50        6.25      10

The audit parses JSON numbers directly as Decimal and performs all cost
arithmetic with Decimal (50 digits for division). Finite cost sums/products are
exact; per-run, per-success and percentage divisions may repeat and use that
precision. Original JSON decimal representations are preserved, including the
last-digit artifacts created by the publisher's floating-point serialization.
Displayed seven-decimal totals suppress only those negligible artifacts.

TTL CHECKS: WHY THE PRIMARY ANSWER DOES NOT ASSUME ONE-HOUR WRITES
The export contains aggregate cache-creation tokens, not separate 5m/1h counts.
114/115 Sonnet components reconstruct within $1e-10 using the old read price
and all one-hour writes. Max freightcore is $0.2524845 lower than that formula.
It is algebraically consistent with 168,323 five-minute and 171,671 one-hour
write tokens. Those counts are inferred, not observed or independently proved.

All 115 Opus 5.5 components and all 21 auxiliary Opus 5 components reconstruct
with current prices and one-hour writes. Of 30 Opus 4.8 fallback components,
29 do; Max cellarcore is $0.2147475 lower and is algebraically consistent with
57,266 five-minute and 307,708 one-hour write tokens. Again, not observed TTLs.

All 281 serving-model components lie within their standard token-pricing 5m/1h
write envelopes. Every exceptional residual permits an integer TTL mixture.
These checks support compatibility, not proof of a billing mechanism. An
unobserved discount, adjustment or offset could also produce compatible costs.
The price-only method deliberately preserves recorded write costs and residuals.

Separately, assuming all writes are 5m versus all 1h gives a Sonnet current-rate
TOKEN-ONLY envelope of $260.4563616 to $286.6822071. This is conditional on the
listed standard prices and categories, not an unconditional bound on invoices.
It is a sensitivity analysis, not the main $286.4297226 result. Unknown per-run
TTL mixtures could change some close effort orderings; the stated unchanged
ordering refers to the fixed-recorded-write cache-read-only counterfactual.

SCOPE AND LIMITS
- Sonnet's 115 rows mean 23 tasks at five effort levels, not 115 models. Opus
  contributes another 115 rows: 230 runs and 281 serving-model components total.
- All runs count toward cost and successes. Opus High has only 22 judged rows,
  but the unjudged run remains in this audit's 23-run cost denominator.
- Public economics excludes judging and describes subscription API-equivalent
  estimates. It explicitly says no Batch, Fast or priority pricing was applied.
- Raw per-request TTL, residency/region flags, service-tier receipts, server-tool
  fees, negotiated terms, subscription allocation, taxes and invoices are not
  exposed by these aggregate records. No missing category is invented or set to
  zero as an empirical finding. The primary adjustment retains other recorded
  costs and assumes only the specified cache-read rate change, no multiplier.
- Current docs distinguish regional/data-residency, fast and batch modifiers.
  This standard-price scenario is not a claim about the user's actual charges.
- Sonnet has no fallback usage. Opus includes Opus 4.8 fallback and Opus 5
  auxiliary usage. Both are retained and checked at their own current rates.
- The models ran on different Claude Code versions and different dates. The
  judging protocol versions also differ. This is a same-suite model-and-harness
  comparison, not an isolated causal comparison of base models.
- One attempt per task per effort; no reruns, confidence test or score regrading.
  No fair ranking against Fable, GPT or other unrepriced models is asserted.

REPRODUCE (NO NETWORK, PROVIDER CALLS OR EXTRA PACKAGES)
1. Extract this ZIP into a writable directory.
2. Use Python 3.10 or newer: python3 audit.py
3. Read results/summary.json, results/model_effort.csv and results/checks.json.

The script verifies source hashes first, then 2,072 assertions in total,
covering factual rate-table consistency, run uniqueness/counts, task coverage, token sums, output sums,
component costs, pricing envelopes, algebraic TTL compatibility, counterfactual
bounds and published aggregate reconciliation. A failed assertion stops it.
CSV files are supporting machine-readable audit evidence, not a prospect list.

FILES
sources/manifest.json       Included numeric sources: URLs, UTC retrieval times,
                            byte counts and SHA-256 hashes
sources/runs.json           Original numerical Sonnet run records
sources/economics.json      Original Sonnet economics data
sources/opus-runs.json       Original numerical Opus run records
sources/opus-economics.json  Original Opus economics data
sources/current-price-facts.json  Our structured factual price table with source
                                  URL and retrieval time; no copied document
external-source-provenance.json  URLs and hashes for supporting documents whose
                                 contents are not distributed in this bundle
audit.py                   Offline deterministic calculation and assertions
results/runs.csv           All 230 runs, original and scenario cost columns
results/components.csv     All 281 components with reconstruction residuals
results/model_effort.csv   Model/effort cost and success summaries
results/comparisons.csv    Same-effort comparison before and after
results/summary.json       Exact Decimal results serialized as strings
results/checks.json        Every check and pass status

The four public data JSON files are frozen for this audit. Some numerical
records include their original explanatory metadata. Full report articles,
reproduction guides, official documentation and raw web-tool responses are not
included. External document hashes record provenance only; the offline script
does not fetch or verify absent documents. The five included source files are
hash-verified before calculation. Hashes verify integrity, not the validity of
withheld raw receipts or benchmark scores. The official price table was read
through the web tool on 2026-10-09; direct raw HTTP retrieval returned 403.

PRIMARY SOURCE LINKS
https://vulcanbench.com/benchmarks/swe-v4-sonnet55-v323.html
https://vulcanbench.com/assets/data/swe-v4-sonnet55-v323/runs.json
https://vulcanbench.com/assets/data/swe-v4-sonnet55-v323/economics.json
https://vulcanbench.com/benchmarks/swe-v4-opus55-v315.html
https://vulcanbench.com/assets/data/swe-v4-opus55-v315/runs.json
https://vulcanbench.com/assets/data/swe-v4-opus55-v315/economics.json
https://platform.claude.com/docs/en/about-claude/pricing
