Our A/B bench, in the open. Measure us.
We publish every run: four batteries on public datasets, 10 models, and both arms metered by the provider's own usage counter. Every figure carries its n, its interval and its date. What doesn't save is here too.
- batteries
- 4
- models tried
- 10
- models with a measurement
- 10
- live reports
- 15
- latest run
- Sep 29, 2026
What is measured, and how
- Two arms, the same request. A goes straight to the provider, without the layer; B goes through the layer with its default configuration. Same prompt, same sampling, same output cap.
- The provider counts the cost. Each arm is priced from the
usagethe API returns on every call, at list price. Savings are 1 − cost(B) / cost(A) of the same run, not what the layer says about itself. - 95% intervals. Each harness computes them (resampling whole question families, Wilson for refusal rates). An interval that crosses zero means “indistinguishable from zero”, and we say so.
- The package you install. Arm B runs the source of tag 0.4.31; its
dist/index.jsis byte-identical to the npm tarball's (sha256b52f99cacf9804ecd3153e56f26ec3adcdc2510d52a8446a369ecf914e0b5731). - No shortcuts. Each harness watches the hosts it talks to and carries a per-model spend cap written in code. A model that could not be measured is reported as “not measured”, with the reason.
Four batteries, public datasets, pinned
| Battery | What it checks | Dataset | Pinned files | License |
|---|---|---|---|---|
| GSM8K | Math word problems with repeats and paraphrases: how much the cache saves and whether accuracy drops. | openai/grade-school-mathcommit 3101c7d5072418e28b9008a6636bde82a006892c | gsm8k-test sha256 3730d312f6e3440559ace48831e51066acaca737f6eabec99bccb9e4b3c39d14gsm8k-trainsha256 17f347dc51477c50d4efb83959dbb7c56297aba886e5544ee2aaed3024813465 | MIT |
| BFCL | Tool calls: the layer must not change a single one, and the whole catalog must arrive. | ShishirPatil/gorillacommit f7cf7359b7ac615a0b294831c5ba2bc95ee4a000 | pinned by commit | Apache-2.0 |
| XSTest | Refusals: the layer must not change when the model refuses, nor serve a refusal from cache. | paul-rottger/exaggerated-safetycommit d7bb5bd738c1fcbc36edd83d5e7d1b71a3e2d84d | xstest_prompts.csv sha256 11783fb294ed017473ee53c207d71f2161c7672c8d0b037501e78387f801cb5a | CC-BY-4.0 |
| MT-Bench | Two-turn chat without repeats: the layer must not make cost, latency or quality worse. | lm-sys/FastChatcommit 587d5cfa1609a43d192cedb8441cac3c17db105d | question.jsonl sha256 119565adbab82227089cefdb44c8d7e2cf04dc0a0ec233634c82e7d4e2a944f7gpt-4.jsonlsha256 f957a5bc977badb66885ec970e6cd08527845780313f0995764260e5777b9b3fjudge_prompts.jsonlsha256 fd283293406d024f44c174b094ef48031d0687a4682fd3a56b29b138f80281b6 | Apache-2.0 |
Measured savings per model
GSM8K, on a request stream where 30% are exact repeats and 20% paraphrases. It is the battery with repeated traffic: your savings depend on how much yours repeats. Without repeats, see “What doesn't save”.
| Model | n | Measured savings [CI] | From the exact cache | In the calls (prefix) | Accuracy A → B | Version · date |
|---|---|---|---|---|---|---|
openai/gpt-4o-mini | 200 requests · 100 unique | 31.5% [25.7% – 36.7%] | $0.0124 | $0.0006 | 95.0% → 96.1%Δ +1.1 pp [0.0, +2.8] | 0.4.31 · Sep 28, 2026 |
openai/gpt-5.4-mini | 176 requests · 88 unique | 31.7% [25.6% – 37.6%] | $0.0734 | $0.0003 | 95.6% → 93.7%Δ -1.9 pp [-5.0, 0.0] | 0.4.31 · Sep 28, 2026 |
openai/gpt-5.5 | 22 requests · 11 unique | 33.4% [6.8% – 53.8%] | $0.0790 | -$0.0023 | 100.0% → 100.0%Δ 0.0 pp [0.0, 0.0] | 0.4.31 · Sep 28, 2026 |
openai/gpt-5.6-sol | 32 requests · 16 unique | 29.2% [9.5% – 46.8%] | $0.0252 | -$0.0003 | 100.0% → 100.0%Δ 0.0 pp [0.0, 0.0] | 0.4.31 · Sep 28, 2026 |
anthropic/claude-haiku-4-5 | 116 requests · 58 unique | 32.2% [23.6% – 40.2%] | $0.0773 | $0.0002 | 99.0% → 99.0%Δ 0.0 pp [0.0, 0.0] | 0.4.31 · Sep 28, 2026 |
anthropic/claude-opus-5-5 | 32 requests · 16 unique | 77.2% [67.1% – 83.9%] | $0.0974 | $0.1391 | 100.0% → 100.0%Δ 0.0 pp [0.0, 0.0] | 0.4.31 · Sep 28, 2026 |
anthropic/claude-sonnet-5 | 62 requests · 31 unique | 75.1% [69.3% – 79.6%] | $0.1054 | $0.1313 | 100.0% → 100.0%Δ 0.0 pp [0.0, 0.0] | 0.4.31 · Sep 28, 2026 |
anthropic/claude-sonnet-5-5 | 62 requests · 31 unique | 77.5% [72.0% – 81.7%] | $0.0924 | $0.1289 | 100.0% → 100.0%Δ 0.0 pp [0.0, 0.0] | 0.4.31 · Sep 28, 2026 |
google/gemini-3.1-pro-preview | the layer failed: requests did not go through the layer | 0.4.31 · Sep 28, 2026 | ||||
google/gemini-3.8-flash | the layer failed: requests did not go through the layer | 0.4.31 · Sep 28, 2026 | ||||
Where the savings come from, position by position: “from the exact cache” is what A paid on the requests the layer served without calling the provider; “in the calls” is the difference on the ones that did reach it, mostly the provider's prefix cache (plus some output variation). Together they make the total.
Accuracy is measured on the questions with a published answer; the variants with a changed number only serve to check that the layer never hands back another question's answer.
Guarantees, counted
The layer never changed a tool call
0 altered calls in 1,325 BFCL comparisons, 6 models, since 0.4.31. If the layer fails (an error in B that A did not have), it is not counted here: it is under “Bugs the bench found”.
No answer from another question since 0.4.31
0 of 198 variants (the same question with another number, hence another answer) served with another question's answer, across 8 live models and every cache configuration measured.
Offline replay of the whole GSM8K test split: 0 in 5,813 requests, in each of the 5 cache configurations.
The same refusal rate, with and without the layer
XSTest, 10 models measured: no A/B difference below p = 0.05 (McNemar, paired prompts). The layer delivered the text of 4,668 responses unaltered (0 altered).
| Model | Prompts | Refusals, safe prompts (A → B) | Refusals, unsafe prompts (A → B) | Version · date |
|---|---|---|---|---|
openai/gpt-4o-mini | 450 | 5.2% → 5.2% (n = 250, p = 1.00) | 81.5% → 81.5% (n = 200, p = 1.00) | 0.4.31 · Sep 28, 2026 |
openai/gpt-5.4-mini | 441 | 11.0% → 9.8% (n = 245, p = 0.63) | 87.2% → 86.7% (n = 196, p = 1.00) | 0.4.31 · Sep 28, 2026 |
openai/gpt-5.5 | 120 | 6.3% → 4.7% (n = 64, p = 1.00) | 80.0% → 80.0% (n = 55, p = 1.00) | 0.4.31 · Sep 29, 2026 |
openai/gpt-5.6-sol | 199 | 2.7% → 2.7% (n = 111, p = 1.00) | 75.6% → 76.7% (n = 86, p = 1.00) | 0.4.31 · Sep 29, 2026 |
anthropic/claude-haiku-4-5 | 450 | 1.6% → 2.0% (n = 250, p = 1.00) | 42.5% → 42.5% (n = 200, p = 1.00) | 0.4.31 · Sep 29, 2026 |
anthropic/claude-opus-5-5 | 81 | — | 67.1% → 67.1% (n = 79, p = 1.00) | 0.4.31 · Sep 29, 2026 |
anthropic/claude-sonnet-5 | 264 | 2.7% → 2.0% (n = 147, p = 1.00) | 68.4% → 71.8% (n = 117, p = 0.34) | 0.4.31 · Sep 29, 2026 |
anthropic/claude-sonnet-5-5 | 311 | 1.5% → 0.0% (n = 135, p = 0.50) | 64.3% → 65.3% (n = 98, p = 1.00) | 0.4.31 · Sep 29, 2026 |
google/gemini-3.1-pro-preview | 58 | 0.0% → 0.0% (n = 32, p = 1.00) | 80.8% → 84.6% (n = 26, p = 1.00) | 0.4.31 · Sep 29, 2026 |
google/gemini-3.8-flash | 450 | 0.4% → 0.4% (n = 249, p = 1.00) | 67.8% → 66.8% (n = 199, p = 0.73) | 0.4.31 · Sep 29, 2026 |
What doesn't save
Without repeats, savings are ≈ 0%. We publish it anyway, because it is what you will see if your traffic does not repeat.
MT-Bench is two-turn chat with a fresh layer per question: there is nothing for the cache to reuse, and the prompt almost never reaches the minimum the provider requires to cache a prefix. In BFCL every tool call is different. What these two batteries measure is that the layer makes nothing worse; the cost difference left over is noise in the model's output, not the layer.
| Battery | Model | n | Measured savings [CI] | Reading | Version · date |
|---|---|---|---|---|---|
| BFCL | openai/gpt-4o-mini | 400 items | 0.1% [0.0% – 0.3%] | indistinguishable from zero | 0.4.31 · Sep 28, 2026 |
| BFCL | openai/gpt-5.4-mini | 400 items | -0.2% [-1.0% – 0.5%] | indistinguishable from zero | 0.4.31 · Sep 28, 2026 |
| BFCL | anthropic/claude-haiku-4-5 | 48 items | 0.1% [0.0% – 0.2%] | indistinguishable from zero | 0.4.31 · Sep 29, 2026 |
| BFCL | anthropic/claude-sonnet-5 | 48 items | 1.5% [-1.6% – 5.8%] | indistinguishable from zero | 0.4.31 · Sep 29, 2026 |
| BFCL | google/gemini-3.1-pro-preview | 97 items | 4.8% [-6.6% – 15.6%] | indistinguishable from zero | 0.4.31 · Sep 29, 2026 |
| BFCL | google/gemini-3.8-flash | 324 items | -10.7% [-27.8% – 3.2%] | indistinguishable from zero | 0.4.31 · Sep 29, 2026 |
| MT-Bench | openai/gpt-4o-mini | 80 questions | 3.3% [0.8% – 6.0%] | savings · not attributable to the layer | 0.4.31 · Sep 28, 2026 |
| MT-Bench | openai/gpt-5.4-mini | 80 questions | -3.5% [-8.6% – 1.6%] | indistinguishable from zero · not attributable to the layer | 0.4.31 · Sep 28, 2026 |
| MT-Bench | openai/gpt-5.5 | 8 questions | 2.1% [-4.4% – 9.7%] | indistinguishable from zero | 0.4.31 · Sep 29, 2026 |
| MT-Bench | openai/gpt-5.6-sol | 13 questions | -0.4% [-21.2% – 13.5%] | indistinguishable from zero · not attributable to the layer | 0.4.31 · Sep 29, 2026 |
| MT-Bench | anthropic/claude-haiku-4-5 | 80 questions | -2.6% [-4.6% – -0.9%] | extra cost · not attributable to the layer | 0.4.31 · Sep 28, 2026 |
| MT-Bench | anthropic/claude-opus-5-5 | 18 questions | 1.1% [-2.0% – 4.9%] | indistinguishable from zero | 0.4.31 · Sep 29, 2026 |
| MT-Bench | anthropic/claude-sonnet-5 | 23 questions | 1.7% [-8.9% – 12.7%] | indistinguishable from zero | 0.4.31 · Sep 28, 2026 |
| MT-Bench | anthropic/claude-sonnet-5-5 | 37 questions | -0.8% [-2.5% – 0.9%] | indistinguishable from zero | 0.4.31 · Sep 29, 2026 |
Status per model and battery
Each cell shows the run on the layer's highest version. If a model could not be measured, we say why instead of leaving a gap.
| Model | GSM8K | BFCL | XSTest | MT-Bench |
|---|---|---|---|---|
openai/gpt-4o-mini | measured · n = 200 0.4.31 · Sep 28, 2026 | measured · n = 400 0.4.31 · Sep 28, 2026 | measured · n = 450 0.4.31 · Sep 28, 2026 | measured · n = 80 0.4.31 · Sep 28, 2026 |
openai/gpt-5.4-mini | measured · n = 176 0.4.31 · Sep 28, 2026 | measured · n = 400 0.4.31 · Sep 28, 2026 | measured · n = 441 0.4.31 · Sep 28, 2026 | measured · n = 80 0.4.31 · Sep 28, 2026 |
openai/gpt-5.5 | measured · n = 22 0.4.31 · Sep 28, 2026 | not measured: spend cap exhausted 0.4.31 · Sep 29, 2026 | measured · n = 120 0.4.31 · Sep 29, 2026 | measured · n = 8 0.4.31 · Sep 29, 2026 |
openai/gpt-5.6-sol | measured · n = 32 0.4.31 · Sep 28, 2026 | not measured: the provider rejected the request in both arms 0.4.31 · Sep 29, 2026 | measured · n = 199 0.4.31 · Sep 29, 2026 | measured · n = 13 0.4.31 · Sep 29, 2026 |
anthropic/claude-haiku-4-5 | measured · n = 116 0.4.31 · Sep 28, 2026 | measured · n = 48 0.4.31 · Sep 29, 2026 | measured · n = 450 0.4.31 · Sep 29, 2026 | measured · n = 80 0.4.31 · Sep 28, 2026 |
anthropic/claude-opus-5-5 | measured · n = 32 0.4.31 · Sep 28, 2026 | not measured: spend cap exhausted 0.4.31 · Sep 29, 2026 | measured · n = 81 0.4.31 · Sep 29, 2026 | measured · n = 18 0.4.31 · Sep 29, 2026 |
anthropic/claude-sonnet-5 | measured · n = 62 0.4.31 · Sep 28, 2026 | measured · n = 48 0.4.31 · Sep 29, 2026 | measured · n = 264 0.4.31 · Sep 29, 2026 | measured · n = 23 0.4.31 · Sep 28, 2026 |
anthropic/claude-sonnet-5-5 | measured · n = 62 0.4.31 · Sep 28, 2026 | not measured: spend cap exhausted 0.4.31 · Sep 29, 2026 | measured · n = 311 0.4.31 · Sep 29, 2026 | measured · n = 37 0.4.31 · Sep 29, 2026 |
google/gemini-3.1-pro-preview | the layer failed: requests did not go through the layer 0.4.31 · Sep 28, 2026 | measured · n = 97 0.4.31 · Sep 29, 2026 | measured · n = 58 0.4.31 · Sep 29, 2026 | not measured: provider quota exhausted 0.4.31 · Sep 29, 2026 |
google/gemini-3.8-flash | the layer failed: requests did not go through the layer 0.4.31 · Sep 28, 2026 | measured · n = 324 0.4.31 · Sep 29, 2026 | measured · n = 450 0.4.31 · Sep 29, 2026 | the layer failed: requests did not go through the layer 0.4.31 · Sep 29, 2026 |
Bugs the bench found
The bench is not there to look good: it is there to find the bugs before you do. It found these, with the version they appeared in and the one that fixes them.
The uncalibrated semantic cache served another question's answer 0.4.30 → 0.4.31
With the semantic cache uncalibrated, a variant with another number (“three baskets” versus “two baskets”) could receive the original question's answer. Counted over the GSM8K variants.
- Found in 0.4.30: 15 of 111. gsm8k.json
- Fixed in 0.4.31 (#383).
- Verified live on 0.4.31: 0 of 198.
The layer trimmed the tool catalog with no signal to do so 0.4.30 → 0.4.31
In multi-turn conversations, arm B sent fewer tools than arm A even though nothing indicated which ones were unneeded. Counted over the BFCL multi-turn conversations.
- Found in 0.4.30: 20 of 20. bfcl.json
- Fixed in 0.4.31 (#382).
- Verified live on 0.4.31: 0 of 83.
Some refusals written as text were cached 0.4.31 → 0.4.32
A refusal the model writes as text (without the API's signal) could be cached, repeated and credited as savings. Counted in XSTest prompts.
- Found in 0.4.31: 39 of 2,824. xstest-0431-claude.jsonxstest-0431-gama-alta.jsonxstest-0431.json
- Fixed in 0.4.32 (#393, #401).
- Waiting to be re-measured live on 0.4.32.
A Gemini content filter was cached as an answer 0.4.31 → 0.4.32
Gemini flags its filter with its own stop reason, which the layer did not recognize as a refusal, so it could cache and repeat it. Counted in XSTest prompts.
- Found in 0.4.31: 6 of 508. xstest-0431-gama-alta.json
- Fixed in 0.4.32 (#402).
- Waiting to be re-measured live on 0.4.32.
Gemini did not go through the layer 0.4.31 → 0.4.32
The layer sent a field specific to the OpenAI API to every compatible endpoint, and Gemini rejects it: every request in arm B failed. Counted in GSM8K requests.
- Found in 0.4.31: 504 of 504. gsm8k-0431-gama-alta.json
- Fixed in 0.4.32 (#402).
- Waiting to be re-measured live on 0.4.32.
With Gemini 3, the second turn with tools failed 0.4.31 → 0.4.32
The layer did not return the thought signature Gemini 3 requires on each tool call, so the next turn was rejected. Counted in BFCL multi-turn continuations.
- Found in 0.4.31: 50 of 50. bfcl-0431-gama-alta.json
- Fixed in 0.4.32 (#402).
- Waiting to be re-measured live on 0.4.32.
How to reproduce it
The bench's code is not public yet; the method is, and it fits in five steps. With your own provider keys you can rerun it and compare figure by figure with our reports.
- The dataset. Download each public dataset at the commit in the table above and check every file's sha256 before using it. If it does not match, it is not the same dataset.
- The layer. Install
@bivelio/savings-layerfrom npm at the version the report names and check the sha256 ofdist/index.js: every report carries its own. - Two arms per prompt. A goes straight to the provider; B sends the same prompt, to the same model and with the same output cap, through the layer with its default configuration. Part of the prompts are repeated or paraphrased, as in real traffic, and the A/B order is randomised.
- The provider states the cost. Add up what each arm is billed according to the
usagethe provider returns, at its published price. Savings are1 − cost B / cost A, with their interval. - Quality, in the same run. Score both arms' answers the same way (exact answer in GSM8K, tool call in BFCL, refusal or not in XSTest, a judge in MT-Bench) and compare them with the JSON report for the same battery and version.
The reports
Every run leaves a JSON report with its design, its limitations and every measured call: usage, cost per arm and how the answer was classified. They are served from this website.
The reports are served as the bench wrote them, with one exception: data about the environment of whoever ran it (disk paths, email addresses) is removed. For XSTest no answer text is published, only its classification.
Measure us yourself
An independent measurement is worth more than ours. The bench's code is not public yet: if you want to rerun the measurement, or measure the layer with your own harness or your own traffic, write to us and we will give you what you need to do it.
Write to us at support@bivelio.com