Savings Layer is part of the BiVelio platform
Public bench · reproducible

Our A/B bench, in the open. Measure us.

We publish every run: four batteries on public datasets, 10 models, and both arms metered by the provider's own usage counter. Every figure carries its n, its interval and its date. What doesn't save is here too.

batteries
4
models tried
10
models with a measurement
10
live reports
15
latest run
Sep 29, 2026

What is measured, and how

  • Two arms, the same request. A goes straight to the provider, without the layer; B goes through the layer with its default configuration. Same prompt, same sampling, same output cap.
  • The provider counts the cost. Each arm is priced from the usage the API returns on every call, at list price. Savings are 1 − cost(B) / cost(A) of the same run, not what the layer says about itself.
  • 95% intervals. Each harness computes them (resampling whole question families, Wilson for refusal rates). An interval that crosses zero means “indistinguishable from zero”, and we say so.
  • The package you install. Arm B runs the source of tag 0.4.31; its dist/index.js is byte-identical to the npm tarball's (sha256 b52f99cacf9804ecd3153e56f26ec3adcdc2510d52a8446a369ecf914e0b5731).
  • No shortcuts. Each harness watches the hosts it talks to and carries a per-model spend cap written in code. A model that could not be measured is reported as “not measured”, with the reason.

Four batteries, public datasets, pinned

BatteryWhat it checksDatasetPinned filesLicense
GSM8KMath word problems with repeats and paraphrases: how much the cache saves and whether accuracy drops.openai/grade-school-mathcommit 3101c7d5072418e28b9008a6636bde82a006892cgsm8k-test
sha256 3730d312f6e3440559ace48831e51066acaca737f6eabec99bccb9e4b3c39d14
gsm8k-train
sha256 17f347dc51477c50d4efb83959dbb7c56297aba886e5544ee2aaed3024813465
MIT
BFCLTool calls: the layer must not change a single one, and the whole catalog must arrive.ShishirPatil/gorillacommit f7cf7359b7ac615a0b294831c5ba2bc95ee4a000pinned by commitApache-2.0
XSTestRefusals: the layer must not change when the model refuses, nor serve a refusal from cache.paul-rottger/exaggerated-safetycommit d7bb5bd738c1fcbc36edd83d5e7d1b71a3e2d84dxstest_prompts.csv
sha256 11783fb294ed017473ee53c207d71f2161c7672c8d0b037501e78387f801cb5a
CC-BY-4.0
MT-BenchTwo-turn chat without repeats: the layer must not make cost, latency or quality worse.lm-sys/FastChatcommit 587d5cfa1609a43d192cedb8441cac3c17db105dquestion.jsonl
sha256 119565adbab82227089cefdb44c8d7e2cf04dc0a0ec233634c82e7d4e2a944f7
gpt-4.jsonl
sha256 f957a5bc977badb66885ec970e6cd08527845780313f0995764260e5777b9b3f
judge_prompts.jsonl
sha256 fd283293406d024f44c174b094ef48031d0687a4682fd3a56b29b138f80281b6
Apache-2.0

Measured savings per model

GSM8K, on a request stream where 30% are exact repeats and 20% paraphrases. It is the battery with repeated traffic: your savings depend on how much yours repeats. Without repeats, see “What doesn't save”.

ModelnMeasured savings [CI]From the exact cacheIn the calls (prefix)Accuracy A → BVersion · date
openai/gpt-4o-mini200 requests · 100 unique31.5% [25.7% – 36.7%]$0.0124$0.000695.0% → 96.1%Δ +1.1 pp [0.0, +2.8]0.4.31 · Sep 28, 2026
openai/gpt-5.4-mini176 requests · 88 unique31.7% [25.6% – 37.6%]$0.0734$0.000395.6% → 93.7%Δ -1.9 pp [-5.0, 0.0]0.4.31 · Sep 28, 2026
openai/gpt-5.522 requests · 11 unique33.4% [6.8% – 53.8%]$0.0790-$0.0023100.0% → 100.0%Δ 0.0 pp [0.0, 0.0]0.4.31 · Sep 28, 2026
openai/gpt-5.6-sol32 requests · 16 unique29.2% [9.5% – 46.8%]$0.0252-$0.0003100.0% → 100.0%Δ 0.0 pp [0.0, 0.0]0.4.31 · Sep 28, 2026
anthropic/claude-haiku-4-5116 requests · 58 unique32.2% [23.6% – 40.2%]$0.0773$0.000299.0% → 99.0%Δ 0.0 pp [0.0, 0.0]0.4.31 · Sep 28, 2026
anthropic/claude-opus-5-532 requests · 16 unique77.2% [67.1% – 83.9%]$0.0974$0.1391100.0% → 100.0%Δ 0.0 pp [0.0, 0.0]0.4.31 · Sep 28, 2026
anthropic/claude-sonnet-562 requests · 31 unique75.1% [69.3% – 79.6%]$0.1054$0.1313100.0% → 100.0%Δ 0.0 pp [0.0, 0.0]0.4.31 · Sep 28, 2026
anthropic/claude-sonnet-5-562 requests · 31 unique77.5% [72.0% – 81.7%]$0.0924$0.1289100.0% → 100.0%Δ 0.0 pp [0.0, 0.0]0.4.31 · Sep 28, 2026
google/gemini-3.1-pro-previewthe layer failed: requests did not go through the layer0.4.31 · Sep 28, 2026
google/gemini-3.8-flashthe layer failed: requests did not go through the layer0.4.31 · Sep 28, 2026

Where the savings come from, position by position: “from the exact cache” is what A paid on the requests the layer served without calling the provider; “in the calls” is the difference on the ones that did reach it, mostly the provider's prefix cache (plus some output variation). Together they make the total.

Accuracy is measured on the questions with a published answer; the variants with a changed number only serve to check that the layer never hands back another question's answer.

Guarantees, counted

The layer never changed a tool call

0 altered calls in 1,325 BFCL comparisons, 6 models, since 0.4.31. If the layer fails (an error in B that A did not have), it is not counted here: it is under “Bugs the bench found”.

No answer from another question since 0.4.31

0 of 198 variants (the same question with another number, hence another answer) served with another question's answer, across 8 live models and every cache configuration measured.

Offline replay of the whole GSM8K test split: 0 in 5,813 requests, in each of the 5 cache configurations.

The same refusal rate, with and without the layer

XSTest, 10 models measured: no A/B difference below p = 0.05 (McNemar, paired prompts). The layer delivered the text of 4,668 responses unaltered (0 altered).

Refusal rate per model, without the layer (A) and with it (B), on the same prompts.
ModelPromptsRefusals, safe prompts (A → B)Refusals, unsafe prompts (A → B)Version · date
openai/gpt-4o-mini4505.2% → 5.2% (n = 250, p = 1.00)81.5% → 81.5% (n = 200, p = 1.00)0.4.31 · Sep 28, 2026
openai/gpt-5.4-mini44111.0% → 9.8% (n = 245, p = 0.63)87.2% → 86.7% (n = 196, p = 1.00)0.4.31 · Sep 28, 2026
openai/gpt-5.51206.3% → 4.7% (n = 64, p = 1.00)80.0% → 80.0% (n = 55, p = 1.00)0.4.31 · Sep 29, 2026
openai/gpt-5.6-sol1992.7% → 2.7% (n = 111, p = 1.00)75.6% → 76.7% (n = 86, p = 1.00)0.4.31 · Sep 29, 2026
anthropic/claude-haiku-4-54501.6% → 2.0% (n = 250, p = 1.00)42.5% → 42.5% (n = 200, p = 1.00)0.4.31 · Sep 29, 2026
anthropic/claude-opus-5-581—67.1% → 67.1% (n = 79, p = 1.00)0.4.31 · Sep 29, 2026
anthropic/claude-sonnet-52642.7% → 2.0% (n = 147, p = 1.00)68.4% → 71.8% (n = 117, p = 0.34)0.4.31 · Sep 29, 2026
anthropic/claude-sonnet-5-53111.5% → 0.0% (n = 135, p = 0.50)64.3% → 65.3% (n = 98, p = 1.00)0.4.31 · Sep 29, 2026
google/gemini-3.1-pro-preview580.0% → 0.0% (n = 32, p = 1.00)80.8% → 84.6% (n = 26, p = 1.00)0.4.31 · Sep 29, 2026
google/gemini-3.8-flash4500.4% → 0.4% (n = 249, p = 1.00)67.8% → 66.8% (n = 199, p = 0.73)0.4.31 · Sep 29, 2026

What doesn't save

Without repeats, savings are ≈ 0%. We publish it anyway, because it is what you will see if your traffic does not repeat.

MT-Bench is two-turn chat with a fresh layer per question: there is nothing for the cache to reuse, and the prompt almost never reaches the minimum the provider requires to cache a prefix. In BFCL every tool call is different. What these two batteries measure is that the layer makes nothing worse; the cost difference left over is noise in the model's output, not the layer.

BatteryModelnMeasured savings [CI]ReadingVersion · date
BFCLopenai/gpt-4o-mini400 items0.1% [0.0% – 0.3%]indistinguishable from zero0.4.31 · Sep 28, 2026
BFCLopenai/gpt-5.4-mini400 items-0.2% [-1.0% – 0.5%]indistinguishable from zero0.4.31 · Sep 28, 2026
BFCLanthropic/claude-haiku-4-548 items0.1% [0.0% – 0.2%]indistinguishable from zero0.4.31 · Sep 29, 2026
BFCLanthropic/claude-sonnet-548 items1.5% [-1.6% – 5.8%]indistinguishable from zero0.4.31 · Sep 29, 2026
BFCLgoogle/gemini-3.1-pro-preview97 items4.8% [-6.6% – 15.6%]indistinguishable from zero0.4.31 · Sep 29, 2026
BFCLgoogle/gemini-3.8-flash324 items-10.7% [-27.8% – 3.2%]indistinguishable from zero0.4.31 · Sep 29, 2026
MT-Benchopenai/gpt-4o-mini80 questions3.3% [0.8% – 6.0%]savings · not attributable to the layer0.4.31 · Sep 28, 2026
MT-Benchopenai/gpt-5.4-mini80 questions-3.5% [-8.6% – 1.6%]indistinguishable from zero · not attributable to the layer0.4.31 · Sep 28, 2026
MT-Benchopenai/gpt-5.58 questions2.1% [-4.4% – 9.7%]indistinguishable from zero0.4.31 · Sep 29, 2026
MT-Benchopenai/gpt-5.6-sol13 questions-0.4% [-21.2% – 13.5%]indistinguishable from zero · not attributable to the layer0.4.31 · Sep 29, 2026
MT-Benchanthropic/claude-haiku-4-580 questions-2.6% [-4.6% – -0.9%]extra cost · not attributable to the layer0.4.31 · Sep 28, 2026
MT-Benchanthropic/claude-opus-5-518 questions1.1% [-2.0% – 4.9%]indistinguishable from zero0.4.31 · Sep 29, 2026
MT-Benchanthropic/claude-sonnet-523 questions1.7% [-8.9% – 12.7%]indistinguishable from zero0.4.31 · Sep 28, 2026
MT-Benchanthropic/claude-sonnet-5-537 questions-0.8% [-2.5% – 0.9%]indistinguishable from zero0.4.31 · Sep 29, 2026

Status per model and battery

Each cell shows the run on the layer's highest version. If a model could not be measured, we say why instead of leaving a gap.

ModelGSM8KBFCLXSTestMT-Bench
openai/gpt-4o-minimeasured · n = 200
0.4.31 · Sep 28, 2026
measured · n = 400
0.4.31 · Sep 28, 2026
measured · n = 450
0.4.31 · Sep 28, 2026
measured · n = 80
0.4.31 · Sep 28, 2026
openai/gpt-5.4-minimeasured · n = 176
0.4.31 · Sep 28, 2026
measured · n = 400
0.4.31 · Sep 28, 2026
measured · n = 441
0.4.31 · Sep 28, 2026
measured · n = 80
0.4.31 · Sep 28, 2026
openai/gpt-5.5measured · n = 22
0.4.31 · Sep 28, 2026
not measured: spend cap exhausted
0.4.31 · Sep 29, 2026
measured · n = 120
0.4.31 · Sep 29, 2026
measured · n = 8
0.4.31 · Sep 29, 2026
openai/gpt-5.6-solmeasured · n = 32
0.4.31 · Sep 28, 2026
not measured: the provider rejected the request in both arms
0.4.31 · Sep 29, 2026
measured · n = 199
0.4.31 · Sep 29, 2026
measured · n = 13
0.4.31 · Sep 29, 2026
anthropic/claude-haiku-4-5measured · n = 116
0.4.31 · Sep 28, 2026
measured · n = 48
0.4.31 · Sep 29, 2026
measured · n = 450
0.4.31 · Sep 29, 2026
measured · n = 80
0.4.31 · Sep 28, 2026
anthropic/claude-opus-5-5measured · n = 32
0.4.31 · Sep 28, 2026
not measured: spend cap exhausted
0.4.31 · Sep 29, 2026
measured · n = 81
0.4.31 · Sep 29, 2026
measured · n = 18
0.4.31 · Sep 29, 2026
anthropic/claude-sonnet-5measured · n = 62
0.4.31 · Sep 28, 2026
measured · n = 48
0.4.31 · Sep 29, 2026
measured · n = 264
0.4.31 · Sep 29, 2026
measured · n = 23
0.4.31 · Sep 28, 2026
anthropic/claude-sonnet-5-5measured · n = 62
0.4.31 · Sep 28, 2026
not measured: spend cap exhausted
0.4.31 · Sep 29, 2026
measured · n = 311
0.4.31 · Sep 29, 2026
measured · n = 37
0.4.31 · Sep 29, 2026
google/gemini-3.1-pro-previewthe layer failed: requests did not go through the layer
0.4.31 · Sep 28, 2026
measured · n = 97
0.4.31 · Sep 29, 2026
measured · n = 58
0.4.31 · Sep 29, 2026
not measured: provider quota exhausted
0.4.31 · Sep 29, 2026
google/gemini-3.8-flashthe layer failed: requests did not go through the layer
0.4.31 · Sep 28, 2026
measured · n = 324
0.4.31 · Sep 29, 2026
measured · n = 450
0.4.31 · Sep 29, 2026
the layer failed: requests did not go through the layer
0.4.31 · Sep 29, 2026

Bugs the bench found

The bench is not there to look good: it is there to find the bugs before you do. It found these, with the version they appeared in and the one that fixes them.

  1. The uncalibrated semantic cache served another question's answer 0.4.30 → 0.4.31

    With the semantic cache uncalibrated, a variant with another number (“three baskets” versus “two baskets”) could receive the original question's answer. Counted over the GSM8K variants.

    • Found in 0.4.30: 15 of 111. gsm8k.json
    • Fixed in 0.4.31 (#383).
    • Verified live on 0.4.31: 0 of 198.
  2. The layer trimmed the tool catalog with no signal to do so 0.4.30 → 0.4.31

    In multi-turn conversations, arm B sent fewer tools than arm A even though nothing indicated which ones were unneeded. Counted over the BFCL multi-turn conversations.

    • Found in 0.4.30: 20 of 20. bfcl.json
    • Fixed in 0.4.31 (#382).
    • Verified live on 0.4.31: 0 of 83.
  3. Some refusals written as text were cached 0.4.31 → 0.4.32

    A refusal the model writes as text (without the API's signal) could be cached, repeated and credited as savings. Counted in XSTest prompts.

  4. A Gemini content filter was cached as an answer 0.4.31 → 0.4.32

    Gemini flags its filter with its own stop reason, which the layer did not recognize as a refusal, so it could cache and repeat it. Counted in XSTest prompts.

  5. Gemini did not go through the layer 0.4.31 → 0.4.32

    The layer sent a field specific to the OpenAI API to every compatible endpoint, and Gemini rejects it: every request in arm B failed. Counted in GSM8K requests.

  6. With Gemini 3, the second turn with tools failed 0.4.31 → 0.4.32

    The layer did not return the thought signature Gemini 3 requires on each tool call, so the next turn was rejected. Counted in BFCL multi-turn continuations.

How to reproduce it

The bench's code is not public yet; the method is, and it fits in five steps. With your own provider keys you can rerun it and compare figure by figure with our reports.

  1. The dataset. Download each public dataset at the commit in the table above and check every file's sha256 before using it. If it does not match, it is not the same dataset.
  2. The layer. Install @bivelio/savings-layer from npm at the version the report names and check the sha256 of dist/index.js: every report carries its own.
  3. Two arms per prompt. A goes straight to the provider; B sends the same prompt, to the same model and with the same output cap, through the layer with its default configuration. Part of the prompts are repeated or paraphrased, as in real traffic, and the A/B order is randomised.
  4. The provider states the cost. Add up what each arm is billed according to the usage the provider returns, at its published price. Savings are 1 − cost B / cost A, with their interval.
  5. Quality, in the same run. Score both arms' answers the same way (exact answer in GSM8K, tool call in BFCL, refusal or not in XSTest, a judge in MT-Bench) and compare them with the JSON report for the same battery and version.

The reports

Every run leaves a JSON report with its design, its limitations and every measured call: usage, cost per arm and how the answer was classified. They are served from this website.

The reports are served as the bench wrote them, with one exception: data about the environment of whoever ran it (disk paths, email addresses) is removed. For XSTest no answer text is published, only its classification.

Measure us yourself

An independent measurement is worth more than ours. The bench's code is not public yet: if you want to rerun the measurement, or measure the layer with your own harness or your own traffic, write to us and we will give you what you need to do it.

Write to us at support@bivelio.com

Send an email